An agent method and system for multi-modal data augmentation for data-constrained scenarios

CN122819296APending Publication Date: 2026-09-25ANHUI ZHIJI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610797567.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0007](1)真实数据获取困难

Benefits of technology

[0058](7)具有良好的场景扩展性和工程可实施性。本发明采用工具化智能体架构,不限定具体的生成模型、质控模型或训练模型,可根据不同数据模态、不同任务目标和不同行业场景替换或扩展执行层节点,具有较强的工程落地能力和应用扩展价值。具体而言,执行层可以根据结构化任务规格和执行查询,调用样本生成、数据增强、质量评价、语义一致性评估、标注校验、规则校验和领域任务分析等功能节点中的一种或多种。其中,规则校验可以用于检查样本格式、标签字段、目标位置、时间信息、空间属性和任务约束条件,领域任务分析可以根据应用场景适配医学影像分析、安防行为识别、遥感目标检测或工业缺陷检测等任务。因此,本发明既能够保持对不同模型和业务场景的适配能力,又能够通过统一的任务规格、执行查询、质量控制和记忆反馈机制保证数据增强流程的可追踪和可复用。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819296A_ABST
    Figure CN122819296A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of artificial intelligence, multi-modal data processing and data enhancement, and particularly relates to an agent method and system for multi-modal data enhancement in a data-restricted scenario, which comprises: a specification analysis tool generating structured task information; a task arrangement tool generating an execution query; an execution layer tool forming candidate multi-modal samples; the execution layer tool performing automatic quality control; a memory storage tool recording execution results; and the task arrangement tool generating an adaptive supplementary sampling strategy. The system comprises the specification analysis tool, the task arrangement tool, the execution layer tool and the memory storage tool. The application can change the data set construction from manual collection and static arrangement to a closed-loop data enhancement process for target task requirements, reduce the dependence on large-scale real data and manual annotation, improve the sample coverage capability of low-frequency classes, complex scenarios and data-restricted tasks, and is suitable for scenarios such as smart medicine, security monitoring, remote sensing interpretation and industrial detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to, but is not limited to, the fields of artificial intelligence, generative artificial intelligence, multimodal data processing, data augmentation and model training, and particularly relates to an intelligent agent method and system for multimodal data augmentation in data-constrained scenarios. Background Technology

[0002] With the development of artificial intelligence technology, basic models and multimodal models have become crucial technological foundations supporting tasks such as intelligent medicine, intelligent security, remote sensing intelligent interpretation, and industrial visual inspection. These models typically require pre-training or fine-tuning with large-scale, high-quality, and diverse data to achieve strong representation learning capabilities, cross-task transfer capabilities, and adaptability to complex scenarios. However, in many real-world business scenarios, the data available for model training is often limited by factors such as collection conditions, privacy compliance, manual annotation, scene coverage, sample quality, and distribution balance, leading to significant data bottlenecks in model training.

[0003] In smart medicine scenarios, training basic medical models typically requires a large amount of two-dimensional or three-dimensional medical imaging data, such as CT, MRI, X-ray, ultrasound, and pathology images, combined with textual information such as examination reports, diagnostic descriptions, lesion annotations, and clinical knowledge. Because medical data involves patient privacy, institutional compliance, and professional annotation, large-scale centralized acquisition of real-world data is difficult. High-quality supervisory information, such as lesion location, anatomical structure, and disease type, usually relies on annotation by doctors or professionals, resulting in high costs, long cycles, and difficulties in reuse. Therefore, the actual construction of basic medical models often faces problems such as insufficient data scale, scarcity of low-frequency disease samples, and significant differences across devices and institutions.

[0004] In security monitoring scenarios, environments such as construction sites, industrial parks, roads, and factories generate a large amount of surveillance video or image data. However, this data often suffers from low resolution, blurry images, poor quality at night or in inclement weather, and a scarcity of abnormal event samples. In particular, tasks such as helmet wearing, personnel intrusion, dangerous behavior identification, equipment malfunction, fire smoke, and construction violations often require a large number of abnormal samples for training. However, real abnormal events themselves occur infrequently and are difficult to collect, making it difficult to form a training dataset with sufficient coverage and reliable annotation.

[0005] Similar issues exist in other scenarios such as remote sensing and industrial inspection. For example, specific target samples are scarce in high-resolution remote sensing images, and data distribution varies significantly across different regions, seasons, lighting conditions, and sensor conditions; the number of defect samples in industrial inspection is limited, and some faults or abnormal states are difficult to repeatedly collect. These problems can lead to unstable model performance in low-frequency categories, complex scenarios, and cross-domain applications.

[0006] Current data construction methods primarily rely on passive collection of real data, manual screening, manual annotation, and static dataset organization, which struggles to meet the continuous demands of basic and multimodal models for large-scale, high-quality, and controllable data. Specifically, existing technologies suffer from at least the following problems:

[0007] (1) Difficulty in obtaining real data. In scenarios such as medicine, security, remote sensing, and industry, high-quality real data is often scattered across different institutions, different devices, or different business systems. Due to limitations such as privacy protection, data security, collection costs, and business compliance requirements, it is difficult to directly share and centrally construct a unified training set.

[0008] (2) High-quality annotation is costly. Multimodal model training usually requires not only the image or video itself, but also supervisory information such as detection boxes, segmentation masks, lesion regions, event categories, text descriptions, and report knowledge. Such annotation often relies on doctors, security personnel, remote sensing interpreters, or industry experts, which is labor-intensive, time-consuming, and difficult to unify annotation standards.

[0009] (3) Imbalanced sample distribution. Real business data usually exhibits a clear long-tail distribution, with a large number of samples in high-frequency scenarios, while there are fewer samples in low-frequency diseases, abnormal events, special targets, extreme weather, and complex environments. During the training process, the model tends to favor high-frequency categories or common scenarios, resulting in insufficient ability to identify low-frequency samples, key anomalies, and complex scenarios.

[0010] (4) Lack of consistency among multimodal data. In the actual data construction process, images, videos, 3D images, text reports, structured labels and knowledge descriptions often come from different sources, have different formats and different granularities, which can easily lead to problems such as semantic inconsistency, incomplete annotation, mismatch between images and text or missing spatiotemporal information, affecting the training effect of multimodal models.

[0011] (5) Existing data augmentation processes are mostly open-loop processes. Existing solutions typically treat data generation, data screening, quality assessment, and model training as independent linear steps, lacking a closed-loop mechanism for adaptive adjustment based on data gaps, failure samples, model feedback, and historical construction results. Therefore, the system struggles to continuously determine which categories in the current dataset are insufficient, which modalities are of poor quality, and which scenarios are still not covered, and it also struggles to automatically fill the data gaps required for model training.

[0012] (6) Manual quality control is difficult to support large-scale construction. For large-scale synthetic image, video or 3D image data, relying mainly on manual review is not only inefficient, but also difficult to guarantee long-term stable consistency standards. Especially in scenarios such as medical imaging, security anomaly events and remote sensing target recognition, data quality directly affects the model training effect. Therefore, there is an urgent need for an automated, interpretable and traceable data quality assessment mechanism.

[0013] Based on the above analysis, the urgent technical problems that need to be solved in the existing technology are:

[0014] There is an urgent need for a multimodal data augmentation agent system designed for data-constrained scenarios. This system should be able to automatically parse data requirements based on different business tasks, invoke appropriate generative models, visual models, language models, or specialized analytical models to generate multimodal samples consistent with the task objectives, and automatically perform quality assessment, filtering, recording, and iterative feedback on the generated results. This approach allows for the continuous construction of high-quality, diverse, and traceable training data even under conditions of insufficient real data, high annotation costs, or uneven sample distribution. This reduces the reliance of model training on large-scale real data and improves the adaptability and engineering application value of basic and industry-specific intelligent models in complex scenarios. Summary of the Invention

[0015] To address the problems existing in the prior art, this invention provides a multimodal data augmentation intelligent agent method and system for data-constrained scenarios, applicable to application scenarios such as smart medicine, security monitoring, remote sensing interpretation, and industrial inspection where data acquisition is difficult, annotation costs are high, or sample distribution is uneven.

[0016] This invention is implemented as follows: a multimodal data augmentation agent method for data-constrained scenarios, the method comprising:

[0017] S1: Receive data augmentation tasks, sample completion tasks, or model building tasks, and generate structured task information through the specification parsing tool;

[0018] S2: The task orchestration tool generates execution queries based on structured task information and historical feedback from memory storage tools, and schedules execution layer tools.

[0019] S3: The execution layer tools call the corresponding functional nodes based on the execution query to generate candidate visual samples and their corresponding text and structured information, forming candidate multimodal samples;

[0020] S4: The execution layer tools perform automatic quality control on candidate multimodal samples to determine whether they meet the preset quality requirements;

[0021] S5: Write the candidate multimodal samples that pass quality control, the failed samples that fail quality control, the quality score, the reason for failure, and the execution status into the memory storage tool;

[0022] S6: The task orchestration tool identifies data gaps based on historical feedback in the memory storage tool and generates an adaptive resampling strategy.

[0023] Furthermore, in S1, the structured task specification includes at least one or more of the following: data modality, application scenario, target object, target category, task type, spatial location, time range, sample format, target sample size, quality threshold, generation constraints, and output requirements: data modality describes the data type to be generated or enhanced; application scenario describes the business domain to which the data belongs; target object and target category describe the sample content that needs to be covered; task type describes the model training or business analysis goal that the data ultimately serves; spatial location and time range constrain the location, region, or time period where the samples appear; sample format constrains the output data format; target sample size and quality threshold constrain the sample size and availability; generation constraints and output requirements constrain the sample generation, annotation generation, or attribute list generation process.

[0024] Further, in S2, the task orchestration tool generates execution queries based on the structured task specifications and schedules the execution layer. The task orchestration tool determines the execution nodes, execution order, and parameter configurations required for the current task based on the structured task specifications and generates execution queries for the execution layer. The task orchestration tool generates execution queries based on structured task information and historical feedback in the memory storage tool, and schedules the execution layer tools to execute the corresponding data construction tasks. The execution query includes at least one or more of the following: task type, data modality, target category, generation constraints, quality control requirements, output format, and calling parameters. The execution layer tools call one or more of the following based on the execution query: sample generation nodes, data augmentation nodes, quality control nodes, attribute list construction nodes, rule verification nodes, and model training nodes. The task orchestration tool also dynamically selects and combines functional nodes in the execution layer tools based on the current sample status, historical failure records, quality feedback, and training requirements to form an execution path adapted to the current data augmentation task.

[0025] Further, in step S3, the execution layer tool is used to call the corresponding sample generation node or data augmentation node according to the execution query, and to generate or augment the target visual data based on one or more of the target object, target category, task type, scene constraints, spatial constraints, temporal constraints, sample format, and quality requirements in the structured task information to obtain candidate visual samples; the execution layer tool is also used to generate text and structured information corresponding to the candidate visual samples, the text and structured information including sample semantic description, structured labels, target location, spatial attributes, temporal attributes, annotation information, generation parameters, quality-related metadata, and one or more data fields used for sample retrieval, filtering, or training; the candidate visual samples and their corresponding text and structured information together constitute candidate multimodal samples.

[0026] Furthermore, in step S4, the execution layer performs automatic quality control on the generated or enhanced multimodal samples. The automatic quality control node is used to determine whether the generated or enhanced samples meet the training data requirements. The automatic quality control includes at least one or more of the following: image quality assessment, video quality assessment, structural rationality assessment, semantic consistency assessment, task relevance assessment, annotation consistency assessment, and metadata consistency assessment. The automatic quality control process can be jointly completed by a visual language model, a quality evaluation model, an object detection model, a segmentation model, a classification model, a rule constraint model, or a professional analysis model. The system can select the appropriate quality control method according to different task specifications, but their unified goal is to determine whether the sample is real and usable, whether it conforms to the task semantics, whether it meets the training requirements, and whether the label, text, or attribute information corresponding to the sample is consistent with the sample content. Automatic quality control can adopt single threshold, double threshold, or multi-dimensional scoring admission rules. Only when the sample meets the preset quality threshold requirements is it included in the training data pool. Samples that do not meet the requirements are written into the failure sample set, and the failure reason, quality score, and related execution status are recorded.

[0027] Furthermore, in step S5, the memory storage tool is used to store and update the execution results returned by the execution layer tool. The execution results include one or more of the following: candidate multimodal samples that have passed quality control, failed samples that have not passed quality control, quality scores, reasons for failure, generation parameters, task specifications, sample categories, sample distribution, execution time, and resampling records. The memory storage tool is also used to organize the execution results according to one or more indexing methods of data modality, target category, scenario conditions, quality status, and execution rounds, and to provide historical execution feedback to the task orchestration tool for data gap identification, resampling scale determination, generation parameter adjustment, and execution node scheduling.

[0028] Further, in step S6, the task orchestration tool statistically analyzes the current training data pool based on historical feedback information stored in the memory storage tool, identifying data gaps in the target task. These data gaps include, but are not limited to, insufficient target sample size, insufficient low-frequency categories, insufficient low-coverage scenarios, insufficient quality pass rate, insufficient morphological distribution, insufficient modality pairing, or insufficient text semantic coverage. For the target category, target object, or scenario condition c, the system first counts the number of samples that have already passed quality control for that category and compares it with the target sample size set in the task specifications. The sample gap can be expressed as:

[0029]

[0030] in, This represents the sample gap for category or scenario condition c. This indicates the target sample size set in the task specifications; This indicates the number of samples that have passed quality control.

[0031] The system can determine the scale of the next round of supplementary sampling by combining historical quality control pass rates. For category or scenario condition c, its historical quality control pass rate can be expressed as:

[0032]

[0033] in This represents the historical quality control pass rate for category or scenario condition c. This indicates the total number of samples already generated in this category or scenario. The task orchestration tool generates an adaptive sampling scale based on the sample gap and historical quality control pass rate.

[0034]

[0035] Indicates the number of samples that need to be generated in the next round under category or scenario condition c; This represents a very small constant to prevent the denominator from being zero; This indicates rounding up to the nearest integer.

[0036] Another object of the present invention is to provide a multimodal data augmentation intelligent system for data-constrained scenarios, the system comprising:

[0037] A specification parsing tool is used to receive data augmentation tasks, sample completion tasks, or model building tasks, and parse the tasks into structured task information that can be recognized and executed by the intelligent agent system. The structured task information includes at least one or more of the following: data modality, application scenario, target object, target category, task type, scenario constraints, spatial constraints, time constraints, sample format, target sample size, quality threshold, and output requirements. The structured task specifications can be represented in YAML, JSON, TOML, database forms, interface parameters, or other structured description forms.

[0038] The task orchestration tool is used to generate execution queries based on the structured task information and historical feedback in the memory storage tool, and to schedule the execution layer tools to execute the corresponding data augmentation tasks. The execution query includes at least one or more of the following: task type, target data modality, target category, generation constraints, quality control requirements, output format, and calling parameters. The task orchestration tool is also used to select, combine, and schedule the functional nodes in the execution layer tools based on the current sample status, historical failure records, quality feedback, and training requirements to form an execution path adapted to the current data augmentation task.

[0039] The execution layer tool is used to invoke the corresponding sample generation node or data augmentation node according to the execution query, and to generate or augment the target visual data based on one or more of the target object, target category, task type, scene constraints, spatial constraints, temporal constraints, sample format, and quality requirements in the structured task information to obtain candidate visual samples. The execution layer tool is also used to generate text and structured information corresponding to the candidate visual samples. The text and structured information includes sample semantic description, structured labels, target location, spatial attributes, temporal attributes, annotation information, generation parameters, quality-related metadata, and data used for sample retrieval, filtering, or training. One or more of the fields; the candidate visual samples and their corresponding text and structured information together constitute candidate multimodal samples; the execution layer tool is also used to call the quality control node to perform automatic quality control on the candidate multimodal samples, the automatic quality control including one or more of sample quality assessment, semantic consistency assessment, annotation consistency assessment, task relevance assessment and rule constraint verification; when the candidate multimodal sample meets the preset quality requirements, it is determined as a multimodal sample that has passed the quality control; when the candidate multimodal sample does not meet the preset quality requirements, it is determined as a failed sample, and the corresponding quality score, failure reason and execution status information are recorded;

[0040] A memory storage tool is used to store and update the execution results returned by the execution layer tools. The execution results include one or more of the following: candidate multimodal samples that have passed quality control, failed samples that have not passed quality control, quality scores, reasons for failure, generation parameters, task specifications, sample categories, sample distribution, execution time, and resampling records. The memory storage tool is also used to organize the execution results according to one or more indexing methods of data modality, target category, scenario conditions, quality status, and execution rounds, and to provide historical execution feedback to the task orchestration tool for data gap identification, resampling scale determination, generation parameter adjustment, and execution node scheduling.

[0041] Furthermore, the general interaction flow between the task orchestration tool and the execution layer tool is as follows:

[0042] (1) The task orchestration tool reads the structured task specifications and historical feedback information in the memory storage tool to determine whether the current task belongs to the initial data construction task, the supplementary sampling task, the failed sample retry task, or the model training feedback driven task.

[0043] (2) The task orchestration tool selects the corresponding execution node based on the task type, data modality, target object and quality requirements, and generates an execution query;

[0044] (3) After the execution layer completes the sample generation or enhancement, it returns the generation results, quality score, pass status and execution status to the task orchestration tool;

[0045] (4) The task scheduling tool then writes the relevant information into the memory storage tool and decides whether to continue to perform supplementary sampling based on the current number of samples, pass rate and quality distribution.

[0046] Furthermore, the memory storage tool and the resampling process are as follows: The memory storage tool records all historical execution results and organizes them according to task category, scenario category, data modality, target object, quality score, and pass status; after reading the memory storage tool, the task orchestration tool can statistically analyze the number of passed samples, the number of generated samples, the quality control pass rate, and the target sample gap for each category or scenario condition; when the number of passed samples for a certain category or scenario condition is lower than the target sample size, the task orchestration tool generates a resampling task; when the pass rate for a certain category or scenario condition is low, the task orchestration tool adjusts the generation parameters or replaces the execution model; when a certain category or scenario condition performs poorly in model training, the task orchestration tool increases its resampling priority.

[0047] Another object of the present invention is to provide a computer device including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the aforementioned multimodal data augmentation agent method for data-constrained scenarios.

[0048] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the described multimodal data augmentation agent method for data-constrained scenarios.

[0049] Another objective of this invention is to provide an information data processing terminal, which includes the aforementioned multimodal data augmentation intelligent agent system for data-constrained scenarios.

[0050] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:

[0051] First, the technical solution of this invention solves the problem that existing data augmentation methods mainly rely on manual design, random amplification, or single generation, making it difficult to continuously fill data gaps according to model training needs. Through "task specification parsing, execution layer node scheduling, automatic quality control, memory feedback, and adaptive supplementary sampling," this invention transforms the data augmentation process from static data amplification to dynamic closed-loop data construction. This allows the system not only to generate samples but also to determine sample usability, record failure reasons, identify data gaps, and adjust subsequent supplementary sampling strategies based on quality feedback and historical execution results. Compared to traditional data augmentation methods, this invention focuses not only on the amplification of the data itself but also on the consistency between generated samples and text descriptions, structured labels, target locations, task semantics, and model training objectives. Compared to a single generation model, this invention, through a proxy orchestration mechanism, organizes one or more of the following into a closed-loop execution layer: sample generation nodes, data augmentation nodes, quality evaluation nodes, semantic consistency evaluation nodes, target detection nodes, segmentation nodes, classification nodes, rule verification nodes, and domain task model nodes. This makes the data augmentation process more controllable, traceable, and engineering-feasible. The specific advantages and positive effects of this invention's technical solution are as follows:

[0052] (1) Achieving a closed-loop data augmentation process. This invention utilizes a closed-loop mechanism of "specification parsing—task orchestration—execution generation—automatic quality control—memory feedback—gap filling" to enable the data construction process to continuously iterate based on target task requirements and historical execution feedback. Specifically, the specification parsing tool transforms data augmentation tasks, sample completion tasks, or model building tasks into structured task specifications. The task orchestration tool generates execution queries based on the structured task specifications and historical feedback. The execution layer tool generates candidate multimodal samples and performs automatic quality control. The memory storage tool records samples that pass quality control, failed samples that fail quality control, quality scores, reasons for failure, generation parameters, and execution status. The task orchestration tool then identifies sample gaps and generates supplementary sampling strategies based on the above records. Through this approach, this invention avoids the problem of traditional one-off, open-loop data augmentation being unable to adapt to real business needs, transforming the data augmentation process from static amplification into a closed-loop data construction process oriented towards target sample size, quality thresholds, category coverage, and task semantic requirements.

[0053] (2) Reduced reliance on large-scale real-world data. This invention can alleviate the dependence of model training on large-scale real-world data collection and manual annotation to some extent, by controlling the generation and automatic screening of high-quality training samples in situations where real-world data is insufficient, privacy compliance is restricted, collection costs are high, or abnormal samples are scarce. This can be combined with limited real-world data for alignment training or model fine-tuning. This technology is particularly effective for task scenarios where low-frequency disease samples, abnormal event samples, rare target samples, and complex environment samples are difficult to collect sufficiently or repeatedly.

[0054] (3) Improved coverage of low-frequency categories and complex scenarios. This invention can proactively identify low-frequency categories, low-coverage areas, abnormal events, complex environments, or difficult samples based on the sample distribution, pass rate, and training feedback in the memory storage tool. It then uses a supplementary sampling strategy to specifically expand the coverage. Specifically, the memory storage tool can record the number of generated samples, the number of samples passing quality control, historical quality control pass rates, and the distribution of failure reasons under different target categories, scenario conditions, data modalities, and execution rounds. The task orchestration tool identifies the gap between the target sample size and the actual sample size based on the above statistical results and generates the next round of supplementary sampling tasks for categories with insufficient samples or low pass rates. Therefore, this invention can improve the coverage and balance of training data, reducing the risk of excessive bias towards high-frequency categories or common scenarios during model training.

[0055] (4) Improve the consistency of multimodal data. While generating images, videos, and 3D images, this invention can simultaneously generate or verify text descriptions, structured labels, bounding boxes, segmentation masks, spatial locations, and event information, thereby improving semantic consistency and training usability across different modalities. Furthermore, the automatic quality control node can perform consistency checks on candidate visual samples and their corresponding text descriptions, structured labels, target locations, spatial attributes, temporal attributes, annotation information, and quality-related metadata to determine if there are issues such as image-text mismatch, missing labels, incorrect target locations, abnormal masks, incomplete metadata, or inconsistent task semantics. Through this method, this invention can improve the usability of multimodal samples in model training, retrieval and filtering, and subsequent analysis.

[0056] (5) Reduce the burden of manual quality control and annotation. This invention uses visual language models, quality evaluation models, detection models, segmentation models, or professional analysis models to automatically review generated samples, which can reduce the workload of large-scale manual screening and expert review, and improve data construction efficiency. Specifically, the system can first filter candidate samples that obviously do not meet quality requirements, have semantic inconsistencies, have annotation errors, or do not match the scenario, and record the corresponding quality scores and reasons for failure. Human reviewers can prioritize reviewing samples that have passed automatic quality control or samples that are in the boundary score range. Therefore, this invention does not completely replace manual review, but achieves sample initial screening, failure attribution, and review priority ranking through automatic quality control, thereby reducing the burden of repetitive manual quality control and improving quality control consistency.

[0057] (6) Enhancing the data efficiency and transferability of model training. This invention expands the scale and scene coverage of training data through high-quality synthetic samples, and combines them with limited real data for alignment training or fine-tuning, enabling the model to achieve better representational ability and downstream task adaptability under data-constrained conditions. Compared with methods that rely solely on random augmentation or manual collection, this invention can proactively fill in missing categories, difficult samples, and low-coverage scenes around the target task, so that limited real data and augmented samples that have passed quality control can jointly form a data pool that is more suitable for model training. Thus, under conditions of limited real sample quantity, unbalanced category distribution, or scarcity of abnormal samples, it can improve the efficiency of training data utilization and downstream task adaptability.

[0058] (7) It has good scenario scalability and engineering feasibility. This invention adopts a tool-based intelligent agent architecture, which does not limit the specific generation model, quality control model or training model. The execution layer nodes can be replaced or extended according to different data modalities, different task objectives and different industry scenarios, which has strong engineering implementation capability and application expansion value. Specifically, the execution layer can call one or more of the following functional nodes according to the structured task specifications and execution queries: sample generation, data augmentation, quality evaluation, semantic consistency assessment, annotation verification, rule verification and domain task analysis. Among them, rule verification can be used to check sample format, label field, target location, time information, spatial attributes and task constraints. Domain task analysis can be adapted to tasks such as medical image analysis, security behavior recognition, remote sensing target detection or industrial defect detection according to the application scenario. Therefore, this invention can maintain the adaptability to different models and business scenarios, and can ensure the traceability and reusability of the data augmentation process through unified task specifications, execution queries, quality control and memory feedback mechanisms.

[0059] This invention can be widely applied to scenarios such as training basic models for smart medicine, supplementing medical image lesion samples, generating abnormal event samples for construction site monitoring, enhancing the quality of security videos, constructing remote sensing target detection data, and supplementing industrial defect samples, thereby achieving the goals of "structured data requirements, controllable sample generation, automated quality assessment, closed-loop data supplementation, and efficient model training".

[0060] Secondly, as supplementary evidence of the inventive step of the claims of this invention, it is also reflected in the following important aspects:

[0061] (1) The expected benefits and commercial value of the technical solution of this invention after transformation are as follows:

[0062] This invention utilizes an intelligent closed-loop data augmentation mechanism to continuously build high-quality, diverse, and traceable training data in scenarios where real data acquisition is difficult and manual annotation is costly. Compared to methods that rely entirely on real data collection and manual annotation, this invention reduces data construction costs, shortens model training cycles, and improves coverage of low-frequency categories and complex scenarios. Specifically, this invention clarifies data modalities, target categories, target sample sizes, and quality thresholds through structured task specifications. It uses execution-layer tools to generate, augment, and automatically control candidate samples, and uses memory storage tools to record sample states, quality scores, failure reasons, and resampling records, enabling the data construction process to be tracked, reused, and continuously optimized. In practical applications, this invention can be used for various scenarios such as basic model training in smart medicine, anomaly detection in security monitoring, remote sensing target detection, and industrial defect detection. For industry tasks requiring large amounts of training data but where real samples are difficult to obtain, this invention provides a replicable, scalable, and engineerable data augmentation solution with significant industrial application value.

[0063] (2) The technical solution of this invention fills a technical gap in the industry both domestically and internationally:

[0064] Existing data augmentation solutions typically focus on random transformations of single-modal data, artificial rule amplification, or single-cycle generation, making it difficult to simultaneously address issues such as structured representation of task requirements, quality assessment of generated samples, recording of failed samples, identification of data gaps, and subsequent resampling adjustments. While existing generative models can generate images, videos, 3D images, or text data, they generally lack the closed-loop capability to continuously assess "which samples are sufficient, which categories are still insufficient, why which samples failed, and how to fill gaps in the next round" around the target task. The technical solution of this invention differs from the aforementioned open-loop or single-model data augmentation methods. Through the collaboration of specification parsing tools, task orchestration tools, execution layer tools, and memory storage tools, the data augmentation process is organized into a closed-loop flow of "task specification parsing—execution node scheduling—candidate sample generation—automatic quality control—historical feedback recording—adaptive resampling." This flow not only generates samples but also adjusts subsequent sample construction strategies based on quality control results and historical execution records, thereby forming a continuous data construction mechanism for data-constrained scenarios. Therefore, this invention does not simply apply generative models to data augmentation, but establishes an executable feedback relationship between generation, quality control, recording, and resampling, thereby improving the controllability, traceability, and task adaptability of multimodal training data construction.

[0065] (3) The technical solution of the present invention solves a technical problem that people have long wanted to solve but have never been able to solve successfully:

[0066] Data constraints have long been a major obstacle to model performance and industry application in the training of existing artificial intelligence models. While existing generative models can generate images, videos, 3D images, or text data, it remains difficult to address whether the generated results meet specific task requirements, possess training value, cover low-frequency categories, and can continuously improve based on model failure feedback using a single generative model. This invention elevates the data augmentation process from simply "generating samples" to "continuously building usable data around model training objectives" by introducing specification parsing, orchestration control, memory feedback, and automatic quality control mechanisms. Specifically, the system can transform business-side data augmentation requirements into structured task specifications and record the quality score, failure reasons, category distribution, generation parameters, and resampling status of candidate samples during execution. The task orchestration tool then identifies sample gaps based on the above information and performs targeted resampling for low-frequency categories, low-pass-rate categories, or samples from complex scenarios. This solution establishes a closed-loop relationship between data generation, quality control, and gap identification, thereby solving the long-standing key problems of "generated data failing to truly serve model training needs" and "data augmentation lacking a continuous feedback mechanism."

[0067] (4) The technical solution of the present invention overcomes technical bias:

[0068] Existing data augmentation processes typically exhibit two common technical biases: one favors local transformations of existing real samples, such as rotation, cropping, scaling, and color perturbation, believing that data augmentation primarily involves limited perturbation of the original samples. The other favors directly calling generative models to generate samples in batches, assuming that simply increasing the number of generated samples will alleviate the problem of insufficient training data. Both approaches tend to overlook issues such as whether the generated samples satisfy the specific task semantics, whether they are consistent with the text or structured labels, whether they cover low-frequency categories and complex scenarios, and whether failed samples can provide feedback for subsequent generation strategies.

[0069] This invention overcomes the conventional approach of primarily "increasing quantity" by redefining the data augmentation process as a closed-loop data construction process oriented towards the target task. The system not only focuses on the quantity of generated samples but also uses automatic quality control to determine sample usability, records failure reasons and sample distribution through memory storage tools, and generates resampling strategies based on gaps and pass rates through task orchestration tools. Thus, this invention achieves a shift from "quantity amplification" to a technical path of "quality-controlled, semantically consistent, gap-driven, and feedback-iterative," better adapting to the model training needs of data-constrained scenarios such as medicine, security, remote sensing, and industrial inspection. Attached Figure Description

[0070] Figure 1 This is a flowchart of the intelligent agent method for multimodal data augmentation in data-constrained scenarios provided in an embodiment of the present invention;

[0071] Figure 2 This is a diagram of a multimodal data augmentation intelligent agent system architecture for data-constrained scenarios provided in this embodiment of the invention;

[0072] Figure 3 Comparison of the number of quality control samples for different anatomical regions;

[0073] Figure 4 A statistical chart showing the pass rate of automatic quality control for candidate samples from different anatomical regions. Detailed Implementation

[0074] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0075] like Figure 1 As shown, this embodiment of the invention provides a multimodal data augmentation agent method for data-constrained scenarios, the method comprising:

[0076] S1: Receive data augmentation tasks, sample completion tasks, or data construction tasks for model training, and generate structured task information through specification parsing tools;

[0077] S2: The task orchestration tool generates execution queries based on structured task information and historical feedback in the memory storage tool, and schedules the corresponding functional nodes in the execution layer tool;

[0078] S3: The execution layer tools call the corresponding functional nodes based on the execution query to generate or enhance candidate visual samples and their corresponding text and structured information, forming candidate multimodal samples;

[0079] S4: The execution layer tools perform automatic quality control on candidate multimodal samples to determine whether they meet the preset quality requirements;

[0080] S5: Write the candidate multimodal samples that pass quality control, the failed samples that fail quality control, the quality score, the reason for failure, the generation parameters, and the execution status into the memory storage tool;

[0081] S6: The task orchestration tool identifies data gaps based on historical feedback in the memory storage tool and generates an adaptive resampling strategy.

[0082] In step S1, the specification parsing tool receives data augmentation requests, sample completion requests, or data construction requests for model training from user input, business systems, model training processes, or dataset construction processes, and parses these requests into structured task specifications. The structured task specifications include at least one or more of the following: data modality, application scenario, target object, target category, task type, spatial location, time range, sample format, target sample size, quality threshold, quality control requirements, generation constraints, and output requirements. Specifically, the data modality describes the data type to be generated or augmented; the application scenario describes the business domain to which the data belongs; the target object and target category describe the sample content that needs to be covered; the task type describes the model training or business analysis objective that the data ultimately serves; the spatial location constrains the location of the target in images, videos, 3D images, anatomical regions, geographical regions, or industrial components; the time range constrains video clips, event occurrence periods, data collection periods, or follow-up periods; the sample format constrains the output data format; the target sample size constrains the sample size; and the quality threshold and quality control requirements constrain whether the candidate samples meet the requirements for using training data. Generate constraints and output requirements for processes that constrain sample generation, augmentation, annotation generation, structured attribute generation, or metadata output.

[0083] In step S2, the task orchestration tool determines the execution nodes, execution order, and parameter configurations required for the current task based on the structured task specifications and historical feedback from the memory storage tool. It then generates an execution query for the execution layer tools to schedule the corresponding data construction tasks. The execution query includes at least one or more of the following: task type, data modality, target category, generation constraints, quality control requirements, output format, and calling parameters. The execution layer tools, based on the execution query, call one or more of the following: sample generation nodes, data augmentation nodes, quality control nodes, structured attribute and metadata generation nodes, rule validation nodes, and training data organization nodes, to complete multimodal sample generation, automatic quality control, sample recording, and training data organization. The task orchestration tool is also used to dynamically select and combine functional nodes in the execution layer tools based on the current sample status, historical failure records, quality feedback, and training requirements to form an execution path adapted to the current data augmentation task.

[0084] In step S3, the execution layer tool, based on the execution query issued by the task orchestration tool, calls one or more of the following: sample generation nodes, data augmentation nodes, annotation generation nodes, and structured attribute and metadata generation nodes, to generate candidate visual samples and their corresponding text and structured information. The text and structured information include one or more of the following: sample semantic description, structured labels, target location, spatial attributes, temporal attributes, annotation information, generation parameters, metadata for quality control, and data fields for sample retrieval, filtering, or training; the candidate visual samples and their corresponding text and structured information together constitute candidate multimodal samples.

[0085] In step S4, the execution layer tool performs automatic quality control on the generated or enhanced multimodal samples. The automatic quality control node is used to determine whether the generated or enhanced candidate multimodal samples meet preset quality requirements and training data usage requirements. The automatic quality control includes at least one or more of the following: image quality assessment, video quality assessment, structural rationality assessment, semantic consistency assessment, task relevance assessment, annotation consistency assessment, and metadata consistency assessment. The automatic quality control process can be completed by one or more of the following: visual language model, quality evaluation model, object detection model, segmentation model, classification model, rule verification module, or domain task model. The domain task model includes one or more of the following: medical image analysis model, security behavior recognition model, remote sensing object detection model, or industrial defect detection model. The system can select appropriate quality control methods according to different task specifications, but their unified goal is to determine whether the samples meet preset quality requirements, conform to task semantics, meet training data usage requirements, and whether the labels, text descriptions, or structured attribute information corresponding to the samples are consistent with the sample content. Furthermore, automatic quality control can employ single-threshold, double-threshold, or multi-dimensional scoring admission rules. Only samples that meet the preset quality threshold or multidimensional scoring admission rules are included in the training data pool; samples that do not meet the requirements are written into the failure sample set, and the reasons for failure, quality scores and related execution status are recorded.

[0086] In step S5, the memory storage tool is used to record the historical state and feedback information of the data augmentation agent during execution. The memory storage tool is used to store and update the execution results returned by the execution layer tools. These execution results include one or more of the following: multimodal samples that passed quality control, failed samples that did not pass quality control, quality scores, failure reasons, generation parameters, task specifications, sample categories, sample distribution, execution time resampling records, quality control pass rate, and sample gaps. The memory storage tool is also used to organize the execution results according to one or more indexing methods among data modality, target category, scenario conditions, quality status, and execution rounds. Furthermore, by statistically analyzing the number of passed samples, the number of generated samples, the distribution of failure reasons, and the quality control pass rate, it provides historical execution feedback to the task orchestration tool for data gap identification, resampling scale determination, generation parameter adjustment, and execution node scheduling.

[0087] In step S6, the task orchestration tool statistically analyzes the current training data pool based on historical feedback information stored in the memory storage tool to identify data gaps in the target task. These data gaps include, but are not limited to, insufficient target sample size, insufficient low-frequency categories, insufficient low-coverage scenarios, insufficient quality control pass rate, insufficient morphological distribution, insufficient modality pairing, or insufficient text semantic coverage. For a target category, target object, or scenario condition c, the system first counts the number of samples that have already passed quality control for that category and compares it with the target sample size set in the task specifications. The sample gap can be expressed as:

[0088]

[0089] in, Indicates the sample gap for category or scenario condition c; This indicates the target sample size set in the task specifications; This indicates the number of samples that have passed quality control. When If this occurs, it indicates that the conditions for that category or scenario have not yet met the requirements for constructing the target sample, and further sampling is needed.

[0090] To avoid repeatedly generating samples in fixed quantities, which could lead to a prolonged lack of replacements for low-pass-rate categories, the system can determine the scale of the next round of supplementary sampling by combining historical quality control pass rates. For category or scenario condition c, its historical quality control pass rate can be expressed as:

[0091]

[0092] in This indicates the historical quality control pass rate for category or scenario condition c; This indicates the total number of samples that have been generated in this category or scenario condition.

[0093] The task orchestration tool generates an adaptive resampling size based on the sample gap and historical quality control pass rate:

[0094]

[0095] Indicates the number of samples that need to be generated in the next round under category or scenario condition c; This represents a very small constant to prevent the denominator from being zero; This indicates rounding up. Using this method, when the historical pass rate for a certain category or scenario condition is low, the system will automatically increase the number of samples generated for that category or scenario condition to ensure that the final number of samples passing quality control is close to the target requirement.

[0096] The adaptive supplementation strategy may also include adjusting the number of samples generated, modifying generation conditions, changing prompts or structured constraints, raising or lowering the quality threshold, switching the execution model, increasing the proportion of samples of specific categories, and adding difficult or boundary samples. Through this step, the data augmentation process is no longer a one-time generation, but can be continuously iterated around the model training objectives, data gaps, and historical quality control feedback.

[0097] like Figure 2 As shown, this embodiment of the invention provides a multimodal data augmentation intelligent agent system for data-constrained scenarios. The system includes a specification parsing tool, a task orchestration tool, an execution layer tool, and a memory storage tool.

[0098] The specification parsing tool is used to receive data augmentation tasks, sample completion tasks, or data construction tasks for model training, and parse these tasks into structured task specifications that can be recognized and executed by the intelligent agent system. The structured task specifications include at least one or more of the following: data modality, application scenario, target object, target category, task type, scenario constraints, spatial constraints, temporal constraints, sample format, target sample size, quality threshold, quality control requirements, generation constraints, and output requirements.

[0099] The task orchestration tool generates execution queries based on the structured task specifications and historical feedback stored in the memory storage tool, and schedules the execution layer tools to execute the corresponding data augmentation tasks. The execution query includes at least one or more of the following: task type, target data modality, target category, generation constraints, quality control requirements, output format, and calling parameters. The task orchestration tool also selects, combines, and schedules functional nodes in the execution layer tools based on the current sample status, historical failure records, quality feedback, sample gaps, and training requirements to form an execution path adapted to the current data augmentation task.

[0100] The execution layer tool is used to invoke the corresponding sample generation node or data augmentation node according to the execution query, and to generate or augment the target visual data based on one or more of the target object, target category, task type, scene constraints, spatial constraints, temporal constraints, sample format, and quality requirements in the structured task specification, to obtain candidate visual samples. The execution layer tool is also used to generate text and structured information corresponding to the candidate visual samples. The text and structured information includes one or more of the following: sample semantic description, structured labels, target location, spatial attributes, temporal attributes, annotation information, generation parameters, metadata for quality control, and data fields for sample retrieval, filtering, or training. The candidate visual samples and their corresponding text and structured information together constitute candidate multimodal samples. Furthermore, the execution layer tool is also used to invoke quality control nodes to perform automatic quality control on the candidate multimodal samples. This automatic quality control includes one or more of the following: sample quality assessment, semantic consistency assessment, annotation consistency assessment, task relevance assessment, and rule verification. When a candidate multimodal sample meets preset quality requirements, it is identified as a multimodal sample that has passed quality control. When a candidate multimodal sample does not meet preset quality requirements, it is identified as a failed sample, and the corresponding quality score, failure reason, and execution status information are recorded. Specific functional nodes or models in the execution layer tool can be replaced or extended according to different business scenarios; therefore, this invention does not limit itself to a fixed model structure, fixed data modality, or fixed industry task.

[0101] The memory storage tool is used to store and update the execution results returned by the execution layer tools. These results include one or more of the following: multimodal samples that passed quality control, failed samples that did not pass quality control, quality scores, reasons for failure, generation parameters, task specifications, sample categories, sample distribution, quality control pass rate, sample gaps, execution time, and resampling records. The memory storage tool is also used to organize the execution results according to one or more indexing methods based on data modality, target category, scenario conditions, quality status, and execution rounds, and to provide historical execution feedback to the task orchestration tool for data gap identification, resampling scale determination, generation parameter adjustment, and execution node scheduling.

[0102] The specification parsing tool receives user-inputted data augmentation requirements, sample completion requirements, data requirements for model training, or business task requirements, and breaks these requirements down into structured task specifications. These structured task specifications can be represented using YAML, JSON, TOML, database forms, interface parameters, or other structured descriptive formats. In this way, the specification parsing tool can convert natural language business requirements or system interface requirements into task representations that are callable at the execution layer, schedulable by task orchestration tools, and traceable by memory storage tools.

[0103] The general interaction flow between the task orchestration tool and the execution layer tool is as follows: The task orchestration tool first reads the structured task specifications and historical feedback information from the memory storage tool to determine whether the current task belongs to an initial data construction task, a supplementary sampling task, a failed sample retry task, or a data completion task after receiving model training feedback. Subsequently, the task orchestration tool selects the corresponding execution node based on the task type, data modality, target object, target category, sample gap, and quality requirements, and generates an execution query. After completing sample generation, enhancement, and automatic quality control, the execution layer tool returns the generation results, quality score, pass status, and execution status to the task orchestration tool. The task orchestration tool then writes the relevant information into the memory storage tool and decides whether to continue supplementary sampling based on the current sample quantity, pass rate, and quality distribution. Through the above interaction flow, this invention can achieve an automated closed loop from task parsing, model invocation, result evaluation to feedback updates.

[0104] The automatic quality control first receives generated or enhanced candidate multimodal samples, which include one or more of the following: images, videos, 3D images, text descriptions, bounding boxes, masks, structured attributes, or metadata. Subsequently, the automatic quality control scores or judges the samples from multiple dimensions. These dimensions include, but are not limited to, sample quality, structural rationality, semantic consistency, task relevance, annotation accuracy, modality matching, and metadata integrity. When a sample's overall score or key dimension score reaches a preset threshold, the sample is added to the quality control sample set; otherwise, the sample is added to the failed sample set, and the reason for failure is recorded. Reasons for failure include, but are not limited to, insufficient image quality, semantic inconsistency, missing targets, annotation errors, scene mismatch, video discontinuity, or structural anomalies.

[0105] The memory storage tool and the resampling process are as follows: The memory storage tool records all historical execution results and organizes them according to task category, scenario category, data modality, target object, quality score, and pass status. After reading the memory storage tool, the task orchestration tool can statistically analyze the number of passed samples, the number of generated samples, the quality control pass rate, and the target sample gap for each category or scenario condition. When the number of passed samples for a certain category or scenario condition is lower than the target sample size, the task orchestration tool generates a resampling task; when the pass rate for a certain category or scenario condition is low, the task orchestration tool adjusts the generation parameters or replaces the execution model; when the external model training or validation results show that a certain category or scenario condition performs poorly, the task orchestration tool can use this result as feedback information to increase its resampling priority.

[0106] This process makes the data augmentation process of the present invention no longer a one-time generation, but a continuous iteration based on the data construction status and model training feedback.

[0107] Figure 3 The comparison of the number of quality-controlled samples across different anatomical regions is shown. In the medical image data enhancement embodiment, the system constructs candidate 3D CT samples around multiple anatomical regions and forms a training data pool through automatic quality control and adaptive supplementation. After supplementation, each of the five anatomical regions—head, chest, abdomen, pelvis, and lower limbs—reached 2200 quality-controlled samples, resulting in a total of 11000 quality-controlled 3D CT samples. This result demonstrates that the task orchestration tool can supplement different categories based on sample gaps in the memory storage tool and historical quality control feedback, thereby improving the balance of sample distribution.

[0108] In the medical image data enhancement embodiment, the execution layer generated a total of 11,159 candidate 3D CT samples. After automatic quality control, 11,000 samples were retained, with an overall quality control pass rate of 98.58%. The quality control pass rates for different anatomical regions ranged from 96.62% to 99.91%, indicating that the automatic quality control node can differentiate the generation quality of different categories of samples and write samples that failed quality control, their quality scores, and reasons for failure into a memory storage tool, providing a basis for determining the subsequent resampling scale and adjusting generation parameters.

[0109] head 2222 2200 99.01% Chest 2246 2200 97.95% abdomen 2202 2200 99.91% pelvic cavity 2212 2200 99.46% lower limbs 2277 2200 96.62%

[0110] Table 1: Statistical table of automatic quality control pass rate for different categories of candidate samples

[0111] Figure 4 The figure shows the automatic quality control pass rate of candidate 3D CT samples from different anatomical regions. The horizontal axis represents five anatomical regions: head, chest, abdomen, pelvis, and lower limbs, while the vertical axis represents the quality control pass rate. As can be seen from the figure, the pass rate for each region is generally high, all above 96%, with the abdomen having the highest pass rate at 99.91%, followed by the pelvis at 99.46%, the head at 99.01%, the chest at 97.95%, and the lower limbs at a relatively low 96.62%. This figure demonstrates that the system can automatically screen samples generated from different regions and identify quality differences based on the quality control results, providing a basis for subsequent resampling and parameter optimization.

[0112] In the description of this invention, unless otherwise stated, "a plurality of" means two or more; the terms "upper," "lower," "left," "right," "inner," "outer," "front end," "rear end," "head," "tail," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0113] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.

[0114] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A multimodal data augmentation agent method for data-constrained scenarios, characterized in that, The method specifically includes: S1: Receive data augmentation tasks, sample completion tasks, or data construction tasks for model training, and generate structured task specifications through the specification parsing tool; S2: The task orchestration tool generates execution queries based on structured task information and historical feedback in the memory storage tool, and schedules the corresponding functional nodes in the execution layer tool; S3: The execution layer tools call the corresponding functional nodes based on the execution query to generate candidate visual samples and their corresponding text and structured information, forming candidate multimodal samples; S4: The execution layer tools perform automatic quality control on candidate multimodal samples to determine whether they meet the preset quality requirements; S5: Write the candidate multimodal samples that pass quality control, the failed samples that fail quality control, the quality score, the reason for failure, the execution status, and the generation parameters into the memory storage tool; S6: The task orchestration tool identifies data gaps based on historical feedback in the memory storage tool and generates an adaptive resampling strategy.

2. The multimodal data augmentation agent method for data-constrained scenarios as described in claim 1, characterized in that, In step S1, the structured task specification includes at least one or more of the following: data modality, application scenario, target object, target category, task type, spatial location, time range, sample format, target sample size, quality threshold, quality control requirements, generation constraints, and output requirements. Specifically, the data modality describes the data type to be generated or enhanced; the application scenario describes the business domain to which the data belongs; the target object and target category describe the sample content that needs to be covered; the task type describes the model training or business analysis objective that the data ultimately serves; the spatial location constrains the location of the target in images, videos, 3D images, anatomical regions, geographical regions, or industrial components; the time range constrains video clips, event occurrence periods, data collection periods, or follow-up periods; the sample format constrains the output data format; the target sample size constrains the sample size; the quality threshold and quality control requirements constrain whether candidate samples meet the requirements for training data usage; and the generation constraints and output requirements constrain the sample generation, enhancement processing, annotation generation, structured attribute generation, or metadata output processes.

3. The multimodal data augmentation agent method for data-constrained scenarios as described in claim 1, characterized in that, In step S2, the task orchestration tool determines the execution nodes, execution order, and parameter configurations required for the current task based on the structured task specifications and historical feedback from the memory storage tool. It then generates an execution query for the execution layer tools to schedule the corresponding data augmentation task. The execution query includes at least one or more of the following: task type, data modality, target category, generation constraints, quality control requirements, output format, and calling parameters. The execution layer tools, based on the execution query, call one or more of the following: sample generation nodes, data augmentation nodes, quality control nodes, structured attribute and metadata generation nodes, rule validation nodes, and training data organization nodes. The task orchestration tool also dynamically selects and combines functional nodes in the execution layer tools based on the current sample status, historical failure records, quality feedback, sample gaps, and training requirements to form an execution path adapted to the current data augmentation task.

4. The multimodal data augmentation agent method for data-constrained scenarios as described in claim 1, characterized in that, In step S3, the execution layer tool is used to invoke one or more of the following based on the execution query: sample generation node, data augmentation node, annotation generation node, and structured attribute and metadata generation node. Based on one or more of the following in the structured task specifications: target object, target category, task type, scene constraints, spatial constraints, temporal constraints, sample format, and quality requirements, the tool generates or enhances target visual data such as images, videos, 3D images, remote sensing images, or industrial images to obtain candidate visual samples. The execution layer tool is also used to generate text and structured information corresponding to the candidate visual samples. This text and structured information includes one or more of the following: sample semantic description, structured labels, target location, spatial attributes, temporal attributes, annotation information, generation parameters, metadata for quality control, and data fields for sample retrieval, filtering, or training. The candidate visual samples and their corresponding text and structured information together constitute candidate multimodal samples.

5. The multimodal data augmentation agent method for data-constrained scenarios as described in claim 1, characterized in that, In step S4, the execution layer tool performs automatic quality control on the generated or enhanced candidate multimodal samples. The automatic quality control node is used to determine whether the generated or enhanced candidate multimodal samples meet the preset quality requirements and training data usage requirements. The automatic quality control includes at least one or more of the following: image quality assessment, video quality assessment, structural rationality assessment, semantic consistency assessment, task relevance assessment, annotation consistency assessment, and metadata consistency assessment. The automatic quality control process can be completed by one or more of the following: visual language model, quality evaluation model, object detection model, segmentation model, classification model, rule verification module, or domain task model. The system can select the appropriate quality control method according to different task specifications, but the unified goal is to determine whether the sample meets the preset quality requirements, whether it conforms to the task semantics, whether it meets the training data usage requirements, and whether the label, text description, or structured attribute information corresponding to the sample is consistent with the sample content. Automatic quality control can adopt single threshold, double threshold, or multidimensional scoring admission rules. Only when a sample meets the preset quality threshold or multidimensional scoring admission rules is it included in the training data pool; samples that do not meet the requirements are written into the failure sample set, and the failure reason, quality score, and related execution status are recorded.

6. The multimodal data augmentation agent method for data-constrained scenarios as described in claim 1, characterized in that, In step S5, the memory storage tool is used to store and update the execution results returned by the execution layer tool. These execution results include one or more of the following: multimodal samples that passed quality control, failed samples that did not pass quality control, quality scores, failure reasons, generation parameters, task specifications, sample categories, sample distribution, quality control pass rate, sample gaps, execution time, and resampling records. The memory storage tool is also used to organize the execution results according to one or more indexing methods based on data modality, target category, scenario conditions, quality status, and execution rounds. It provides historical execution feedback to the task orchestration tool by statistically analyzing the number of generated samples, the number of passed samples, the distribution of failure reasons, and the quality control pass rate. This feedback is used for identifying data gaps, determining the resampling scale, adjusting generation parameters, and scheduling execution nodes. For new tasks with the same data modality, the same target category, or similar scenario conditions, the historical quality control pass rate, the distribution of failure reasons, and the effective generation parameters in the memory storage tool are used as a reference for the task orchestration tool to initialize execution nodes, generation parameters, or the resampling scale.

7. The multimodal data augmentation agent method for data-constrained scenarios as described in claim 1, characterized in that, In step S6, the task orchestration tool statistically analyzes the current training data pool based on the historical feedback information stored in the memory storage tool, identifying data gaps in the target task. These data gaps include, but are not limited to, insufficient target sample size, insufficient low-frequency categories, insufficient low-coverage scenarios, insufficient quality control pass rate, insufficient morphological distribution, insufficient modality pairing, or insufficient text semantic coverage. For the target category, target object, or scenario condition c, the system first counts the number of samples that have passed quality control for that category and compares it with the target sample size set in the task specifications. The sample gap can be expressed as: ; in, This represents the sample gap for category or scenario condition c. This indicates the target sample size set in the task specifications. This indicates the number of samples that have passed quality control. The system can determine the scale of the next round of supplementary sampling by combining historical quality control pass rates. For category or scenario condition c, its historical quality control pass rate can be expressed as: ; in This represents the historical quality control pass rate for category or scenario condition c. This indicates the total number of samples already generated in this category or scenario. The task orchestration tool generates an adaptive supplementation sampling scale based on the sample gap and historical quality control pass rate. ; Indicates the number of samples that need to be generated in the next round under category or scenario condition c; This represents a very small constant to prevent the denominator from being zero; This indicates rounding up. When the pass rate of quality control for a certain category or scenario condition is lower than the preset threshold for multiple consecutive rounds, the task orchestration tool adjusts the generation constraints, calling parameters, or execution nodes.

8. A multimodal data augmentation agent system for data-constrained scenarios based on the method of any one of claims 1-7, characterized in that, The system specifically includes: A specification parsing tool is used to receive data augmentation tasks, sample completion tasks, or data construction tasks for model training, and parse the tasks into structured task specifications that can be recognized and executed by the intelligent agent system. The structured task specifications include at least one or more of the following: data modality, application scenario, target object, target category, task type, scenario constraints, spatial constraints, temporal constraints, sample format, target sample size, quality threshold, quality control requirements, generation constraints, and output requirements. The structured task specifications can be represented in YAML, JSON, TOML, database forms, interface parameters, or other structured description forms. The task orchestration tool is used to generate execution queries based on the structured task specifications and historical feedback in the memory storage tool, and to schedule the execution layer tools to execute the corresponding data augmentation tasks. The execution query includes at least one or more of the following: task type, target data modality, target category, generation constraints, quality control requirements, output format, and calling parameters. The task orchestration tool is also used to select, combine, and schedule the functional nodes in the execution layer tools based on the current sample status, historical failure records, quality feedback, sample gaps, and data augmentation requirements to form an execution path adapted to the current data augmentation task. The execution layer tool is used to invoke one or more of the following nodes based on the execution query: sample generation node, data augmentation node, annotation generation node, and structured attribute and metadata generation node. Based on one or more of the target object, target category, task type, scene constraints, spatial constraints, temporal constraints, sample format, and quality requirements in the structured task specifications, it generates or enhances target visual data such as images, videos, 3D images, remote sensing images, or industrial images to obtain candidate visual samples. The execution layer tool is also used to generate text and structured information corresponding to the candidate visual samples; the text and structured information includes sample semantic description, structured labels, target location, spatial attributes, temporal attributes, annotation information, generation parameters, etc. Metadata used for quality control and one or more data fields used for sample retrieval, filtering, or training; the candidate visual samples and their corresponding text and structured information together constitute candidate multimodal samples; the execution layer tool is also used to call the quality control node to perform automatic quality control on the candidate multimodal samples, the automatic quality control including one or more of sample quality assessment, semantic consistency assessment, annotation consistency assessment, task relevance assessment, and rule verification; when the candidate multimodal sample meets the preset quality requirements, it is determined as a multimodal sample that has passed quality control; when the candidate multimodal sample does not meet the preset quality requirements, it is determined as a failed sample, and the corresponding quality score, failure reason, and execution status information are recorded; A memory storage tool is used to store and update the execution results returned by the execution layer tools. The execution results include one or more of the following: multimodal samples that have passed quality control, failed samples that have not passed quality control, quality scores, failure reasons, generation parameters, task specifications, sample categories, sample distribution, quality control pass rate, sample gaps, execution time, and resampling records. The memory storage tool is also used to organize the execution results according to one or more indexing methods of data modality, target category, scenario conditions, quality status, and execution rounds, and to provide historical execution feedback to the task orchestration tool for data gap identification, resampling scale determination, generation parameter adjustment, and execution node scheduling by statistically analyzing the number of generated samples, the number of passed samples, the distribution of failure reasons, and the quality control pass rate.

9. The multimodal data augmentation agent system for data-constrained scenarios as described in claim 8, characterized in that, The general interaction flow between the task orchestration tool and the execution layer tool is as follows: (1) The task orchestration tool reads the structured task specifications and historical feedback information in the memory storage tool to determine whether the current task belongs to the initial data construction task, the supplementary sampling task, the failed sample retry task, or the data completion task after receiving model training or validation feedback. (2) The task orchestration tool selects the corresponding execution node based on the task type, data modality, target object, target category, sample gap and quality requirements, and generates an execution query; (3) After the execution layer tool completes sample generation, enhancement and automatic quality control, it returns the generation results, quality score, pass status and execution status to the task orchestration tool; (4) The task scheduling tool then writes the relevant information into the memory storage tool and decides whether to continue to perform supplementary sampling based on the current sample quantity, quality control pass rate and quality distribution.

10. The multimodal data augmentation agent system for data-constrained scenarios as described in claim 8, characterized in that, The memory storage tool and the resampling process are as follows: The memory storage tool records all historical execution results and organizes them according to task category, scenario category, data modality, target object, quality score, and pass status. After reading the memory storage tool, the task orchestration tool counts the number of passed samples, the number of generated samples, the quality control pass rate, the distribution of failure reasons, and the target sample gap for each category or scenario condition. When the number of passed samples for a certain category or scenario condition is lower than the target sample size, the task orchestration tool generates a resampling task. When the quality control pass rate for a certain category or scenario condition is low, the task orchestration tool adjusts the generation parameters, calls the parameters, or replaces the execution node. When the failure reasons are concentrated in one or more of the following: target missing, semantic inconsistency, labeling errors, scenario mismatch, or metadata missing, the task orchestration tool adjusts the generation constraints, structured constraints, labeling generation nodes, or rule verification conditions according to the failure reasons. When the external model training or validation results show that a certain category or scenario condition performs poorly, the task orchestration tool uses this result as feedback information and increases its resampling priority.