Automated Training Method for Small Models Driven by Cognitive Intelligence of Multimodal Large Models

By adopting exclusive and shared sample annotation strategies in the model training system, the training data is automatically allocated and marked, and the problem of labeling task processing pressure during peak periods of model training demand is solved, achieving high-quality and continuous training data feedback.

CN119903442BActive Publication Date: 2025-05-30GONGYEYUN MFG (SICHUAN) INNOVATION CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510400017.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-05-30
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

During the peak period of model training demand, the platform side cannot reasonably allocate and process labeling tasks, resulting in inaccurate labeling data, insufficient feedback on training data, or overloaded hardware equipment.

Method used

By obtaining the training sample generation task in the target period and the task execution status of the model driver's task, we can determine whether the exclusive sample annotation strategy is met. If it is not met, we will adopt the shared sample annotation strategy, allocate and automatically label the training data to reduce the task processing volume of the model-driven group.

Benefits of technology

While ensuring the accuracy of user labeling data and the feedback amount of training data, avoid overloading of hardware equipment and balance the training quality and sustainability during peak periods of model training demand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119903442B_ABST
    Figure CN119903442B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of model automated training, and discloses a small model automated training method driven by multi-modal large model cognitive intelligence. By obtaining each training sample generation task received during the target period and the task execution status of each model driving end, it is determined whether the model driving group meets the time limit requirements for executing the exclusive sample annotation strategy, and the exclusive sample annotation strategy or the shared sample annotation strategy is selected to execute the allocation and automatic annotation actions for all multi-modal training data in each training sample generation task, and the training sample set obtained by automatic annotation is sent to the model application end for small model training for the corresponding detection scenario. Thus, by selectively sharing the annotation image ratio for the training sample generation task, a dual behavior determination system of the model application end and the model driving end is established, which reduces the task processing volume while minimizing the impact on the quality of the annotated data and the recognition accuracy of the training model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model automated training, and particularly to a small model automated training method driven by multi-modal large model cognitive intelligence. Background Art

[0002] Currently, there are a large number of application requirements for behavior recognition models in various model application terminals, such as the security monitoring field, the intelligent transportation field, the industrial manufacturing field, etc. These fields can deploy recognition models based on artificial intelligence at the model application terminal to improve the supervision and control of the detection scenario.

[0003] Due to the large number of model deployment requirements in different application scenarios, how to improve the training and deployment efficiency when building small models at the application terminal has become a major difficulty. In the traditional model training scheme, it is usually necessary to manually label the target categories and behavior categories of a large number of scene images, and then use the labeled images to train the model in the corresponding detection scenario. This process is time-consuming and laborious. Manual labeling may take several months, and manual labeling is subjective and affects the labeling quality.

[0004] Based on the defects in the above traditional model training scheme, some platforms have proposed an automatic labeling of the data set using a large model and using the labeled training set to implement the small model training scheme in the corresponding scenario. Users only need to provide basic multi-modal data and model application requirements to obtain labeled training data samples, and can quickly implement the training and deployment of small models in the local corresponding scenario. However, with the continuous growth of model training requirements, the task volume of data labeling executed by the platform side gradually approaches the load limit of the hardware device. Especially during the peak period of model training requirements, the platform side often needs to face short-term and overloaded task volumes. At this time, if the platform side cannot reasonably allocate and process the labeling tasks, problems such as inaccurate labeling data feedback to users, insufficient training data feedback, or overloaded operation of the hardware device for deploying the large model will occur.

[0005] Therefore, how to achieve the automated training of small models driven by multi-modal large model cognitive intelligence, while ensuring the accuracy of user-labeled data and the amount of training data feedback to a certain extent, and avoiding overloaded operation of the hardware device for deploying the large model, so as to balance the training quality and training sustainability during the peak period of model training requirements, is a technical problem that needs to be solved urgently. Summary of the Invention

[0006] The present invention provides a small model automated training method driven by multi-modal large model cognitive intelligence, aiming to solve at least one of the above technical problems.

[0007] To achieve the above object, the present invention provides a small model automatic training method based on multimodal large model cognitive intelligent driving, which is used in a model automatic training system. The model automatic training system includes a model driving group having several model driving terminals, a model application group having several model application terminals, and a labeling task allocation server. The method includes the following steps:

[0008] S1: The labeling task allocation server receives the training sample generation task sent by each model application end in the target period and the task execution status sent by each model driver end; wherein the training sample generation task includes the labeling task deadline association information and several detection scene features;

[0009] S2: The labeling task allocation server determines whether the model-driven group meets the deadline requirements for executing the exclusive sample labeling strategy based on the task execution status and the labeling task deadline association information;

[0010] S3: If yes, the labeling task allocation server distributes the training sample generation task to the corresponding model driver according to the exclusive sample labeling strategy, and uses the large model deployed by the model driver to perform the automatic labeling action of the multimodal training data;

[0011] S4: If not, the labeling task allocation server divides the training sample generation task into several general level categories based on several detection scene features, generates a shared sample labeling strategy for the training sample generation task for the associated general level based on the general level of each training sample generation task, and distributes the training sample generation task to the corresponding model driver, and uses the large model deployed by the model driver to perform the automatic labeling action of the multimodal training data;

[0012] S5: Each model driver sends the training sample set obtained after performing the automatic labeling action of the multimodal training data to the model application end corresponding to the training sample generation task, driving the model application end to use the training sample set to train and deploy the terminal category recognition model and the terminal behavior recognition model;

[0013] S6: Each model application terminal uses the deployed terminal category recognition model and terminal behavior recognition model to perform scene target category recognition and scene target behavior recognition of the target scene image in the corresponding detection scene, and determines whether the behavior of the scene target is a prohibited behavior;

[0014] S7: If yes, the target scene image is sent to the verification model driver of the model driver group, and the prohibited behavior judgment verification is performed using the large model deployed by the verification model driver.

[0015] Optionally, step S1 specifically includes:

[0016] S11: Each model application terminal collects multi-modal training data and scenario detection description information of the target detection scenario, extracts the annotation task deadline association information and several scenario detection features in the scenario detection description information, constructs the training sample generation task together with the multi-modal training data as the task label, and sends it to the annotation task allocation server;

[0017] S12: Each model driver terminal calls the training sample generation task list stored by itself, extracts the estimated execution period of each training sample generation task to be executed in the training sample generation task list, and sends it to the annotation task allocation server as the task execution status of each model driver terminal;

[0018] S13: The annotation task allocation server receives the training sample generation tasks sent by each model application terminal during the target period and the task execution status sent by each model driver terminal.

[0019] Optionally, in step S11, extracting the annotation task deadline association information and several scenario detection features in the scenario detection description information, constructing the training sample generation task together with the multi-modal training data as the task label specifically includes:

[0020] Parse the scenario detection description information, extract the sample generation deadline and several scenario detection features in the scenario detection description information, and generate the training data demand quantity of each training sample generation task based on the several scenario detection features;

[0021] Summarize the training data demand quantity and the sample generation deadline as the annotation task deadline association information, and construct the training sample generation task together with the annotation task deadline association information and several scenario detection features as the task label and the multi-modal training data, and send it to the annotation task allocation server.

[0022] Optionally, generating the training data demand quantity of each training sample generation task based on the several scenario detection features specifically includes:

[0023] Extract the number of recognized object types of the recognized object features and the recognized scene categories determined by the recognized scene features in the several scenario detection features;

[0024] Call the predefined relationship comparison table of the required number of training images corresponding to different numbers of recognized object types and recognized scene categories, and match the training data demand quantity of each training sample generation task.

[0025] Optionally, step S2 specifically includes:

[0026] S21: The annotation task allocation server extracts the estimated execution period of each training sample generation task to be executed for each model-driven end from the task execution status, and extracts the training data demand and sample generation deadline of each training sample generation task from the annotation task deadline association information;

[0027] S22: Perform an exhaustive allocation on all the multimodal training data in each training sample generation task received within the target period, and obtain several exclusive sample annotation schemes for the multimodal training data in the training sample generation task to be allocated to the corresponding model-driven end;

[0028] S23: According to the estimated execution period of each training sample generation task to be executed, the training data demand and sample generation deadline of each training sample generation task, and in accordance with the standard annotation speed of each model-driven end and the total training data demand of each training sample generation task, calculate the completion time after each model-driven end performs the annotation action on the allocated multimodal training data;

[0029] S24: Determine whether there is an exclusive sample annotation scheme among the several exclusive sample annotation schemes where the completion times of all the model-driven ends for executing the training sample generation tasks within the target period all meet the sample generation deadline.

[0030] Optionally, step S3 specifically includes:

[0031] S31: If so, the annotation task allocation server extracts the target exclusive sample annotation scheme with the farthest distance from the sample generation deadline among the exclusive sample annotation strategies;

[0032] S32: Based on the strategy of allocating all the multimodal training data of each training sample generation task to the model-driven end in the target exclusive sample annotation scheme, distribute the corresponding multimodal training data in each training sample generation task to the corresponding model-driven end, and use the large model deployed on the model-driven end to perform the automatic annotation action on the multimodal training data.

[0033] Optionally, step S4 specifically includes:

[0034] S41: If not, the annotation task allocation server calculates the detection scenario similarity between any two training sample generation tasks based on the several detection scenario features of each training sample generation task;

[0035] S42: Sort the detection scenario similarities between each training sample generation task and other training sample generation tasks from high to low, and determine whether the highest-ranked detection scenario similarity after sorting exceeds the preset similarity threshold. If so, define the training sample generation task as a general scenario task; if not, define the training sample generation task as a non-general scenario task;

[0036] S43: Divide all general-purpose scenario tasks into several general-purpose level categories according to the highest detected scenario similarity in the general-purpose scenario tasks. Based on the training data demand and sample generation deadline of each training sample generation task, considering the estimated execution time period of each training sample generation task to be executed at each model driver end and the general-purpose level of each general-purpose scenario task, perform traversal allocation on the multi-modal training data in the training sample generation tasks received within the target time period. According to the principle of the minimum shared generalization quantization value, generate a shared sample annotation strategy for the training sample generation task for the associated general-purpose level;

[0037] Among them, the principle of the minimum shared generalization quantization value is specifically:

[0038] First, perform traversal allocation on all the multi-modal training data of the non-general-purpose scenario tasks in the training sample generation task to obtain several exclusive sample annotation schemes in which the multi-modal training data in the non-general-purpose scenario tasks are cut and allocated to the model driver ends according to different ratios;

[0039] Then, perform traversal allocation on all the multi-modal training data of each general-purpose scenario task in the training sample generation task according to the corresponding shared sample annotation ratio in each exclusive sample annotation scheme, so that the completion times of all the model driver ends of the training sample generation tasks within the execution target time period all meet the sample generation deadline, and the shared generalization quantization value of all general-purpose scenario tasks is the smallest. The shared generalization quantization value is configured as: the cumulative sum of the products of the shared data volume determined by the shared sample annotation ratio and the training data demand of each general-purpose scenario task and the general-purpose level of the general-purpose scenario task;

[0040] S44: Distribute the training sample generation tasks to the corresponding model driver ends according to the shared sample annotation strategy, and use the large model deployed at the model driver end to perform automatic annotation actions on the multi-modal training data. Extract the corresponding number of training image data according to the shared sample annotation ratio of the general-purpose scenario task from the training image data set of the corresponding general-purpose scenario task with the highest detected scenario similarity of each general-purpose scenario task and put it into the training image data set of the general-purpose scenario task.

[0041] Optionally, using the large model deployed at the model driver end to perform automatic annotation actions on the multi-modal training data includes:

[0042] Use the large model deployed at the model driver end to perform data modality conversion and data augmentation processing on the multi-modal training data to obtain a training image data set for performing automatic annotation;

[0043] The large model deployed on the model driver is used to perform target category detection and behavior detection for each target category on the training image dataset. The detection results are annotated in the training images and converted into a text file in YOLO format. A training sample set is constructed and divided into a training set and a validation set.

[0044] Optionally, in step S5, the model application end is driven to use the training sample set to train and deploy the terminal category recognition model and the terminal behavior recognition model, specifically: obtain the official pre-trained model yolo11n.pt and import it into the model application end, drive the model application end to use the training sample set of target category detection and the training sample set of behavior detection of each target category to train the pre-trained model yolo11n.pt respectively, and deploy the generated terminal category recognition model and the terminal behavior recognition model of each target category on the model application end.

[0045] Optionally, step S6 specifically includes:

[0046] S61: Each model application terminal uses the deployed terminal category recognition model to perform scene target category recognition of the target scene image in the corresponding detection scene;

[0047] S62: Perform target segmentation on the target scene image according to the scene target category recognition result, and use the terminal behavior recognition model of each target category deployed on each model application end to perform scene target behavior recognition on the category in each image obtained by segmentation, and determine whether the behavior of the scene target is a prohibited behavior.

[0048] The beneficial effects of the present invention are as follows: A small model automated training method driven by the cognitive intelligence of a multi-modal large model is proposed. By obtaining each training sample generation task received during the target period and the task execution status of each model driving end, it is determined whether the model driving group meets the deadline requirement for executing the exclusive sample annotation strategy. If so, the allocation and automatic annotation actions of all multi-modal training data in each training sample generation task are performed according to the exclusive sample annotation strategy. If not, the allocation and automatic annotation actions of all multi-modal training data in each training sample generation task are performed according to the shared sample annotation strategy. The training sample set obtained by automatic annotation is sent to the model application end for small model training for the corresponding detection scenario. Thus, when it is detected that the task processing ability of the model driving group does not meet the current task execution requirements, by selectively sharing the annotation image ratio of the training sample generation tasks, the task processing volume of the model driving group is reduced, and by establishing a dual behavior determination system for the model application end and the model driving end, the impact of shared sample annotation on the quality of annotation data and the recognition accuracy of the training model is minimized as much as possible. While ensuring the accuracy of user annotation data and the amount of training data feedback to a certain extent, it avoids overloading the hardware devices for deploying large models, thereby balancing the training quality and training sustainability during the peak period of model training requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic flowchart of an embodiment of the small model automated training method driven by the cognitive intelligence of a multi-modal large model according to the present invention.

[0050] The realization, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] In order to make the object, technical solution, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0052] An embodiment of the present invention provides a small model automated training method driven by the cognitive intelligence of a multi-modal large model. Refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the small model automated training method driven by the cognitive intelligence of a multi-modal large model according to the present invention.

[0053] In this embodiment, a small model automatic training method based on multimodal large model cognitive intelligent driving is used in a model automatic training system, wherein the model automatic training system includes a model driving group having a plurality of model driving terminals, a model application group having a plurality of model application terminals, and a labeling task allocation server, and the method includes the following steps:

[0054] S1: The labeling task allocation server receives the training sample generation task sent by each model application end in the target period and the task execution status sent by each model driver end; wherein the training sample generation task includes the labeling task deadline association information and several detection scene features;

[0055] S2: The labeling task allocation server determines whether the model-driven group meets the deadline requirements for executing the exclusive sample labeling strategy based on the task execution status and the labeling task deadline association information;

[0056] S3: If yes, the labeling task allocation server distributes the training sample generation task to the corresponding model driver according to the exclusive sample labeling strategy, and uses the large model deployed by the model driver to perform the automatic labeling action of the multimodal training data;

[0057] S4: If not, the labeling task allocation server divides the training sample generation task into several general level categories based on several detection scene features, generates a shared sample labeling strategy for the training sample generation task for the associated general level based on the general level of each training sample generation task, and distributes the training sample generation task to the corresponding model driver, and uses the large model deployed by the model driver to perform the automatic labeling action of the multimodal training data;

[0058] S5: Each model driver sends the training sample set obtained after performing the automatic labeling action of the multimodal training data to the model application end corresponding to the training sample generation task, driving the model application end to use the training sample set to train and deploy the terminal category recognition model and the terminal behavior recognition model;

[0059] S6: Each model application terminal uses the deployed terminal category recognition model and terminal behavior recognition model to perform scene target category recognition and scene target behavior recognition of the target scene image in the corresponding detection scene, and determines whether the behavior of the scene target is a prohibited behavior;

[0060] S7: If yes, the target scene image is sent to the verification model driver of the model driver group, and the prohibited behavior judgment verification is performed using the large model deployed by the verification model driver.

[0061] It should be noted that with the continuous growth of model training requirements, the task volume of data annotation executed on the platform side is gradually approaching the load upper limit of the hardware device. Especially during the peak period of model training requirements, the platform side often has to face the situation of short-term and overloaded task volume. At this time, if the platform side cannot reasonably allocate and process the annotation tasks, it will lead to problems such as inaccurate annotation data fed back to users, insufficient training data feedback, or overloaded operation of the hardware device for deploying large models.

[0062] To solve the above problems, in this embodiment, by obtaining each training sample generation task received during the target period and the task execution status of each model driver end, it is determined whether the model driver group meets the deadline requirements for executing the exclusive sample annotation strategy, and the exclusive sample annotation strategy or the shared sample annotation strategy is selected to perform the allocation and automatic annotation actions for all the multimodal training data in each training sample generation task, and the training sample set obtained by automatic annotation is sent to the model application end for small model training for the corresponding detection scenario. Thus, by selectively sharing the annotation image ratio for the training sample generation tasks and establishing a dual behavior determination system for the model application end and the model driver end, while reducing the task processing volume, the impact on the quality of the annotation data and the recognition accuracy of the training model is minimized as much as possible.

[0063] In a preferred embodiment, step S1 specifically includes:

[0064] S11: Each model application end collects the multimodal training data and the scene detection description information of the target detection scenario, extracts the annotation task deadline association information and several scene detection features in the scene detection description information, constructs them as a training sample generation task together with the multimodal training data, and sends it to the annotation task allocation server;

[0065] S12: Each model driver end calls the training sample generation task list stored by itself, extracts the estimated execution period of each training sample generation task to be executed in the training sample generation task list, uses it as the task execution status of each model driver end, and sends it to the annotation task allocation server;

[0066] S13: The annotation task allocation server receives the training sample generation tasks sent by each model application end during the target period and the task execution status sent by each model driver end.

[0067] Further, in step S11, the annotation task deadline correlation information and several scene detection features in the scene detection description information are extracted as task tags and jointly constructed with the multi-modal training data into a training sample generation task, specifically including: parsing the scene detection description information, extracting the sample generation deadline and several scene detection features in the scene detection description information, and generating the training data demand quantity for each training sample generation task based on the several scene detection features; summarizing the training data demand quantity and the sample generation deadline into the annotation task deadline correlation information, and using the annotation task deadline correlation information and the several scene detection features as task tags and jointly constructing with the multi-modal training data into a training sample generation task, and sending it to the annotation task allocation server.

[0068] Further, based on the several scene detection features, generating the training data demand quantity for each training sample generation task, specifically including: extracting the number of recognized object types of the recognized object features and the recognized scene categories determined by the recognized scene features in the several scene detection features; calling the pre-defined relationship comparison table of the corresponding required number of training images under different numbers of recognized object types and recognized scene categories, and matching the training data demand quantity for each training sample generation task.

[0069] In this embodiment, by driving each model application end and each model driving end to separately send the training sample generation task and the estimated execution period of each to-be-executed training sample generation task to the annotation task allocation server, and using the task execution status determined by the scene detection description information and the estimated execution period recorded in the training sample generation task, the model training requirements and task execution capabilities of the model application end are summarized to the annotation task allocation server, providing data support for the subsequent annotation task allocation server to judge the task completion deadline and select the sample annotation strategy.

[0070] In a preferred embodiment, step S2 specifically includes:

[0071] S21: The annotation task allocation server extracts the estimated execution period of each to-be-executed training sample generation task of each model driving end from the task execution status, and extracts the training data demand quantity and the sample generation deadline of each training sample generation task from the annotation task deadline correlation information;

[0072] S22: Perform traversal allocation on all the multi-modal training data in each training sample generation task received within the target period, and obtain several exclusive sample annotation schemes for the multi-modal training data in the training sample generation task to be allocated to the corresponding model driving end;

[0073] S23: Generate the estimated execution time period for the task to be performed for each training sample generation task, the training data demand for each training sample generation task, and the sample generation deadline. Calculate the completion time for each model driver after performing the annotation action on the multimodal training data assigned to it, according to the standard annotation speed of each model driver and the total training data demand for each training sample generation task.

[0074] S24: Determine whether there is an exclusive sample annotation scheme among several exclusive sample annotation schemes where the completion times of all model drivers for the training sample generation tasks within the execution target time period all meet the sample generation deadline.

[0075] In this embodiment, first, perform traversal allocation on all the multimodal training data in each training sample generation task received within the target time period to obtain several exclusive sample annotation schemes (that is, all the multimodal training data in each training sample generation task need to perform annotation actions after being assigned to the model driver). By calculating the completion time of each model driver after performing the annotation action on the assigned multimodal training data and comparing it with the sample generation deadline corresponding to the training sample generation task, it is determined whether the exclusive sample annotation scheme can be used to perform the training sample generation task allocation within the target time period.

[0076] In a preferred embodiment, step S3 specifically includes:

[0077] S31: If so, the annotation task allocation server extracts the target exclusive sample annotation scheme with the farthest completion time from the sample generation deadline in the exclusive sample annotation strategy.

[0078] S32: Based on the strategy of allocating all the multimodal training data of each training sample generation task to the model driver in the target exclusive sample annotation scheme, distribute the corresponding multimodal training data of each training sample generation task to the corresponding model driver, and use the large model deployed on the model driver to perform the automatic annotation action on the multimodal training data.

[0079] In this embodiment, when the completion time after each model driver performs the annotation action on the assigned multimodal training data can meet the sample generation deadline corresponding to the training sample generation task, the exclusive sample annotation scheme can be used to control all the multimodal training data in each training sample generation task to perform annotation actions after being assigned to the model driver, and the multimodal training data after performing the annotation action is used as the training image dataset for the model application side to perform small model training for the detection scenario.

[0080] In a preferred embodiment, step S4 specifically includes:

[0081] S41: If not, the annotation task allocation server generates several detection scenario features of the task based on each training sample, and calculates the detection scenario similarity between any two training sample-generated tasks;

[0082] S42: Sort the detection scenario similarities between each training sample-generated task and other training sample-generated tasks from high to low, and determine whether the highest-ranked detection scenario similarity after sorting exceeds a preset similarity threshold. If so, define the training sample-generated task as a general-purpose scenario task; if not, define the training sample-generated task as a non-general-purpose scenario task;

[0083] S43: According to the highest-ranked detection scenario similarity in the general-purpose scenario tasks, divide all general-purpose scenario tasks into several general-purpose level categories. Based on the training data demand and sample generation deadline of each training sample-generated task, considering the estimated execution time period of each to-be-executed training sample-generated task of each model driver and the general-purpose level of each general-purpose scenario task, perform traversal allocation on the multi-modal training data in the training sample-generated tasks received within the target time period, and generate a shared sample annotation strategy for the training sample-generated task for the associated general-purpose level according to the principle of the minimum shared general-purpose quantification value;

[0084] Among them, the principle of the minimum shared general-purpose quantification value is specifically:

[0085] First, perform traversal allocation on all the multi-modal training data of the non-general-purpose scenario tasks in the training sample-generated tasks to obtain several exclusive sample annotation schemes in which the multi-modal training data in the non-general-purpose scenario tasks are cut and allocated to the model drivers according to different ratios;

[0086] Then, perform traversal allocation on all the multi-modal training data of each general-purpose scenario task in the training sample-generated tasks according to the corresponding shared sample annotation ratio in each exclusive sample annotation scheme, so that the completion times of all the model drivers for the training sample-generated tasks within the execution target time period all meet the sample generation deadline, and the shared general-purpose quantification value of all general-purpose scenario tasks is the smallest. The shared general-purpose quantification value is configured as: the cumulative sum of the product of the training data sharing amount determined by the shared sample annotation ratio and the training data demand of each general-purpose scenario task and the general-purpose level of this general-purpose scenario task;

[0087] S44: According to the shared sample annotation strategy, distribute the training sample-generated tasks to the corresponding model drivers, use the large model deployed on the model driver to perform automatic annotation actions on the multi-modal training data, and extract the corresponding number of training image data of the shared sample annotation ratio of this general-purpose scenario task from the training image dataset of the general-purpose scenario task with the highest detection scenario similarity of each general-purpose scenario task and put it into the training image dataset of this general-purpose scenario task.

[0088] In this embodiment, by obtaining the task of each training sample generation received during the target period and the task execution status of each model-driven end, it is determined whether the model-driven group meets the deadline requirement for executing the exclusive sample annotation strategy. If not, according to the shared sample annotation strategy, the allocation and automatic annotation actions of all multi-modal training data in each training sample generation task are performed (that is, all multi-modal training data in each training sample generation task is distributed to the model-driven end for annotation actions after reducing the shared sample annotation ratio. The missing multi-modal training data of this shared sample annotation ratio is extracted from the training image dataset of the corresponding general-purpose scenario task with the highest detected scene similarity for supplementation. It should be noted that this shared sample annotation ratio usually has a limit threshold in practical applications. On the one hand, to ensure the quality of the output training image dataset, and on the other hand, to ensure that the missing multi-modal training data is not too much to avoid the situation where the data volume in the supplementary training image dataset is insufficient). The training sample set obtained by automatic annotation is sent to the model application end for small model training for the corresponding detection scene. Thus, when it is detected that the task processing ability of the model-driven group does not meet the current task execution requirements, by selectively sharing the annotation image ratio of the training sample generation task, the task processing volume of the model-driven group is reduced. While ensuring the accuracy of user-annotated data and the amount of training data feedback to a certain extent, it avoids the overloading operation of the hardware device for deploying the large model.

[0089] In a preferred embodiment, the automatic annotation action of multi-modal training data using the large model deployed on the model-driven end includes: using the large model deployed on the model-driven end to perform data modality conversion and data augmentation processing on the multi-modal training data to obtain a training image dataset for performing automatic annotation; using the large model deployed on the model-driven end to perform target category detection and behavior detection for each target category on the training image dataset, annotating the detection results in the training image and converting them into a text file in YOLO format, constructing a training sample set and dividing it into a training set and a validation set.

[0090] On this basis, in step S5, the model application end is driven to use the training sample set for training and deployment of the terminal category recognition model and the terminal behavior recognition model. Specifically, the official pre-trained model yolo11n.pt is obtained and imported into the model application end. The model application end is driven to train the pre-trained model yolo11n.pt using the training sample set of target category detection and the training sample set of behavior detection for each target category, and deploy the generated terminal category recognition model and the terminal behavior recognition model for each target category at the model application end.

[0091] In a preferred embodiment, step S6 specifically includes:

[0092] S61: Each model application terminal uses the deployed terminal category recognition model to perform scene target category recognition of the target scene image in the corresponding detection scene;

[0093] S62: Perform target segmentation on the target scene image according to the scene target category recognition result, and use the terminal behavior recognition model of each target category deployed on each model application end to perform scene target behavior recognition on the category in each image obtained by segmentation, and determine whether the behavior of the scene target is a prohibited behavior.

[0094] In this embodiment, by establishing a dual behavior judgment system of the model application end and the model driving end, the impact of shared sample annotation on the quality of annotation data and the recognition accuracy of the training model is minimized, thereby balancing the training quality and training continuity during the peak period of model training demand.

[0095] In addition, the present invention also proposes a small model automatic training device based on multimodal large model cognitive intelligence driving, and the small model automatic training device based on multimodal large model cognitive intelligence driving includes: a memory, a processor, and a small model automatic training program based on multimodal large model cognitive intelligence driving stored in the memory and executable on the processor. When the small model automatic training program based on multimodal large model cognitive intelligence driving is executed by the processor, the steps of the small model automatic training method based on multimodal large model cognitive intelligence driving as described above are implemented.

[0096] The specific implementation of the small model automatic training device based on multimodal large model cognitive intelligence driving in the present application is basically the same as the above-mentioned embodiments of the small model automatic training method based on multimodal large model cognitive intelligence driving, and will not be repeated here.

[0097] It is understood that, in the description of this specification, the description with reference to the terms "one embodiment", "another embodiment", "other embodiments", or "first to Nth embodiments" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0098] It should be noted that in this text, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or system that includes a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or system that includes such an element.

[0099] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A small model automated training method based on multimodal large model cognitive intelligence drive, characterized in that: For a model automation training system, the model automation training system includes a model driving group having a plurality of model driving terminals, a model application group having a plurality of model application terminals, and a labeling task allocation server, the method includes the following steps: S1: The labeling task allocation server receives the training sample generation task sent by each model application end in the target period and the task execution status sent by each model driver end; wherein the training sample generation task includes the labeling task deadline association information and several detection scene features; S2: The labeling task allocation server determines whether the model-driven group meets the deadline requirements for executing the exclusive sample labeling strategy based on the task execution status and the labeling task deadline association information; S3: If yes, the labeling task allocation server distributes the training sample generation task to the corresponding model driver according to the exclusive sample labeling strategy, and uses the large model deployed by the model driver to perform the automatic labeling action of the multimodal training data; S4: If not, the labeling task allocation server divides the training sample generation task into several general level categories based on several detection scene features, generates a shared sample labeling strategy for the training sample generation task for the associated general level based on the general level of each training sample generation task, and distributes the training sample generation task to the corresponding model driver, and uses the large model deployed by the model driver to perform the automatic labeling action of the multimodal training data; S5: Each model driver sends the training sample set obtained after performing the automatic labeling action of the multimodal training data to the model application end corresponding to the training sample generation task, driving the model application end to use the training sample set to train and deploy the terminal category recognition model and the terminal behavior recognition model; S6: Each model application terminal uses the deployed terminal category recognition model and terminal behavior recognition model to perform scene target category recognition and scene target behavior recognition of the target scene image in the corresponding detection scene, and determines whether the behavior of the scene target is a prohibited behavior; S7: If yes, the target scene image is sent to the verification model driver of the model driver group, and the prohibited behavior judgment verification is performed using the large model deployed by the verification model driver.

2. The small model automatic training method based on multimodal large model cognitive intelligence drive as claimed in claim 1 is characterized in that: Step S1 specifically includes: S11: Each model application end collects multimodal training data and scene detection description information of the target detection scene, extracts the labeling task deadline association information and several scene detection features in the scene detection description information, and constructs them as task labels together with the multimodal training data to generate training sample tasks, and sends them to the labeling task allocation server; S12: Each model driver calls the training sample generation task list stored in itself, extracts the estimated execution period of each training sample generation task to be executed in the training sample generation task list, and sends it to the labeling task allocation server as the task execution status of each model driver; S13: The labeling task allocation server receives the training sample generation task sent by each model application end in the target period and the task execution status sent by each model driver end.

3. The small model automatic training method based on multimodal large model cognitive intelligence drive as claimed in claim 2 is characterized in that: In step S11, the labeling task deadline association information and several scene detection features in the scene detection description information are extracted as task labels and constructed together with the multimodal training data to form a training sample generation task, which specifically includes: Parsing the scene detection description information, extracting the sample generation period and several scene detection features in the scene detection description information, and generating the training data requirement for each training sample generation task based on the several scene detection features; The training data demand and sample generation deadline are summarized as labeling task deadline association information, and the labeling task deadline association information and several scene detection features are used as task labels together with multimodal training data to form a training sample generation task, which is sent to the labeling task allocation server.

4. The small model automated training method based on multimodal large model cognitive intelligence drive as claimed in claim 3, characterized in that: Based on several scene detection features, the training data requirements for each training sample generation task are generated, including: Extracting the number of recognized object types of recognized object features and the recognized scene categories determined by recognized scene features from a plurality of scene detection features; Call the predefined comparison table of the relationship between the number of different types of recognized objects and the corresponding number of training images required under the recognition scene category to match the training data requirement of each training sample generation task.

5. The small model automatic training method based on multimodal large model cognitive intelligence drive as claimed in claim 1, characterized in that: Step S2 specifically includes: S21: The labeling task allocation server extracts the estimated execution period of each to-be-executed training sample generation task of each model driver from the task execution status, and extracts the training data demand and sample generation deadline of each training sample generation task from the labeling task deadline association information; S22: performing ergodic allocation on all multimodal training data in each training sample generation task received within the target period, and obtaining a plurality of exclusive sample labeling schemes for allocating the multimodal training data in the training sample generation task to the corresponding model driver end; S23: according to the estimated execution period of each training sample generation task to be executed, the training data requirement and sample generation period of each training sample generation task, according to the standard annotation speed of each model driver and the total training data requirement of each training sample generation task, calculate the completion time after each model driver executes the annotation action of the allocated multimodal training data; S24: Determine whether there is an exclusive sample labeling scheme among several exclusive sample labeling schemes, in which the completion time of all model drivers executing the training sample generation task within the target period meets the sample generation deadline.

6. The small model automatic training method based on multimodal large model cognitive intelligence drive as claimed in claim 5, characterized in that: Step S3 specifically includes: S31: If yes, the labeling task allocation server extracts the target exclusive sample labeling scheme whose last completion time is the farthest from the sample generation deadline in the exclusive sample labeling strategy; S32: Based on the strategy of allocating all multimodal training data for each training sample generation task to the model driving end in the target exclusive sample labeling scheme, the corresponding multimodal training data in each training sample generation task is distributed to the corresponding model driving end, and the large model deployed by the model driving end is used to perform automatic labeling of the multimodal training data.

7. The small model automatic training method based on multimodal large model cognitive intelligence drive as claimed in claim 1, characterized in that: Step S4 specifically includes: S41: If not, the labeling task allocation server calculates the detection scene similarity between any two training sample generation tasks based on a number of detection scene features of each training sample generation task; S42: sorting the detection scene similarities between each training sample generation task and other training sample generation tasks from high to low, and judging whether the detection scene similarity with the highest ranking after sorting exceeds a preset similarity threshold; if so, defining the training sample generation task as a general-purpose scene task; if not, defining the training sample generation task as a non-general-purpose scene task; S43: according to the highest ranking detection scene similarity in the general-purpose scene task, all general-purpose scene tasks are divided into several general-purpose level categories, based on the training data demand and sample generation period of each training sample generation task, considering the estimated execution period of each to-be-executed training sample generation task of each model driver and the general-purpose level of each general-purpose scene task, ergodic allocation is performed on the multimodal training data in the training sample generation task received within the target period, and according to the principle of minimum shared general-purpose quantization value, a shared sample labeling strategy for the training sample generation task for the associated general-purpose level is generated; The minimum shared universal quantization value principle is specifically: First, all multimodal training data of non-general-purpose scenario tasks in the training sample generation task are ergodicly distributed to obtain several exclusive sample labeling schemes in which the multimodal training data in the non-general-purpose scenario tasks are cut in different proportions and distributed to the model driver end; Then, all multimodal training data of each general-purpose scenario task in the training sample generation task are ergodicly allocated in each exclusive sample annotation scheme according to the corresponding shared sample annotation ratio, so that the completion time of all model driving ends executing the training sample generation task within the target period meets the sample generation deadline, and the shared general-purpose quantization value of all general-purpose scenario tasks is minimized, and the shared general-purpose quantization value is configured as: the product and accumulation of the shared sample annotation ratio of each general-purpose scenario task and the training data sharing amount determined by the training data demand amount and the general level of the general-purpose scenario task; S44: According to the shared sample labeling strategy, the training sample generation task is distributed to the corresponding model driving end, and the large model deployed by the model driving end is used to perform automatic labeling of the multimodal training data. From the training image data set of the corresponding general-purpose scenario task with the highest detection scene similarity for each general-purpose scenario task, the training image data of a corresponding number of shared sample labeling ratios of the general-purpose scenario task are extracted and put into the training image data set of the general-purpose scenario task.

8. The small model automatic training method based on multimodal large model cognitive intelligence drive as described in claim 6 or 7, characterized in that: The automatic labeling actions of multimodal training data using the large model deployed on the model driver include: Use the large model deployed on the model driver to perform data mode conversion and data enhancement processing on the multimodal training data to obtain a training image dataset for automatic annotation. The large model deployed on the model driver is used to perform target category detection and behavior detection for each target category on the training image dataset. The detection results are annotated in the training images and converted into a text file in YOLO format. A training sample set is constructed and divided into a training set and a validation set.

9. The small model automated training method based on multimodal large model cognitive intelligence drive as claimed in claim 8, characterized in that: In step S5, the model application end is driven to use the training sample set to train and deploy the terminal category recognition model and the terminal behavior recognition model, specifically: obtain the official pre-trained model yolo11n.pt and import it into the model application end, drive the model application end to use the training sample set of target category detection and the training sample set of behavior detection of each target category to train the pre-trained model yolo11n.pt respectively, and deploy the generated terminal category recognition model and the terminal behavior recognition model of each target category on the model application end.

10. The small model automatic training method based on multimodal large model cognitive intelligence drive according to claim 9, characterized in that: Step S6 specifically includes: S61: Each model application terminal uses the deployed terminal category recognition model to perform scene target category recognition of the target scene image in the corresponding detection scene; S62: Perform target segmentation on the target scene image according to the scene target category recognition result, and use the terminal behavior recognition model of each target category deployed on each model application end to perform scene target behavior recognition on the category in each image obtained by segmentation, and determine whether the behavior of the scene target is a prohibited behavior.

Citation Information

Patent Citations

  • Machine learning model obtaining method and obtaining device, apparatus, and storage medium

    CN109034188A

  • Targeted data acquisition for model training

    US20230016082A1