Data generation method and device based on multi-modal large language model

By training a multimodal large language model to generate a pyramid-shaped thinking chain of entities, attributes, rules, and inductive information, the systemic problem of data generation in real business scenarios using multimodal large language models is solved, achieving more efficient data generation and improved model performance.

CN121168641APending Publication Date: 2025-12-19HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511261336.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing multimodal large language models lack systematic data generation capabilities in real business scenarios, and cannot effectively handle complex and open business scenarios, resulting in illusory output and poor basic performance.

Method used

By training a multimodal large language model, a pyramid-shaped thinking chain of entity information, attribute information, rule information, and inductive information is generated. Combined with perception and standardized text description, a standardized multimodal data generation method is formed. The trained model is then used for multi-dimensional monitoring and parameter adjustment to form a standardized thinking chain.

Benefits of technology

It improves the adaptability and accuracy of multimodal large language models in complex business scenarios, avoids hallucination output, and enhances model performance and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168641A_ABST
    Figure CN121168641A_ABST
Patent Text Reader

Abstract

The invention discloses a data generation method based on a multi-modal large language model, and the method comprises the steps: obtaining inquiry text information used for representing a business task in a business application scene, and image information associated with the business task, the inquiry text information comprises at least one of task definition information used for representing a service task needing to be completed and task logic information used for representing service logic, reasoning the obtained inquiry text information and the image information based on the trained multi-mode large language model, and obtaining the inquiry text information and the image information. Generating prediction text information used for representing response information of the inquiry text information based on the image information, the prediction text information comprises at least one of entity information used for representing a target image in the image information and image position information where an entity is located, attribute information used for representing additional information of the entity, and rule information used for representing a rule disassembled based on task logic. According to the invention, the generated data has pyramid dimension information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, in particular, to a data generation method based on a multi-modal large language model. BACKGROUND

[0002] With the development of multi-modal large language models, there are many works at present that apply them to various fields, such as image-based target behavior analysis.

[0003] However, there is still a big gap between the real scene business and the landing application. The complexity of rules and the openness of the scene in the real business scenario lead to a lack of systematic data generation for multi-modal applications. SUMMARY

[0004] The present application provides a data generation method based on a multi-modal large language model to improve the systematization of data generation.

[0005] The first aspect of the present application provides a data generation method based on a multi-modal large language model, which comprises:

[0006] Obtaining inquiry text information for representing a business task in a business application scenario and image information associated with the business task, wherein the inquiry text information includes at least one of task definition information for representing a required completed business task and task logic information for representing business logic,

[0007] Based on the trained multi-modal large language model with perception and normalized text description, the obtained inquiry text information and image information are inferred to generate predicted text information for representing the response information of the inquiry text information based on the image information, wherein the predicted text information includes at least one of entity information for representing a target image in the image information and the image position information of the entity, attribute information for representing additional information of the entity, and rule information for representing rules decomposed based on the task logic.

[0008] As a possible implementation, the inquiry text information further includes output mode information for representing the expected output mode,

[0009] The predicted text information further includes at least one of explanation information for representing the inference process of the multi-modal large language model and / or the rule boundary of the representation rule, and induction information for representing the induction of the business task based on each task rule,

[0010] Each information in the predicted text information is sequentially adjacent according to the entity information, the attribute information, the rule information, the explanation information, and the induction information, and the information dimension of each information is sequentially increased, forming a standardized pyramid thinking chain paradigm.

[0011] As a possible implementation, the multi-modal large language model is trained in the following manner:

[0012] Obtain the query text sample, the image sample, and the expected text sample of the business task sample, and add a first identifier of the text information position of the image position information identifying the entity and a second identifier of the text information position of the attribute information in the expected text sample, wherein the expected text sample is used to represent the response text of the query text sample based on the image sample,

[0013] Input the obtained query text sample, image sample, and expected text sample into the multi-modal large language model, and generate a predicted text sample through inference of the multi-modal large language model,

[0014] In the supervised fine-tuning training phase of the multi-modal large language model, each text information in the predicted text sample is monitored according to the entity dimension, the attribute dimension, the rule dimension, the explanation dimension, and the induction dimension,

[0015] Adjust the model parameters of the multi-modal large language model according to the monitoring result,

[0016] Repeat until the monitoring result reaches the expected situation and stop training.

[0017] As a possible implementation, the query text sample is obtained in the following manner:

[0018] Based on any business task sample, the task definition, task logic, and output mode of the business task sample are described in natural language,

[0019] As a possible implementation, the expected text sample is obtained in the following manner:

[0020] According to the business logic of the business task sample, the business task sample is disassembled into at least one rule,

[0021] For each rule, the entity associated with the rule, the attribute associated with the entity and / or the rule, and the first explanation associated with the rule are sequentially described in natural language based on the image sample of the business task sample, obtaining the entity information, attribute information, and rule information and first explanation information of the rule,

[0022] Based on the disassembled rules, the task is induced, and the performed task induction and the second explanation associated with the induction are described in natural language, wherein the first explanation and the second explanation are used to guide the inference process of the multi-modal model,

[0023] According to the output mode in the query text sample, the expected conclusion is described.

[0024] As a possible implementation, the natural language description of the task sample based on the image sample of the business task, the entity associated with the rule, the attribute associated with the entity and / or the rule, and the first explanation associated with the rule in turn includes:

[0025] Based on the image sample of the business task sample, the image positioning of each entity in the image sample is performed to obtain the image position information of the entity, and the image position information is the image coordinate information of the image frame where the entity is located,

[0026] Obtain each entity associated with the rule and the sub-entity of the local entity used to represent the entity associated with the rule, and obtain the entity set of the rule,

[0027] Obtain the entity description of each element in the entity set and the attribute description associated with the rule, and sequentially describe the elements in the entity set with the first explanation description, and obtain the rule description information of the rule;

[0028] The task is summarized based on the decomposed rule, and the task summary and the second explanation associated with the summary are described in natural language, including:

[0029] Based on the rule description information of each rule, the task is described, and the summary description information is obtained,

[0030] The rule description information of each rule is sequentially adjacent, and then the summary description information is sequentially adjacent, to obtain the thinking chain description information of the business task sample.

[0031] As a possible implementation, the obtaining of each entity associated with the rule and the sub-entity of the local entity used to represent the entity associated with the rule includes:

[0032] According to the rule, the entity associated with the rule is filtered from the image sample,

[0033] The business logic analysis is sequentially performed on each independent entity associated with the rule, and if the entity has a sub-entity of the local entity used to represent the entity, the sub-entity and the entity are bound,

[0034] The filtered entity and the bound sub-entity are taken as elements in the entity set of the rule;

[0035] The obtaining of the entity description of each entity in the entity set and the attribute description associated with the rule includes:

[0036] The entity and its image position information are described in the form of image frame interlacing, the sub-entity and its image position information are described in the form of image frame interlacing, and the entity description information of each element in the entity set is obtained,

[0037] Based on the rule, the attribute description information is combined with the entity description information. If the attribute boundary is simple, the attribute description information is placed before the text information of the image position information. If the attribute boundary is complex and fuzzy, the attribute description information is spliced to the tail of the entity description information to obtain attribute description information.

[0038] The entity description information and the attribute description information of each entity in the entity set, the rule and the description information of the first explanation are sequentially connected.

[0039] As a possible implementation, the prediction text sample is monitored in the entity dimension, the attribute dimension, the rule dimension, the explanation dimension, and the induction dimension, respectively, including:

[0040] The follow-up monitoring is performed on each dimension,

[0041] In the case where the loss function value of the segmentation follow-up loss monitored by the follow-up monitoring is less than the set loss threshold, it is determined that the multi-modal model has learned the paradigm description, and the validation set monitoring is performed on each dimension. The accuracy of each dimension is evaluated using an evaluation function. In the case where the evaluation function value converges, the training is stopped.

[0042] As a possible implementation, the accuracy of each dimension includes:

[0043] The target box regression accuracy is used to represent the accuracy of the entity dimension,

[0044] The attribute category accuracy is used to represent the accuracy of the attribute dimension,

[0045] The conclusion accuracy is used to represent the accuracy of the rule dimension, the explanation dimension, and the induction dimension.

[0046] The evaluation function is used to weight average the accuracy of each dimension.

[0047] As a possible implementation, the method further includes:

[0048] The historical business tasks of each historical business application scenario, and the historical thinking chain and the historical conclusion thereof are accumulated as a thinking chain set of the multi-modal large model. Any historical business application scenario corresponds to multiple historical business tasks, each historical business task corresponds to a historical thinking chain, and each historical thinking chain corresponds to a historical conclusion. The historical thinking chain is used to provide a thinking mode for the multi-modal large model to reason, and the historical conclusion is used to provide a non-thinking mode for the multi-modal large model to reason.

[0049] In the case where the number of historical business application scenarios accumulates to reach a set accumulation threshold, the data in the thinking chain set is used as training sample data to train the current multi-modal large model.

[0050] In the case where the performance evaluation of the current multi-modal large model reaches the expectation, the current multi-modal large model is used to generate the thought chain and conclusion of each business task in the next business application scenario.

[0051] The second aspect of the present application provides a data generation device based on a multi-modal large language model, which comprises:

[0052] An input module is configured to obtain inquiry text information for representing a business task in a business application scenario and image information associated with the business task, wherein the inquiry text information comprises at least one of task definition information for representing a required completed business task and task logic information for representing business logic,

[0053] A generation module is configured to perform reasoning on the obtained inquiry text information and image information based on the trained multi-modal large language model with perception and normalized text description, and generate predicted text information for representing response information of the inquiry text information based on the image information, wherein the predicted text information comprises at least one of entity information for representing a target image in the image information and an image position information of the entity, attribute information for representing additional information of the entity, and rule information for representing a rule based on task logic.

[0054] The data generation method based on the multi-modal large language model provided by the present application generates predicted text information comprising at least one of entity information and image position information of the entity, attribute information, and rule information by using the trained multi-modal large language model with perception and normalized text description, so that each information in the predicted text information has an increasing dimension in the information dimension, which is beneficial to improve the performance upper limit of the multi-modal large language model and avoid hallucination output, and the entity information and attribute information are beneficial to supplement multi-modal perception. Further, through multi-dimensional monitoring during the training process, a balance between training resources and model performance can be achieved. BRIEF DESCRIPTION OF DRAWINGS

[0055] Figure 1 FIG. 1 is a flowchart of a data generation method based on a multi-modal large language model according to an embodiment of the present application.

[0056] Figure 2 FIG. 2 is a flowchart of obtaining an inquiry text sample of a business task sample according to an embodiment of the present application.

[0057] Figure 3 FIG. 3 is a flowchart of obtaining expected text sample data according to an embodiment of the present application.

[0058] Figure 4 FIG. 4 is a schematic diagram of obtaining rule information of each rule according to an embodiment of the present application.

[0059] Figure 5An example of a diagram for inquiring about the relationship between a text sample, an expected text sample, and a task.

[0060] Figure 6 An example of a diagram for adjacent text description information in expected text sample data.

[0061] Figure 7 An example of input sample data and expected text sample data based on image sample acquisition business task samples.

[0062] Figure 8 An example of a diagram for training a multi-modal large language model.

[0063] Figure 9 An example of a diagram for multi-dimensional monitoring of a predicted text sample of a multi-modal large language model by the embodiment of the present application.

[0064] Figure 10 An example of a diagram for obtaining a thought chain set by decomposing a business task in a business application scenario.

[0065] Figure 11 An example of a diagram for a data generation device based on a multi-modal large language model according to the embodiment of the present application.

[0066] Figure 12 An example of a diagram for a training device for training a multi-modal large language model according to the embodiment of the present application.

[0067] Figure 13 An example of a diagram for a data generation device based on a multi-modal large language model or a training device for training a multi-modal large language model. DETAILED DESCRIPTION

[0068] In order to make the purposes, technical means and advantages of the present application clearer and more apparent, the present application will be further described in detail below with reference to the accompanying drawings.

[0069] Applicants have found that:

[0070] 1. Compared with general detection, description, selection and other tasks, business scenarios are more of a combination of multiple dimensional information. Since some entities of the business scenario have not been presented in the application of the existing multi-modal large language model, the illusion of the multi-modal large language model generated based on the existing data is very large, and the basic performance is very low.

[0071] 2. Due to the increasing complexity of business tasks in business scenarios, involving multiple levels such as perception, rules, reasoning, and conclusions, and not a single dimension, the existing multimodal large language model's thinking link decomposition mode cannot be directly applied to various different business scenarios. The path illusion of the thinking link is relatively large. Moreover, the current multimodal large language model separates the ability to decompose business tasks, text expression, and target image boxes, and cannot construct a thinking link that interweaves text and image boxes.

[0072] In view of this, this application provides a data generation method based on a multimodal large language model. By using a multimodal task description paradigm in complex open scenarios, visual capabilities are used to supplement the perception of the multimodal large language model. Combined with a thought chain description paradigm, a standardized descriptive thought chain with interwoven text frames is formed, providing guidance for the practical application in vertical fields.

[0073] See Figure 1 As shown, Figure 1 This is a flowchart illustrating a data generation method based on a multimodal large language model, as described in an embodiment of this application. The method includes:

[0074] Step 101: Obtain query text information that represents business tasks in the business application scenario, as well as image information associated with the business tasks.

[0075] The query text information includes at least one of the following: task definition information representing the business task to be completed, and task logic information representing the business logic. The task logic information may include information such as task boundaries and task constraints.

[0076] As an example, the query text information may also include: output pattern information to characterize the desired output method, such as output format.

[0077] Step 102: Based on the trained multimodal large language model with perceptual and normalized text description, reason about the acquired query text information and image information to generate predicted text information. This predicted text information is used to represent the response information of the query text information based on the image information.

[0078] The predicted text information includes at least one of the following: entity information and image location information of the target image in the image information, attribute information for representing additional information of the entity, and rule information for representing the rules decomposed based on task logic.

[0079] As an example, the predicted text information further includes at least one of the following: explanatory information for characterizing the reasoning process of the multimodal large language model and / or the rule boundaries for characterizing the rules, and inductive information for characterizing the generalization of business tasks based on the rules for each task.

[0080] The information in the predicted text information is sequentially adjacent according to entity information, attribute information, rule information, interpretation information, and induction information, and the information dimensions of each information are sequentially increased, forming a standardized pyramid thinking chain paradigm.

[0081] As an example, the multi-modal large language model is trained in the following manner:

[0082] An inquiry text sample, an image sample, and an expected text sample of a business task sample are obtained, and a first identifier of a text information position of image position information of an entity and a second identifier of a text information position of attribute information are added in the expected text sample, where the expected text sample is used to represent a response text of the inquiry text sample based on the image sample,

[0083] The obtained inquiry text sample, image sample, and expected text sample are input into the multi-modal large language model, and a predicted text sample is generated through inference of the multi-modal large language model,

[0084] In the supervised fine-tuning training phase of the multi-modal large language model, each text information in the predicted text sample is monitored according to the entity dimension, the attribute dimension, the rule dimension, the interpretation dimension, and the induction dimension, thereby realizing multi-dimensional monitoring,

[0085] According to the multi-dimensional monitoring result, the model parameters of the multi-modal large language model are adjusted,

[0086] The training is repeatedly performed until the monitoring result reaches the expected situation.

[0087] The task definition information and the task logic information included in the inquiry text information are used to paradigmically describe the inquiry text information, and the multi-modal large language model after training is used to generate predicted text information including entity information and image position information of the entity, attribute information, and rule information, thereby realizing paradigmatic description and thinking chain information description of the predicted text, and improving the adaptability to various business scenarios compared with a general multi-modal large model.

[0088] To facilitate understanding of the embodiments of the present application, the following business task is taken as an example of alarm output, and in this embodiment, the thinking chain includes entities, attributes, rules, first interpretations thereof, inductions, and second interpretations thereof. It should be understood that the embodiments of the present application are not limited thereto.

[0089] The following describes a process of obtaining sample data for training a multi-modal large language model.

[0090] To enable the multi-modal large language model to implement the paradigm task with perception, rule, reasoning process, and conclusion as the paradigm, the predicted text information output by the multi-modal large language model is described as a pyramid thinking chain composed of entities, attributes, rules, explanations, and inductions. The inquiry text sample and the expected text sample of the business task sample are described in a paradigm natural language.

[0091] Referring to Figure 2 as shown, Figure 2 A flowchart for obtaining an inquiry text sample of a business task sample. It includes:

[0092] Step 201, obtaining business task sample data and image sample data,

[0093] As an example, the business task sample data and image sample data are obtained from a business accumulation data set, wherein the image sample data includes pre-labeled image targets and / or manually corrected image targets,

[0094] Step 202, based on the business task sample data, describing the task definition, task logic, and output mode in natural language to obtain an inquiry text sample containing task definition information, task logic information, and output mode information.

[0095] The task definition is used to represent the required business task, the task logic is used to represent the business logic required in the business task completion process, including but not limited to task boundaries, task constraints, etc., and the task definition and task logic can be adjusted based on the image sample. The output mode is used to represent the self-defined expected output mode of the conclusion, for example, the output format.

[0096] Since the capabilities of the general multi-modal large language model cannot represent entities and / or attribute levels in the business task sample, the perception task of the business task is taken as a cornerstone, which is described by entity information and / or attribute information, wherein the entity information is used to represent the entity of the target image in the image sample data, and the attribute information is used to represent the state, size, color, and other additional information of the entity.

[0097] The rules of the business task are described by rule information and first explanation information for explaining the rules, wherein the rule information is used to describe the definition of the rules, and the first explanation information can be used to describe the boundaries of the rule definition.

[0098] The reasoning process of the business task is guided by the rule information and the explanation information to generate an intermediate reasoning process of the multi-modal large language model,

[0099] The conclusion of the business task is described by induction information and second explanation information for explaining the induction information.

[0100] Referring to Figure 3shown, Figure 3 A flowchart for obtaining desired text sample data. It includes:

[0101] Step 301, obtaining business task sample data and its image sample data,

[0102] Step 302, for each business task sample, obtaining at least one rule of the business task sample,

[0103] As an example, the business task sample is disassembled according to the business logic to obtain each rule of the business task sample,

[0104] Step 303, for each rule, based on the image sample of the business task sample, sequentially performing natural language description on the entity associated with the rule, the attribute associated with the entity and / or the rule, and the first explanation associated with the rule to obtain the entity information, attribute information, rule information and first explanation information of the rule, i.e., the rule description information of the rule,

[0105] Wherein, the entity information and attribute information can be used to represent the perception dimension of the business task to identify the image location of the visual element in the image sample data and its attribute, and the rule information and the first explanation information are used to guide the multi-modal large language model to generate the intermediate reasoning process to disassemble the business task into sub-steps associated with the task logic, and finally integrate to form a complete thinking chain, so as to analyze the relationship between the visual elements obtained by perception and the rules.

[0106] Referring to Figure 4 shown, Figure 4 A schematic diagram for obtaining rule information of each rule. The business task sample is disassembled into multiple rules, based on the image sample of the business task sample, each entity in the image sample is image positioned to obtain image position information of the entity, the image position information is image coordinate information of the image block where the entity is located, each entity associated with the rule and each sub-entity associated with the rule for representing the local entity of the entity are obtained, the entity set of the rule is obtained, the entity description of each element in the entity set and the attribute description associated with the rule are obtained, and the elements in the entity set are sequentially described by the first explanation to obtain the rule description information of the rule.

[0107] As an example,

[0108] According to the rule, the entity associated with the rule is filtered out from the image sample,

[0109] The business logic analysis is sequentially performed on each independent entity associated with the rule, and if the entity has a sub-entity for representing the local entity of the entity, the sub-entity and the entity are bound,

[0110] screened entity and its bound sub-entity as an element in an entity set of the rule,

[0111] The entity and its image position information are described in a picture-text box interleaved manner, for example, the entity description is: there is an entity

coordinate information

[0112] The sub-entity and its image position information are described in a picture-text box interleaved manner to obtain entity description information of each element in the entity set, for example, the sub-entity description is: the entity part contains

coordinate information

[0113] Based on the rule, the attribute description is carried out combined with each entity description information, if the attribute boundary is simple, for example, the attribute is a classification label, the attribute description information is placed before the text information of the image position information of the entity, if the attribute boundary is complex and fuzzy, for example, the attribute has no classification label, the attribute description information is spliced to the tail of the entity description information, to obtain the attribute description information, which is generated by using a multi-modal large language model on the basis of the attribute label of the entity and then corrected by artificial to explain and summarize the attribute.

[0114] The entity description information and its attribute description information in the entity set, the rule and its first explanation description information are sequentially adjacent to form a thinking chain.

[0115] Step 304, based on the entity information, attribute information and rule information of each rule and its first explanation information, the task induction under multiple rules is carried out, and the task induction and the second explanation associated with the induction are described in natural language to obtain induction information,

[0116] Among them, the first explanation and the second explanation can be used to guide the reasoning process of the multi-modal model,

[0117] As an example, the rule description information of each rule is sequentially adjacent and then sequentially adjacent with the induction description information to obtain the thinking chain description information of the business task sample.

[0118] Step 305, the expected conclusion is described according to the output mode in the inquiry text sample,

[0119] Among them, the expected conclusion is used to represent the expected result of the multi-modal large language model output, and the expected conclusion information can be natural language description, computer language description, or a combination of the two, which is not limited by the present application.

[0120] Step 306, the information obtained in steps 303-305 is used as an expected text sample for training the multi-modal large language model,

[0121] Expected text sample example 1

[0122] Rule: Alert if not wearing gloves with both hands

[0123] Entity set involved: human, hand (from visual annotation)

[0124] Attribute: wearing gloves and whether the hand is visible (different for different number of visible hands)

[0125] First explanation: Analyze each human body in turn

[0126] Induction: Based on the above analysis, whether this picture meets the rule

[0127] Expected text sample example 2

[0128] Business task: When a human body is working at a high place, check if not wearing gloves or not wearing a safety helmet needs to be alarmed Rule: working at a high place / not wearing gloves / not wearing a safety helmet

[0129] Rule 1: working at a high place

[0130] Entity 1: human body (visual annotation)

[0131] Attribute: working at a high place, why (combine with large model to give spatial information, such as where, what is nearby)

[0132] First explanation: whether each human body is working at a high place, summarize and push the human body working at a high place

[0133] Rule 2: not wearing gloves

[0134] Attribute: wearing gloves and whether the hand is visible (different for different number of visible hands)

[0135] First explanation: analyze each human body in turn

[0136] Rule 3: not wearing a safety helmet

[0137] Attribute: wearing a safety helmet and whether the head is visible

[0138] First explanation: analyze each human body in turn

[0139] Inductive summary: only when rule 1 and rule 2 are met at the same time, or only when rule 1 and rule 3 are met at the same time, an alarm is given.

[0140] Reference Figure 5 shown, Figure 5 is a schematic diagram of the relationship between the inquiry text sample, the expected text sample and the task. An inquiry text sample includes task definition information, task logic information and output mode information, and an expected text sample includes entity information and / or attribute information, rule information and its first explanation information, induction information and its second explanation information,

[0141] wherein,

[0142] The entity information and / or the attribute information correspond to the perception task to identify the image location of the visual element in the image sample and the attribute thereof;

[0143] The rule information and the first explanation information thereof correspond to the rule task and the reasoning task to describe the rule definition of the rule and the boundary of the rule definition, and analyze the relationship between the visual element obtained by the perception task and the rule,

[0144] The induction information and the second explanation information thereof correspond to the conclusion task to describe the expected conclusion.

[0145] In order to make the expected text sample data have a standardized thinking link specification paradigm, each information in the expected text sample data is adjacent in a pyramid thinking link mode.

[0146] Referring to Figure 6 as shown, Figure 6 An illustrative view of adjacent text description information in the expected text sample data. Each text description information in the expected text sample data is sequentially adjacent in the order of entity, attribute, rule, explanation, and induction, wherein the information dimension contained by the sequentially adjacent information is sequentially increased, showing a pyramid shape, and the entity and the attribute can be derived from a business accumulation data set, for example, a conventional detection and classification data set, so as to play a visual role and focus on the entity associated with the business task; the rule, the explanation, and the induction can be derived from a data set of text information generated by a pre-trained multi-modal large language model based on image data, which can be pre-annotated, artificially corrected, and formally rewritten, and the data source and construction difficulty thereof are greater than those of the perception entity and the attribute, and the closer to the bottom of the pyramid, the less clear the boundary is, and the more generalized it is.

[0147] For ease of understanding, an example of a business task sample based on image detection of non-double gloves is taken to illustrate.

[0148] Referring to Figure 7 as shown, Figure 7 An example of input sample data and expected text sample data based on image sample acquisition of a business task sample. In the figure, the image on the right bottom is an image sample, the table on the right top is a business task, the question part on the left is an inquiry text sample, which is described according to the business task in the table, including task definition, task boundary, and output format, and the response part is an expected text sample, which sequentially includes: entity and coordinate information, attribute and coordinate information, rule and first explanation, and induction and second explanation.

[0149] Referring to Figure 8 as shown, Figure 8A schematic diagram of training of the multimodal large language model of the embodiment. The image sample, the query text sample data, and the expected text sample data are input in parallel to the multimodal large model, and the multimodal large language model is inferred to generate the predicted text sample. In order to facilitate feature extraction, in the expected text sample, the first identifier of the text information position for identifying the image position information of the entity, and the second identifier of the text information position for identifying the attribute information are added, wherein the expected text sample is used to represent the response text of the query text sample based on the image sample,

[0150] In the supervised fine-tuning training phase of the multimodal large language model, the predicted text sample is monitored in multiple dimensions according to entity information, attribute information, rule information, explanation information, and induction information, and the model parameters are adjusted according to the multi-dimensional monitoring results, so that the multimodal large language model has instruction following ability and meets the multimodal application requirements of the vertical business field.

[0151] As an example, the convergence ability of the model is monitored through instruction following monitoring and validation set monitoring to complement the business task requirements on the basis of the general ability of the multimodal large language model. Since the expected text sample data is formed according to the specified thinking link of entity-attribute-rule-explanation-induction, compared with the open learning paradigm, the knowledge absorption path of the multimodal large language model is more fine and controllable. When monitoring the validation set, the model's thinking link learning consistency can be evaluated in turn for token following, and for entity and attribute, the IOU and ACC evaluation of detection and classification can be referred to for joint monitoring, thereby facilitating the balance between training resources and model performance.

[0152] Referring to Figure 9 as shown, Figure 9 A schematic diagram of the multi-dimensional monitoring of the predicted text sample of the multimodal large language model of the embodiment. According to each dimension, the following monitoring is performed on each information in the predicted text sample. In the case where the token following loss function value monitored by the following monitoring is less than the set loss threshold, it is determined that the multimodal large language model has learned the paradigm description. Then, the validation set monitoring is performed on each dimension, and the accuracy of each dimension is evaluated using the evaluation function. In the case where the evaluation function value converges, it is indicated that the performance remains unchanged, the model learning is sufficient, and the training is stopped.

[0153] The accuracy of each dimension includes:

[0154] The target box regression accuracy for representing the accuracy of the entity dimension,

[0155] The attribute category accuracy for representing the accuracy of the attribute dimension,

[0156] A conclusion accuracy rate for characterizing the rule dimension, the explanation dimension, and the induction dimension;

[0157] The evaluation function is used to weight average the accuracy of each dimension, that is, to weight average the target frame regression accuracy, the attribute category accuracy, and the conclusion accuracy.

[0158] Referring to Figure 10 as shown, Figure 10 A schematic diagram for obtaining a thought chain set for disassembling business tasks of a business application scenario. With the accumulation of business tasks of the business scenario, the historical thought chains and historical conclusions constructed by combining the historical business rules of the historical business tasks in each historical business application scenario are accumulated into the thought chain set of the multi-modal large model, wherein any historical business application scenario corresponds to multiple historical business tasks, each historical business task corresponds to a historical thought chain, and each historical thought chain corresponds to a historical conclusion. The historical thought chain is used to provide a thinking mode for the multi-modal large model to reason, and the historical conclusion is used to provide a non-thinking mode for the multi-modal large model to reason,

[0159] In the case where the number of historical business application scenarios accumulates to a set accumulation threshold, the data in the thought chain set is used as training sample data to train the current multi-modal large model, so that the multi-modal large language model absorbs knowledge. When the performance of the multi-modal large language model does not meet the user's expectation, the training sample data is manually corrected. When the performance of the multi-modal large language model meets the user's expectation, the current multi-modal large model is used to generate thought chains and conclusions for each business task of the next business application scenario, and the generated data is returned to the thought chain set, thereby opening the use of the capabilities of the multi-modal large language model. In this way, a learnable implicit thought chain link paradigm can be provided. According to the differences in current perception, rules, reasoning, and conclusions, each business task is split differently. When the data in the thought chain set accumulates to a certain extent, it is returned to the model for learning, thereby enabling the business tasks in the new business scenario to form an automated link.

[0160] Referring to Figure 11 as shown, Figure 11 A schematic diagram of a data generation apparatus based on a multi-modal large language model according to an embodiment of the present application. The apparatus comprises:

[0161] An input module configured to obtain inquiry text information for characterizing a business task in a business application scenario and image information associated with the business task, wherein the inquiry text information comprises at least one of task definition information for characterizing a required completed business task and task logic information for characterizing business logic,

[0162] The generating module is configured to infer the obtained inquiry text information and image information based on the trained multi-modal large language model with the perception and normalized text description, and generate predicted text information for representing response information of the inquiry text information based on the image information, wherein the predicted text information comprises at least one of entity information for representing a target image in the image information and an image position information of the entity, attribute information for representing additional information of the entity, and rule information for representing a rule based on task logic.

[0163] Referring to Figure 12 as shown, Figure 12 is a schematic diagram of a training device for training a multi-modal large language model according to an embodiment of the present application. The training device comprises:

[0164] The sample obtaining module is configured to obtain inquiry text samples, image samples, and expected text samples of business task samples, and add a first identifier of a text information position of entity image position information and a second identifier of a text information position of attribute information in the expected text samples, wherein the expected text samples are used to represent response text of the inquiry text samples based on the image samples.

[0165] The prediction module is configured to input the obtained inquiry text samples, image samples, and expected text samples into the multi-modal large language model, and generate predicted text samples through inference of the multi-modal large language model.

[0166] The supervision fine-tuning module is configured to monitor each text information in the predicted text samples according to entity dimensions, attribute dimensions, rule dimensions, explanation dimensions, and induction dimensions in a supervision fine-tuning training phase of the multi-modal large language model, so as to realize multi-dimensional monitoring.

[0167] According to the multi-dimensional monitoring result, the model parameters of the multi-modal large language model are adjusted.

[0168] The training is repeatedly performed until the monitoring result reaches an expected situation and the training is stopped

[0169] Referring to Figure 13 as shown, Figure 13 is a schematic diagram of a data generation device based on a multi-modal large language model or a training device for training a multi-modal large language model. The device comprises a memory and a processor, the memory stores a computer program, and the processor is configured to execute the computer program to implement the steps of the data generation method based on the multi-modal large language model or the training method of the multi-modal large language model.

[0170] The memory can include a random access memory (RAM) and can also include a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.

[0171] The processor described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc., and can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.

[0172] The embodiment of the present application further provides a computer readable storage medium, wherein the storage medium stores a computer program, and the computer program is executed by a processor to implement steps of the data generation method based on the multi-modal large language model.

[0173] For the device / network side equipment / storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts refer to the part of the method embodiment.

[0174] In this document, the terms "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0175] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A data generation method based on a multimodal large language model, characterized in that, The method includes: Acquire query text information used to characterize business tasks in a business application scenario, as well as image information associated with the business tasks. The query text information includes at least one of: task definition information characterizing the business task to be completed, and task logic information characterizing the business logic. Based on the trained multimodal large language model with perception and normalized text description, reasoning is performed on the acquired query text information and image information to generate predicted text information that represents the query text information based on the image information. The predicted text information includes at least one of the following: entity information and image location information of the target image in the image information, attribute information for representing additional information of the entity, and rule information for representing the rules decomposed based on the task logic.

2. The data generation method according to claim 1, characterized in that, The query text information further includes: output pattern information used to characterize the desired output method. The predicted text information further includes at least one of the following: explanatory information for characterizing the reasoning process of the multimodal large language model and / or the rule boundaries for characterizing the rules, and inductive information for characterizing the summarization of business tasks based on the rules of each task. The information in the predicted text is arranged in a sequential order of entity information, attribute information, rule information, explanation information, and induction information, with the information dimension of each piece of information increasing sequentially, forming a standardized pyramid thinking chain paradigm.

3. The data generation method according to claim 1, characterized in that, The multimodal large language model is trained in the following manner: The system acquires query text samples, image samples, and expected text samples from a business task sample. It then adds a first identifier to the expected text sample to identify the location of text information in the image and a second identifier to identify the location of text information containing attribute information. The expected text sample is used to characterize the response text of the query text sample based on the image sample. The acquired query text samples, image samples, and expected text samples are input into a multimodal large language model. The multimodal large language model then infers and generates predicted text samples. During the supervised fine-tuning training phase of the multimodal large language model, the text information in the predicted text samples is monitored according to the entity dimension, attribute dimension, rule dimension, interpretation dimension, and inductive dimension. Adjust the model parameters of the multimodal large language model based on the monitoring results. Repeat the process until the monitoring results meet expectations, at which point training should be stopped.

4. The data generation method according to claim 3, characterized in that, The query text sample was obtained in the following manner: Based on any business task sample, describe the task definition, task logic, and output mode of the business task sample in natural language. The desired text sample is obtained in the following manner: Based on the business logic of this business task sample, the business task sample is broken down into at least one rule. For each rule, based on the image samples of the business task sample, the entities associated with the rule, the attributes associated with the entities and / or the rule, and the first interpretation associated with the rule are sequentially described using natural language to obtain the entity information, attribute information, rule information, and the first interpretation information of the rule. Based on the decomposed rules, the task is summarized, and the task summary and the associated second explanation are described in natural language. The first and second explanations are used to guide the reasoning process of the multimodal model. Describe the expected conclusion based on the output pattern in the query text sample.

5. The data generation method according to claim 4, characterized in that, The image samples based on the business task sample are sequentially described in natural language for the entities associated with the rule, the attributes associated with the entities and / or the rule, and the first interpretation associated with the rule, including: Based on the image samples of this business task, image localization is performed on each entity in the image samples to obtain the image position information of the entity. This image position information is the image coordinate information of the text box where the entity is located. Retrieve the entities associated with the rule, as well as the sub-entities of the local entities associated with the rule that represent the entities, to obtain the entity set of the rule. Obtain the entity description of each element in the entity set and the attribute description associated with the rule, and perform the first interpretation description on each element in the entity set in turn to obtain the rule description information of the rule; The process of summarizing the tasks based on the decomposed rules, and providing natural language descriptions of the task summarization and the associated second interpretation, includes: Based on the rule description information of each rule, the task is summarized and described to obtain summary description information. The rule description information of each rule is sequentially concatenated, and then sequentially concatenated with the inductive description information to obtain the thought chain description information of the business task sample.

6. The data generation method according to claim 5, characterized in that, The step of obtaining each entity associated with the rule, and the sub-entities of the local entities associated with the rule that represent the entities, includes: Based on this rule, entities associated with the rule are selected from the image samples. The business logic of each independent entity associated with the rule is analyzed sequentially. If an entity has sub-entities that represent local entities, then the sub-entities and their respective entities are bound together. The selected entities and their bound sub-entities are used as elements in the entity set of this rule; The step of obtaining the descriptions of each entity in the entity set and the attribute descriptions associated with the rules includes: The entity and its image location information are described using an interleaved text and image frame method, and the sub-entities and their image location information are also described using an interleaved text and image frame method, thus obtaining the entity description information of each element in the entity set. Based on this rule, attribute descriptions are generated by combining the entity description information. If the attribute boundaries are simple, the attribute description information is placed before the text information of the image location information. If the attribute boundaries are complex and ambiguous, the attribute description information is appended to the end of the entity description information to obtain the attribute description information. The description information of each entity in the entity set and its attribute description information, the rule and its first interpretation description information are sequentially connected.

7. The data generation method according to claim 4, characterized in that, The monitoring of each text information in the predicted text sample is performed according to entity dimension, attribute dimension, rule dimension, interpretation dimension, and inductive dimension, including: Follow-up monitoring is conducted for each dimension separately. If the word segmentation following loss function value monitored by the follow-up monitoring is less than the set loss threshold, it is determined that the multimodal model has learned the paradigm description. Then, validation set monitoring is performed on each dimension, and the accuracy of each dimension is evaluated using the evaluation function. Training stops when the evaluation function value converges.

8. The data generation method according to claim 7, characterized in that, The accuracy of each dimension includes: The accuracy of bounding box regression used to characterize the accuracy of entity dimensions. Attribute category accuracy is used to characterize the accuracy of attribute dimensions. The accuracy of conclusions used to characterize the accuracy of the rule dimension, explanation dimension, and inductive dimension; The evaluation function is used to perform a weighted average of the accuracy of each dimension.

9. The data generation method according to claim 2, characterized in that, The method further includes: Each historical business task, its historical thought chain, and its historical conclusions from various historical business application scenarios are accumulated into a set of thought chains for the multimodal large model. Each historical business application scenario corresponds to multiple historical business tasks, each historical business task corresponds to a historical thought chain, and each historical thought chain corresponds to a historical conclusion. Historical thought chains provide the multimodal large model with a thinking mode for reasoning, while historical conclusions provide the multimodal large model with a non-thinking mode for reasoning. When the cumulative number of historical business application scenarios reaches a set cumulative threshold, the data in the thought chain set will be used as training sample data to train the current multimodal large model. Given that the performance evaluation of the current multimodal large model meets expectations, this paper generates the thought process and conclusions for applying the current multimodal large model to various business tasks in the next business application scenario.

10. A data generation device based on a multimodal large language model, characterized in that, The device includes: The input module is used to acquire query text information representing business tasks in a business application scenario, as well as image information associated with the business tasks. The query text information includes at least one of: task definition information representing the business task to be completed, and task logic information representing the business logic. The generation module is used to reason about the acquired query text information and image information based on the trained multimodal large language model with perception and normalized text description, and generate predicted text information to represent the response information of the query text information based on the image information. The predicted text information includes at least one of the following: entity information and image location information of the target image in the image information, attribute information to represent the additional information of the entity, and rule information to represent the rules decomposed based on the task logic.

Citation Information

Cited By

  • Large language model generation process real-time intervention method based on thinking chain verification

    CN121615798A