Method for generating multimodal data samples and electronic device

By introducing prompt words and tool descriptors in LLM, and using dynamic regular expressions and reward verification, an enhanced multimodal data sample is generated, which solves the problem of unguided LLM processing chaos and improves the accuracy and efficiency of multimodal data processing.

CN120086600BActive Publication Date: 2025-08-08HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510582392.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-08
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The unguided large language model (LLM) is confusing and unreasonable instructor generation when processing multimodal data, resulting in inaccuracy and inefficiency in data processing.

Method used

By introducing prompt words in LLM, obtaining tool call instructions, adding tool description words and verification methods, selecting appropriate models to perform tasks, and generating enhanced multimodal data samples through reinforcement learning, and using dynamic regular expressions and reward verification to ensure the accuracy of data processing results.

Benefits of technology

It improves the accuracy and efficiency of multimodal data processing, ensures the logic and coherence of task processing, and enhances the adaptability and performance of LLM in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086600B_ABST
    Figure CN120086600B_ABST
Patent Text Reader

Abstract

The present application discloses a method and electronic device for generating multimodal data samples, relating to the field of data processing technology. The method comprises: inputting original multimodal data into an LLM, obtaining a tool call instruction for the original multimodal data based on a prompt, adding a tool description to the prompt according to the tool call instruction, selecting a model to be called according to the tool call instruction, executing a task corresponding to the task information, obtaining a data processing result, performing data verification on the data processing result using a verification method in the tool description and a tool call response format in the tool description, and if the data verification is successful, performing reinforcement learning on the data processing result and the original multimodal data to generate an enhanced multimodal data sample. This method overcomes the defects of chaotic and unreasonable generation of unguided LLM instructions and improves the accuracy and efficiency of multimodal data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method for generating multimodal data samples and an electronic device. Background Art

[0002] Artificial intelligence is widely used in data processing, and some AI technologies rely on the Large Language Model (LLM). When processing multimodal data (such as text, images, audio, and radar data), unguided LLMs often generate confusing and irrational instructions. Summary of the Invention

[0003] The purpose of this application is to provide a method and electronic device for generating multimodal data samples, so as to overcome the defects of chaotic and unreasonable generation of unguided LLM instructions and improve the accuracy and efficiency of multimodal data processing.

[0004] In a first aspect, the present application provides a method for generating a multimodal data sample, the method comprising:

[0005] Inputting the original multimodal data into the large language model (LLM), and obtaining a tool calling instruction for the original multimodal data based on a prompt word, wherein the tool calling instruction includes task information;

[0006] Adding a tool description word to the prompt word according to the tool call instruction, wherein the tool description word includes a tool call response format and a verification method;

[0007] According to the tool calling instruction, a called model is selected, the task corresponding to the task information is executed, and a data processing result is obtained;

[0008] Performing data verification on the data processing result by referring to the tool call response format through the verification method;

[0009] If the data verification is successful, reinforcement learning is performed on the data processing results and the original multimodal data to generate enhanced multimodal data samples.

[0010] In one embodiment, if the model selected for invocation according to the tool invocation instruction is the external tool corresponding to the tool invocation instruction, the tool execution result of the external tool is used as the data processing result:

[0011] The data processing result is verified by the verification method with reference to the tool call response format, including:

[0012] The data processing result is used as a tool parameter and put into the parameter data set corresponding to the external tool. Based on the parameter data set of the external tool, it is determined whether the tool name and tool parameters in the tool description word conform to the dynamic regular expression.

[0013] If the tool name and tool parameters in the tool description both conform to the dynamic regular expression, the data verification is determined to be successful; if at least one of the tool name and tool parameters in the tool description does not conform to the dynamic regular expression, the data verification is determined to be failed.

[0014] In one embodiment, after verifying the data processing result by referring to the tool call response format through the verification method,

[0015] If the data verification fails: the original multimodal data is set as a rejection sample.

[0016] In one embodiment, if the model selected for invocation according to the tool invocation instruction is the LLM, the tool invocation response format is reward verification:

[0017] The data processing result is verified by the verification method with reference to the tool call response format, including:

[0018] Based on reward verification, determine the degree of match between the data processing result and the true sample value of the original multimodal data:

[0019] If the matching degree reaches a preset matching threshold, the data verification is determined to be successful;

[0020] If the matching degree does not reach the preset matching threshold, it is determined that the data verification has failed.

[0021] In one embodiment, determining the degree of match between the data processing result and the sample true value of the original multimodal data based on reward verification includes:

[0022] Based on reward verification, the matching degree between the data processing result and the sample true value of the original multimodal data is determined by at least one of lexical semantic matching, syntactic structure similarity, semantic context similarity, and task constraint matching.

[0023] In one embodiment, determining the degree of match between the data processing result and the sample true value of the original multimodal data based on reward verification includes:

[0024] The degree of matching between the data processing result and the sample true value of the original multimodal data is determined by a weighted similarity function of lexical semantic matching, syntactic structure similarity, semantic context similarity, and task constraint matching, wherein:

[0025] The lexical semantic matching degree is determined by calculating the similarity between the data processing result and the ground truth corresponding to the original multimodal data through word embedding and / or word frequency statistics TF-IDF;

[0026] The syntactic structure similarity is determined by calculating the syntactic structure similarity between the data processing result and the original multimodal data through dependency parsing and tree edit distance.

[0027] The semantic context similarity is determined by calculating the similarity between the data processing result and the original multimodal data in the context context by using sentence embedding;

[0028] The task constraint matching degree is determined by the named entity matching rate and / or key field consistency between the data processing result and the original multimodal data.

[0029] In one embodiment, if the data verification fails, the method further includes:

[0030] Adjust the prompt words;

[0031] The tool calling instruction of the original multimodal data is re-acquired based on the adjusted prompt word, and the external tool to be called is selected according to the tool calling instruction of the original multimodal data, and the task corresponding to the task information is executed to obtain the data processing result.

[0032] In one embodiment, performing reinforcement learning on the data processing results and the original multimodal data includes:

[0033] Reinforcement learning is performed on the data processing results and the original multimodal data through a random sampling strategy and / or a shuffle random sorting strategy to obtain enhanced multimodal data samples, wherein the original multimodal data includes: at least two of: text data, image data, audio data, and radar data.

[0034] In a second aspect, the present application provides a method for processing multimodal data, the method comprising: training the LLM using the enhanced multimodal data samples generated in the first aspect, and processing a majority of the multimodal data using the trained LLM.

[0035] In a third aspect, the present application provides an electronic device comprising at least one processor and at least one memory, wherein the at least one processor is used to execute a computer program in the at least one memory to implement: the method described in any one of the first aspect or the method described in the second aspect.

[0036] In one embodiment, the electronic device further comprises:

[0037] an input device for acquiring the original multimodal data,

[0038] A display device is used to display the processing result of the LLM.

[0039] In a fourth aspect, the present application provides a device for generating a multimodal data sample, the device comprising:

[0040] an acquisition module, configured to input the original multimodal data into a large language model (LLM), and acquire a tool call instruction for the original multimodal data based on a prompt word, wherein the tool call instruction includes task information;

[0041] an adding module, configured to add a tool description word to the prompt word according to the tool calling instruction, wherein the tool description word includes a tool calling response format and a verification method;

[0042] A selection module is used to select a model to be called according to the tool calling instruction, execute the task corresponding to the task information, and obtain a data processing result;

[0043] A verification module, configured to verify the data processing result by using the verification method and referring to the tool call response format;

[0044] The generation module is used to perform reinforcement learning on the data processing results and the original multimodal data to generate enhanced multimodal data samples if the data verification is successful.

[0045] The embodiment of the present application utilizes LLM to determine the tool call instructions of multimodal data through prompt words, selects the calling model based on the tool call instructions, and generates enhanced multimodal data samples through reinforcement learning when the data processing results match the tool call response format in the tool call instructions. Since the thinking chain of multimodal tool calls is utilized in the prompt words, the tool call instructions are accurately generated, guiding the LLM to process complex instructions in an orderly manner. Compared with the unguided LLM, the logic and accuracy of LLM task processing are greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1A schematic diagram of a flow chart of a method for generating multimodal data samples provided in an embodiment of the present application;

[0047] Figure 2 An internal structure diagram of an electronic device provided in an embodiment of the present application;

[0048] Figure 3 This is a structural block diagram of a device for generating multimodal data samples provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] The present application will be described in detail below in conjunction with the specific embodiments shown in the accompanying drawings, but these embodiments do not limit the present application. Structural, methodological, or functional changes made by ordinary technicians in this field based on these embodiments are included in the scope of protection of the present application.

[0050] Please refer to Figure 1 , an embodiment of the present application provides a method for generating a multimodal data sample, the method for generating a multimodal data sample comprising the following steps:

[0051] Step 101: Inputting the original multimodal data into the LLM, and obtaining a tool call instruction of the original multimodal data based on a prompt word, wherein the tool call instruction includes task information;

[0052] Step 102: Adding a tool description to the prompt word according to the tool call instruction, wherein the tool description includes a tool call response format and a verification method;

[0053] Step 103: According to the tool calling instruction, select the called model, execute the task corresponding to the task information, and obtain the data processing result;

[0054] Step 104: Perform data verification on the data processing result using the verification method in the tool description and the tool call response format in the tool description;

[0055] Step 105: If the data verification is successful, reinforcement learning is performed on the data processing results and the original multimodal data to generate enhanced multimodal data samples.

[0056] The formatting of prompt words helps to respond to text instructions and determine whether to call external tools, thereby dividing the subsequent task processing logic.

[0057] The subsequent task processing logic corresponding to LLM in the two cases of calling external tools and not calling external tools includes: after calling the external tool, LLM can prepare to answer subsequent instructions or questions based on the output of the external tool; when there is no need to call the tool, it relies on LLM's own knowledge base and reasoning ability to directly give answers or execution results, so that the model has a clear action guide regardless of the situation in the entire task process, ensuring the consistency and integrity of task processing.

[0058] For example, when it is determined that an external tool needs to be called, the prompt word can be used<function call / > Tags package tool call information. Tool call information can include API_name and API_params. API_name can be the tool name, and API_params can be the tool parameter name. This improves the standardization and parsability of prompts. API_params can also be a dictionary containing tool parameter names and values. LLM can parse the tool call information in prompts to generate tool call instructions.

[0059] Exemplarily, the tool call instruction may include at least one of task information, LLM model role, and tool descriptors. The tool descriptors may include, but are not limited to, at least one of the following information: tool name, operating parameters, tool call response format, verification method, subsequent processing logic, etc.

[0060] For example, the tool descriptor can appear after the initial prompt and inform the LLM of the available external tools. The tool descriptor can be formatted as a JSON document. The tool call response format in the tool descriptor can specify each tool's function, its parameters, and the corresponding meanings for each parameter.

[0061] For example, using the semantic-sam and clip tools as examples, the description format of each tool can adopt a fixed format, including type, function, name, description, parameters, and required parts. For example, type represents the tool function type; function represents the function of the tool function called by the tool, and function can include name and description, where name can be the tool name and description can be a description of the tool's core functions; parameters represent the parameters included in the tool, assuming that the parameters include image and boxes, where: image represents the path of the input image, and its format is string type (such as a local path or URL); boxes represents the coordinates of the target area, and supports string format or JSON array (such as [[x1,y1,x2,y2], ...]); required indicates that the image and boxes parameters must be provided when calling.

[0062] For example, the description is "According to the coordinate information in the input boxes parameter, the image is segmented, and then different fine-grained segmentation results are output." For another example, the description is "The path to the image used to pass to the model as input."

[0063] Before inputting the original multimodal data into LLM, the multimodal data can also be preprocessed through data cleaning to improve the quality of the original multimodal data.

[0064] If the tool calling instruction indicates calling a corresponding external tool to process the original multimodal data, the above verification method may include a Rule-based verification submodule.

[0065] The Rule-based verification submodule extracts tool call information from the tool call instruction and matches the tool call response format through a regular expression. If the match is successful, the data verification is determined to be successful. If the match fails, the data verification is determined to be failed.

[0066] The subsequent processing logic may include, but is not limited to, introducing data enhancement strategies such as rejection strategy, random sample strategy, and shuffle random sorting. These data enhancement strategies can improve the quality of the sample data set, thereby further improving model performance.

[0067] ; ; ;

[0068] The pre-trained LLM is fed as input X, and the LLM generates tool call instructions Y based on the multimodal data and the prompt word. ;

[0069] Where Y represents a tool call instruction, X represents multimodal data, and Prompt represents a prompt word. Based on the prompt word, an external tool is used to process the multimodal data to obtain the tool call instruction. The tool description word may include, but is not limited to, at least one of the following information: tool name, operating parameters, tool call response format, verification method, subsequent processing logic, etc. Executing the tool call instruction Y yields the tool execution result T.

[0070] Input the tool execution result T into the LLM model to generate the answer P corresponding to the multimodal data. The expression is as follows:

[0071] ;

[0072] Where Y represents a tool call instruction, X represents multimodal data, and Prompt represents a prompt. Based on the prompt, an external tool is used to process the multimodal data to obtain the tool call instruction. The tool description may include, but is not limited to, at least one of the following information: tool name, operating parameters, tool call response format, verification method, and subsequent processing logic. Executing the tool call instruction Y yields the tool execution result T. If both the tool name and tool parameters in the tool description conform to the dynamic regular expression, data verification is considered successful. If at least one of the tool name and tool parameters in the tool description does not conform to the dynamic regular expression, data verification is considered a failure.

[0073] Static rules typically predefine a fixed set of regular expressions to perform pattern matching on input data. However, static rules are limited in that they are difficult to handle diverse input formats, especially when tool parameter requirements are subject to certain variability. Therefore, the embodiments of this application introduce dynamic regular expressions to automatically adjust matching rules based on the requirements of different tools, thereby enhancing the adaptability of tool parameters.

[0074] For example, the grounding_dino tool requires the tool parameter caption, while the sam tool (image segmentation tool) may require the tool parameters boex (segmentation box, indicating the specified area in the segmented image) and image_url (indicating the segmented image path). The function matching rules described below will be adjusted according to different tool names.

[0075] For example, take the tool sam as an example as follows:

[0076] Large model output:<Function Calling / > [{'API_name': 'sam', 'API_params': {'boex': [0.1,0.0.25,0.4,0.6]}}]<Function Calling / > ; Dynamic regular rule: "API_name"\s*:\s*"(?P <name>[^"]+)"\s*,\s*"API_params"\s*:\s*\{(?P <params>.*?)\}, this rule will check whether API_name matches sam and whether API_params contains boex.

[0077] For example, for a tool name t and a parameter dataset P={p1, p2, …, pn}, the following adaptive validation function can be used to determine whether the tool name t and the parameter dataset P conform to the dynamic regular expression:

[0078] ;

[0079] Where Regex(pi) represents the regular expression rule for each parameter pi, and ~ represents a match operation. If the tool name t and all parameters pi satisfy their corresponding regular expression rules, the data validation result is considered "Valid"; otherwise, the data validation result is considered "Invalid".

[0080] In an optional embodiment of the present application, it is assumed that there is the following tool call information:

[0081] <Function Calling / > [{'API_name': 'grounding_dino', 'API_params': {'caption': 'man on a horse'}}]<Function Calling / > .

[0082] The matching rules for tool parameters are:

[0083] API_name must be grounding_dino, and API_params must contain at least the caption field. The value of the caption field must be a non-empty string and conform to the basic description format (for example, it cannot be random characters).

[0084] The dynamic regular expression corresponding to the matching rule of the above tool parameters is expressed as: "API_name"\s*:\s*"grounding_dino"\s*,\s*"API_params"\s*:\s*\{\s*"caption"\s*:\s*".+?"\s*\}

[0085] The dynamic regularization rules for matching the above tool parameters include detection:

[0086] API_name matches grounding_dino.

[0087] Whether API_params contains a caption, and the caption value is a string containing at least one character.

[0088] The caption value "man on a horse" is reasonable and meets the semantic requirements. Therefore, the data validation succeeds and the tool call command is judged as Valid.

[0089] In an optional embodiment of the present application, the dynamic regular rule matching process corresponding to the matching rule of the above tool parameters is expressed as follows:<Function Calling / > [{'API_name': 'unknown_tool', 'API_params':{'caption': 'man on a horse'}}]<Function Calling / >

[0090] Because the API_name is unknown_tool (a tool other than grounding_dino), it is not in the list of tools allowed by the matching rule of the tool parameter API_name. Since grounding_dino cannot be matched, the matching fails and the verification fails.

[0091] In an optional embodiment of the present application, the dynamic regular rule matching process corresponding to the matching rule of the above tool parameters is expressed as follows:<Function Calling / > [{'API_name': 'grounding_dino', 'API_params':{}}]<Function Calling / >

[0092] Because API_params is empty and caption is missing, the match fails and the data verification fails.

[0093] In this embodiment of the present application, if the dynamic regular expression data verification fails: the original multimodal data and the corresponding external tool are set as rejection samples. This can avoid overfitting of the model and improve the adaptability of the model in complex multimodal scenarios.

[0094] If, based on the above tool call instructions, it is determined that no external tool needs to be called and only the LLM model is selected for calling, the corresponding tool call response format is reward verification.

[0095] Exemplarily, based on reward verification, determining the degree of match between the data processing results of LLM and the sample true value of the original multimodal data can be achieved as follows: calculating the cosine similarity between the data processing results of LLM and the sample true value groundtruth. If the cosine similarity exceeds the preset matching threshold, it is determined that the data verification is successful and the data processing results generated by LLM are correct.

[0096] Exemplarily, based on reward verification, the degree of match between the data processing result and the true sample value of the original multimodal data may be determined by at least one of the following four dimensions:

[0097] Lexical semantic matching, syntactic structure similarity, semantic context similarity, and task constraint matching.

[0098] Optionally, the semantics of these multiple dimensions can be weighted and aligned to achieve a match between the data processing results and the true sample values of the original multimodal data. This can more comprehensively consider the differences between different modalities and semantic levels, making the matching more accurate and flexible.

[0099] For example, among the above four dimensions:

[0100] Lexical similarity can be determined by calculating the similarity between the processed data and the ground truth corresponding to the original multimodal data using word embedding (such as Word2Vec and FastText) and / or TF-IDF. For example, "car" and "automobile" are highly similar at the lexical semantic level, while "car" and "banana" are unrelated at the lexical semantic level.

[0101] Syntactic similarity can be determined by calculating the syntactic similarity between the data processing result and the original multimodal data through dependency parsing and tree edit distance.

[0102] For example, "The boy eats an apple." and "An apple is eaten by the boy." are similar in vocabulary but different in grammatical structure. Syntactic structure similarity analysis can better capture this difference.

[0103] Contextual Similarity can be determined by using sentence embedding (such as pre-trained language models such as BERT or GPT) to calculate the similarity between the data processing results and the original multimodal data in the context of each sentence.

[0104] For example, "He went to the bank to withdraw money" and "He sat by the bank of the river" are similar in vocabulary but different in context. Contextual embeddings can better distinguish these contextual semantic differences. Task-specific similarity can be determined by the named entity matching rate and / or key field consistency between the processed data and the original multimodal data.

[0105] For specific tasks (such as question answering, code generation, summarization, etc.), matching metrics based on task features are introduced. For example, in information extraction tasks, named entity overlap or key field matching can be used as task features to calculate matching.

[0106] For example, in a table information extraction task, "Price: $100" and "The cost is 100 dollars" may not be similar under traditional vocabulary matching, but based on task characteristics, it can be identified that both represent the same price information and are therefore considered a high match.

[0107] In the embodiment of the present application, for the LLM data processing result r and the sample true value g, we define a weighted similarity function S(r,g):

[0108] ;

[0109] Simlex(r,g) represents the lexical semantic similarity, Simsyn(r,g) represents the syntactic structure similarity, Sim ctx (r,g) represents the semantic context similarity, Sim task (r, g) represents the task constraint matching degree, w1, w2, w3, and w4 are the weight coefficients of the corresponding dimensions, which can be adjusted according to the specific application scenario. For example, the sum of w1, w2, w3, and w4 is 1, and the granularity adjusted according to the specific application scenario is 0.01.

[0110] This weighting function can more comprehensively evaluate the degree of match between the generated results and the real data by considering the semantic importance of each dimension. If S(r,g) exceeds the predetermined matching threshold θ, the data processing result is considered to meet the requirements and can be output. Otherwise, it is judged as a failure result and requires further adjustment and verification. For example, the prompt word can be adjusted and the process returns to the step of inputting the multimodal data into the LLM. Based on the adjusted prompt word, the tool call instruction for the multimodal data is retrieved. According to the new tool call instruction, the external tool to be called is selected, the task corresponding to the task information is executed, and the data processing result is obtained.

[0111] In the embodiments of this application,

[0112] Suppose the true answer is "Paris is the capital of France."

[0113] The following is a successful example:

[0114] Suppose the answer generated by the LLM (i.e., the result of data processing) is as follows: "The capital of France is Paris."

[0115] The matching calculation results in different dimensions are as follows:

[0116] Lexical semantic similarity: High (the sentence contains the same keywords "capital" and "Paris")

[0117] Syntactic structure similarity: High (although the word order is different, the grammatical components are the same)

[0118] Semantic context similarity: High (BERT indicates that the two sentences convey the same semantics)

[0119] Task Constraint Match: High (In the question-answering task, the core information "Paris" is correct)

[0120] Since S(r,g) exceeds the set threshold θ, the output is judged to be (Valid) and the data processing result is successful, and a valid output is obtained.

[0121] The following are examples of failures:

[0122] Suppose the answer generated by the LLM (i.e., the result of data processing) is as follows: "France is a beautiful country with many tourists visiting Paris."

[0123] Although "Paris" still appears,

[0124] Lexical semantic similarity: partial match (contains "Paris" but adds extra irrelevant information)

[0125] Syntactic structure similarity: Low (the sentence structures are completely different)

[0126] Semantic context similarity: Medium (although it mentions "Paris", it does not directly answer the question about "capital")

[0127] Task Constraint Match: Low (question not answered correctly)

[0128] Because S(r, g) does not reach the threshold θ, the output is considered invalid. The data processing result failed and no result can be output. Further adjustment of the LLM generation strategy is required. The LLM generation strategy can be adjusted based on the multimodal data samples and / or the prompt words.

[0129] If the data verification is successful, reinforcement learning is performed on the data processing results and the original multimodal data to generate enhanced multimodal data samples, which may include:

[0130] Reinforcement learning is performed on the data processing results and the original multimodal data through a random sampling strategy and / or a shuffle random sorting strategy to obtain enhanced multimodal data samples, wherein the original multimodal data includes: at least two of: text data, image data, audio data, and radar data.

[0131] For example, the random sampling strategy can include randomly sampling external tools added to the prompt word. In addition to the external tools required for the current conversation, the tool call instruction randomly samples several additional tools from the external toolset to increase the adaptability of the LLM to different tool set variants.

[0132] For example, the shuffle random sorting strategy can be: randomly shuffling the order of samples of multiple original multimodal data in the data set. This can break the original order of the data, prevent overfitting during the LLM training process, and improve the generalization ability of the LLM under different data arrangements.

[0133] For example, a large amount of raw data covering multiple modalities (such as text data, image data, audio data, radar data, etc.) is collected and clarified.

[0134] For example, large amounts of raw data are standardized to meet the model's input requirements. All image data is uniformly adjusted to a specific resolution, and text is preprocessed through language and layout standardization to obtain cleaned raw multimodal data.

[0135] Based on the type and application scenario of the original multimodal data, an appropriate LLM (such as the GPT series, BERT, etc.) is selected. Following the aforementioned method for generating multimodal data samples, prompts and corresponding tool invocation instructions are designed for the LLM, forming a tool invocation chain. Matching checks based on both rule-based and reward-model approaches enhance the LLM's reasoning capabilities. For samples that refuse to answer, a dedicated sample library is established to manage them and avoid overfitting. During data augmentation, random sampling and shuffle operations are used to control the degree and scope of perturbations, ensuring the rationality and usability of data sampling while ensuring the randomness and sufficiency of samples and avoiding excessive duplication or omission of data. The generation of the aforementioned multimodal data samples can successfully address the shortcomings of multimodal dataset construction methods, improve the quality and diversity of datasets, and ultimately enhance the performance of LLMs, promoting the development of AI in multimodal data processing.

[0136] Based on the same inventive concept, the embodiment of the present application further provides an electronic device, such as Figure 2 As shown, the electronic device includes at least one processor and at least one memory. The memory can be used to store computer programs. The computer programs can include instructions and data to implement the steps of any of the above methods. The memory can be a random access memory, read-only memory, non-volatile, programmable ROM, erasable PROM, electrically erasable, flash memory, optical memory, and registers. The processor can be a general-purpose processor. The general-purpose processor can be a processor that performs specific steps and / or operations by reading and executing computer programs stored in the memory. The general-purpose processor may use data stored in the memory during the execution of the steps and / or operations. The general-purpose processor can be a central processing unit, an ASIC, an FPGA, etc. During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or by instructions in software form. The method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.

[0137] The above-mentioned electronic device may further include:

[0138] An input device for acquiring raw multimodal data. Exemplarily, the input device includes, but is not limited to, at least one of a keyboard, a touch panel, a voice input device, and an image sensor. The raw multimodal data may include, but is not limited to, text, graphics, voice, images, and video.

[0139] The display device is used to display the processing results of the LLM. The display device can be a display screen.

[0140] The memory, processor, input device, and display device may be connected via a system bus.

[0141] The present application also provides a device for generating multimodal data samples, such as Figure 3 As shown, the device includes:

[0142] An acquisition module 301 is configured to input the original multimodal data into a large language model (LLM) and acquire a tool call instruction for the original multimodal data based on a prompt word, wherein the tool call instruction includes task information.

[0143] An adding module 302 is configured to add a tool description word to the prompt word according to the tool call instruction, wherein the tool description word includes a tool call response format and a verification method;

[0144] The selection module 303 is used to select a model to be called according to the tool calling instruction, execute the task corresponding to the task information, and obtain the data processing result;

[0145] Verification module 304, used to verify the data processing result by referring to the tool call response format through verification mode;

[0146] The generation module 305 is used to perform reinforcement learning on the data processing results and the original multimodal data to generate enhanced multimodal data samples if the data verification is successful.

[0147] In an optional embodiment, if the model selected for calling is the external tool corresponding to the tool calling instruction according to the tool calling instruction, the tool execution result of the external tool is used as the data processing result, and the verification module 304 performs data verification on the data processing result by referring to the tool calling response format through a verification method, specifically including: putting the data processing result as a tool parameter into the parameter data set corresponding to the external tool; judging whether the tool name and tool parameters in the tool description word conform to the dynamic regular expression based on the parameter data set of the external tool; if the tool name and tool parameters in the tool description word conform to the dynamic regular expression, the data verification is judged to be successful; if at least one of the tool name and tool parameters in the tool description word does not conform to the dynamic regular expression, the data verification is judged to be failed.

[0148] In an optional embodiment, after the verification module 304 verifies the data processing result by referring to the tool call response format through a verification method, if the data verification fails: the verification module 304 is further configured to set the original multimodal data as a rejection sample.

[0149] In an optional embodiment, if the selection module 303 selects the LLM model to be called according to the tool call instruction, the tool call response format is reward verification: through the verification method, the data processing result is verified with reference to the tool call response format, including: based on the reward verification, determining the matching degree between the data processing result and the sample true value of the original multimodal data: if the matching degree reaches the preset matching threshold, the data verification is judged to be successful; if the matching degree does not reach the preset matching threshold, the data verification is judged to be failed.

[0150] In an optional embodiment, based on reward verification, the matching degree between the data processing result and the sample true value of the original multimodal data is determined, including: based on reward verification, the matching degree between the data processing result and the sample true value of the original multimodal data is determined through at least one of lexical semantic matching, syntactic structure similarity, semantic context similarity, and task constraint matching.

[0151] In an optional embodiment, based on reward verification, the matching degree between the data processing result and the sample true value of the original multimodal data is determined, including: determining the matching degree between the data processing result and the sample true value of the original multimodal data through a weighted similarity function of lexical semantic matching, syntactic structure similarity, semantic context similarity and task constraint matching, wherein: lexical semantic matching is determined by calculating the similarity between the data processing result and the ground truth corresponding to the original multimodal data through word vector embedding Word Embedding and / or word frequency statistics TF-IDF; syntactic structure similarity is determined by calculating the similarity between the data processing result and the original multimodal data in syntactic structure through dependency parsing and tree edit distance Tree Edit Distance; semantic context similarity is determined by calculating the similarity between the data processing result and the original multimodal data in the context context through sentence embedding Sentence Embedding; task constraint matching is determined by the named entity matching rate and / or key field consistency between the data processing result and the original multimodal data.

[0152] In an optional embodiment, the device also includes an adjustment module. If the data verification fails, the adjustment module adjusts the prompt word; based on the adjusted prompt word, the tool call instruction of the original multimodal data is re-obtained, and according to the tool call instruction of the original multimodal data, the external tool to be called is selected to execute the task corresponding to the task information to obtain the data processing result.

[0153] In an optional embodiment, the generation module 305 performs reinforcement learning on the data processing results and the original multimodal data, specifically including: performing reinforcement learning on the data processing results and the original multimodal data through a random sampling strategy and / or a shuffle random sorting strategy to obtain enhanced multimodal data samples, wherein the original multimodal data includes: at least two of: text data, image data, audio data, and radar data.

[0154] The present application also provides a multimodal data processing device, which includes:

[0155] A training module, configured to train the LLM using the enhanced multimodal data samples generated by the multimodal data sample generating device as described above;

[0156] The processing module is used to process the majority of multimodal data using the trained LLM.

[0157] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or solid-state drive (SSD).

[0158] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise," "include," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0159] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.

[0160] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application are included in the scope of protection of the present application.< / params> < / name>

Claims

1. A method for generating multimodal data samples, characterized in that: The method comprises: Inputting the original multimodal data into the large language model (LLM), and obtaining a tool calling instruction for the original multimodal data based on a prompt word, wherein the tool calling instruction includes task information; Adding a tool description word to the prompt word according to the tool call instruction, wherein the tool description word includes a tool call response format and a verification method; According to the tool call instruction, the called model is selected, the task corresponding to the task information is executed, and the data processing result is obtained. The data processing result is verified by referring to the tool call response format through the verification method. If, according to the tool calling instruction, the model selected for calling is the external tool corresponding to the tool calling instruction, the tool execution result of the external tool is used as the data processing result: The data processing result is verified by the verification method with reference to the tool call response format, including: The data processing result is used as a tool parameter and put into the parameter data set corresponding to the external tool. Based on the parameter data set of the external tool, it is determined whether the tool name and tool parameters in the tool description word conform to the dynamic regular expression. If the tool name and tool parameters in the tool description both conform to the dynamic regular expression, the data verification is determined to be successful; if at least one of the tool name and tool parameters in the tool description does not conform to the dynamic regular expression, the data verification is determined to be failed; If the data verification is successful, reinforcement learning is performed on the data processing results and the original multimodal data to generate enhanced multimodal data samples.

2. The method according to claim 1, wherein After verifying the data processing result by referring to the tool call response format through the verification method, If the data verification fails: the original multimodal data is set as a rejection sample.

3. The method according to claim 1, wherein If the model selected for invocation according to the tool invocation instruction is the LLM, the tool invocation response format is reward verification: The data processing result is verified by the verification method with reference to the tool call response format, including: Based on reward verification, determine the degree of match between the data processing result and the true sample value of the original multimodal data: If the matching degree reaches a preset matching threshold, the data verification is determined to be successful; If the matching degree does not reach the preset matching threshold, it is determined that the data verification has failed.

4. The method according to claim 3, wherein Determining a degree of match between the data processing result and a true sample value of the original multimodal data based on reward verification includes: Based on reward verification, the matching degree between the data processing result and the sample true value of the original multimodal data is determined by at least one of lexical semantic matching, syntactic structure similarity, semantic context similarity, and task constraint matching.

5. The method according to claim 3, wherein Determining a degree of match between the data processing result and a true sample value of the original multimodal data based on reward verification includes: The degree of matching between the data processing result and the sample true value of the original multimodal data is determined by a weighted similarity function of lexical semantic matching, syntactic structure similarity, semantic context similarity, and task constraint matching, wherein: The lexical semantic matching degree is determined by calculating the similarity between the data processing result and the sample true value corresponding to the original multimodal data through word vector embedding and / or word frequency statistics TF-IDF; The syntactic structure similarity is determined by calculating the syntactic structure similarity between the data processing result and the true value of the sample of the original multimodal data through dependency parsing and tree edit distance. The semantic context similarity is determined by using sentence embedding to calculate the similarity between the data processing result and the true value of the sample corresponding to the original multimodal data in the context of the sentence; The task constraint matching degree is determined by the named entity matching rate and / or key field consistency between the data processing result and the sample true value corresponding to the original multimodal data.

6. The method according to claim 3, wherein If the data verification fails, the method further includes: Adjust the prompt words; The tool calling instruction of the original multimodal data is re-acquired based on the adjusted prompt word, and the external tool to be called is selected according to the tool calling instruction of the original multimodal data, and the task corresponding to the task information is executed to obtain the data processing result.

7. The method according to claim 1, wherein Performing reinforcement learning on the data processing results and the original multimodal data, including: Reinforcement learning is performed on the data processing results and the original multimodal data through a random sampling strategy and / or a shuffle random sorting strategy to obtain enhanced multimodal data samples, wherein the original multimodal data includes: at least two of: text data, image data, audio data, and radar data.

8. A method for processing multimodal data, characterized in that: include: The LLM is trained by the enhanced multimodal data samples according to any one of claims 1 to 7, The trained LLM is used to process most multimodal data.

9. An electronic device, characterized in that: The system comprises at least one processor and at least one memory, wherein the at least one processor is configured to execute a computer program in the at least one memory to implement the method according to any one of claims 1 to 8.

10. The electronic device according to claim 9, wherein Also includes: an input device for acquiring the original multimodal data, A display device is used to display the processing result of the LLM.

Citation Information

Patent Citations

  • Dangerous behavior identification and early warning method based on multi-modal analysis

    CN119360278A

  • Question answering method and system based on tool calling, electronic equipment and storage medium

    CN119623630A