Sample construction method, large model training method, device, electronic device and medium

By constructing scrambled reasoning chains containing abnormal conditions and training samples of reference text, the problems of insufficient error recognition and complex reasoning capabilities of large multimodal models are solved, and their performance in multimodal understanding and high-order reasoning tasks is improved.

CN119761515BActive Publication Date: 2025-09-26BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411898025.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-09-26
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Existing large multimodal models lack the ability to identify errors, reflect and correct errors, and perform poorly in dealing with complex tasks, which limits their application potential in high-level reasoning tasks.

Method used

By constructing training samples, including a target reasoning chain, which contains a scrambled reasoning chain of predetermined abnormal conditions and a reference text for describing abnormal information, the comprehensive reasoning ability of the multimodal large model is improved.

Benefits of technology

It improves the error recognition, reflection and correction capabilities of large multimodal models, and enhances their reasoning capabilities in multimodal understanding, especially in complex tasks and high-order reasoning tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119761515B_ABST
    Figure CN119761515B_ABST
Patent Text Reader

Abstract

The present disclosure provides a sample construction method, a large model training method, an apparatus, an electronic device, and a medium, relating to the fields of computer vision, deep learning, large models, and the like. A specific implementation scheme comprises: determining an initial reasoning chain for answering a question text based on associated images and question text; scrambling the initial reasoning chain to obtain a scrambled reasoning chain that satisfies a predetermined abnormality condition; determining a target reasoning chain based on the initial and scrambled reasoning chains, based on the large model; the target reasoning chain includes a reference text and the scrambled reasoning chain, the reference text being used to describe abnormal information in the scrambled reasoning chain; and constructing a training sample based on the image, question text, and target reasoning chain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of computer vision, deep learning, large models, and more specifically, provides a sample construction method, a multimodal large model training method, a sample construction device, a multimodal large model training device, an electronic device, a storage medium, and a computer program product. Background Art

[0002] With the development of artificial intelligence, more and more application scenarios will use large multimodal models to perform tasks, such as using large multimodal models for question answering, translation, and other tasks. It is understandable that improving the reasoning capabilities of large multimodal models can further enhance their application prospects. Summary of the Invention

[0003] The present disclosure provides a sample construction method, a training method for a large multimodal model, a sample construction device, a training device for a large multimodal model, an electronic device, a storage medium, and a computer program product.

[0004] According to one aspect of the present disclosure, a sample construction method is provided, comprising: determining an initial reasoning chain for answering a question text based on an associated image and question text; scrambling the initial reasoning chain to obtain a scrambled reasoning chain that satisfies a predetermined abnormality condition; determining a target reasoning chain based on the initial reasoning chain and the scrambled reasoning chain based on a large model, the target reasoning chain comprising a reference text and the scrambled reasoning chain, the reference text being used to describe abnormal information in the scrambled reasoning chain; and constructing a training sample based on the image, the question text, and the target reasoning chain.

[0005] According to another aspect of the present disclosure, a method for training a multimodal large model is provided, comprising: obtaining training samples, the training samples comprising images, question texts, and target reasoning chains; and training the multimodal large model using the training samples; wherein the training samples are obtained using the above-mentioned sample construction method.

[0006] According to another aspect of the present disclosure, a sample construction device is provided, comprising: an initial reasoning chain determination module, a scrambling module, a target reasoning chain determination module, and a construction module. The initial reasoning chain determination module is used to determine an initial reasoning chain for answering a question text based on an associated image and question text. The scrambling module is used to scramble the initial reasoning chain to obtain a scrambled reasoning chain that meets a predetermined abnormality condition. The target reasoning chain determination module is used to determine a target reasoning chain based on a large model, the initial reasoning chain, and the scrambled reasoning chain. The target reasoning chain includes a reference text and a scrambled reasoning chain. The reference text is used to describe abnormal information in the scrambled reasoning chain. The construction module is used to construct a training sample based on the image, the question text, and the target reasoning chain.

[0007] According to another aspect of the present disclosure, a training apparatus for a large multimodal model is provided, comprising an acquisition module and a training module. The acquisition module is configured to acquire training samples, which include images, question text, and target reasoning chains. The training module is configured to train the large multimodal model using the training samples, wherein the training samples are constructed using the sample construction module.

[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.

[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method provided by the present disclosure.

[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the method provided in the present disclosure when executed by a processor.

[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0013] Figure 1 Schematic diagram of an application scenario of a sample construction method, a multimodal large model training method, and an apparatus according to an embodiment of the present disclosure;

[0014] Figure 2 is a schematic flow chart of a sample construction method according to an embodiment of the present disclosure;

[0015] Figure 3 is a schematic diagram of a sample construction method according to an embodiment of the present disclosure;

[0016] Figure 4 is a schematic flow chart of a method for training a multimodal large model according to an embodiment of the present disclosure;

[0017] Figure 5 is a schematic structural block diagram of a sample construction device according to an embodiment of the present disclosure;

[0018] Figure 6is a schematic structural block diagram of a training device for a multimodal large model according to an embodiment of the present disclosure; and

[0019] Figure 7 It is a structural block diagram of an electronic device used to implement the sample construction method and the multimodal large model training method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0021] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0022] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0023] Training a large multimodal model requires the use of multimodal training samples. One type of training sample includes images, question text, and answer text. The answer text is relatively simple, for example, the answer text only contains the answer, or also includes some simple reasoning chains. This simple reasoning chain generally only shows the direct path from the question to the answer.

[0024] Because training samples lack information for error recognition, reflection, and error correction, using such training samples to train large multimodal models results in them lacking the ability to perform complex reasoning. For example, large multimodal models lack the ability to recognize errors and are less sensitive to potential contradictions or errors in reasoning. Another example is that large multimodal models lack reflection and correction capabilities, lacking mechanisms to reflect on and correct errors when faced with complex tasks. Furthermore, large multimodal models have limited logical reasoning capabilities, making them inadequate for tasks requiring in-depth reasoning, such as solving science problems or complex multi-step question-answering tasks.

[0025] In summary, the lack of training samples results in poor reasoning capabilities of large multimodal models in areas such as error recognition, reflective correction, and multi-path reasoning. These shortcomings are particularly pronounced in domain-specific tasks (such as educational assessment and scientific problem-solving), limiting their potential for application in high-level reasoning tasks.

[0026] The disclosed embodiment aims to provide a sample construction method, which can construct a training sample including a target reasoning chain, and the target reasoning chain includes a scrambled reasoning chain of predetermined abnormal conditions, and the scrambled reasoning chain may, for example, contain some errors, information that is inconsistent with the image or text, or incomplete information. At the same time, the target reasoning chain also includes a reference text, and the reference text may, for example, explain the abnormal information in the scrambled reasoning chain, so that the target thinking chain in the training sample is a long thinking chain with the characteristics of thinking, error recognition, reflection, and correction. The deeper reasoning steps in the long thinking chain can improve the comprehensive reasoning ability of the multimodal large model. Therefore, using the training sample to train the multimodal large model can improve the reasoning ability of the multimodal large model in multimodal understanding, such as improving the error recognition, reflection, correction and other capabilities of the multimodal large model.

[0027] The technical solutions provided by the present disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0028] Figure 1 It is a schematic diagram of an application scenario of the sample construction method, multimodal large model training method and device according to the embodiment of the present disclosure.

[0029] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.

[0030] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0031] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with display screens and support web browsing, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers, etc.

[0032] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests and feedback the processing results to the terminal device.

[0033] The system architecture 100 may include a database 106 , which may store pre-collected image and question text data pairs, and may also store constructed training samples.

[0034] For example, during the sample generation phase, users can input image and question text data pairs through terminal devices 101, 102, and 103, or retrieve pre-stored image and question text data pairs from database 106. A large model, such as a large language model or a multimodal large model, can be deployed in server 105. The large model can be used to process images and question text to generate a target reasoning chain, and then construct training samples. The training samples can also be stored in database 106.

[0035] For example, during the model training phase, the server 105 may obtain pre-built training samples from the database 106 and then train the multimodal large model to be trained until the multimodal large model converges. After obtaining the trained multimodal large model, the multimodal large model may be deployed on the server 105.

[0036] For example, during the model application phase, server 105 can obtain the image to be processed and the question text through terminal devices 101, 102, and 103, and then use the pre-deployed, trained multimodal large model to process the obtained image and question text to generate a response text. Server 105 can also return the response text to terminal devices 101, 102, and 103 for display to the user.

[0037] It should be noted that the sample construction method and the multimodal large model training method provided in the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the sample construction device and the multimodal large model training device provided in the embodiments of the present disclosure can generally be set in the server 105.

[0038] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0039] Figure 2 is a schematic flowchart of a sample construction method according to an embodiment of the present disclosure.

[0040] like Figure 2 As shown, the sample construction method 200 may include operations S210 to S240.

[0041] In operation S210 , an initial reasoning chain for answering the question text is determined based on the associated image and question text.

[0042] For example, the question text may ask a question about the image, such as "What does the picture tell you?" For example, the initial reasoning chain may include multiple reasoning steps, and the initial reasoning chain may be determined based on a large model. This embodiment does not limit the process of determining the initial reasoning chain.

[0043] In operation S220 , the initial inference chain is scrambled to obtain a scrambled inference chain that meets a predetermined abnormality condition.

[0044] For example, scrambling processing is used to modify the content of the initial reasoning chain. If the scrambled reasoning chain meets the predetermined abnormal condition, it can mean that the scrambled initial reasoning chain contains some abnormal content, such as logical errors, common sense errors, inconsistency with the image content, inconsistency with the question text, and other abnormal content.

[0045] In operation S230 , based on the large model, a target reasoning chain is determined according to the initial reasoning chain and the scrambled reasoning chain. The target reasoning chain includes a reference text and the scrambled reasoning chain. The reference text is used to describe abnormal information in the scrambled reasoning chain.

[0046] For example, the initial reasoning chain and the scrambled reasoning chain can be compared to determine the location of the error in the scrambled reasoning chain, the cause of the error, the answer to correct the error, and other information. The reference text can explain this information.

[0047] For example, both the initial reasoning chain and the scrambled reasoning chain can be input into a large model, and prompt information can be used to guide the large model to determine abnormal information in the scrambled reasoning chain and output a target reasoning chain, or to output a reference text and then combine the reference text with the scrambled reasoning chain to form a target reasoning chain. The large model can be, for example, a large language model.

[0048] For another example, the image, question text, initial reasoning chain, and scrambled reasoning chain can all be input into the large model. Another prompt can be used to guide the large model to identify abnormal information in the scrambled reasoning chain and output a target reasoning chain. Alternatively, the large model can output reference text and then combine the reference text with the scrambled reasoning chain to form a target reasoning chain. For example, the large model can be a multimodal large model.

[0049] In operation S240 , a training sample is constructed based on the image, the question text, and the target reasoning chain.

[0050] For example, an image, question text, and target reasoning chain can be combined into a training sample. Alternatively, the target reasoning chain can be post-processed and then the image, question text, and target reasoning chain after post-processing can be combined into a training sample.

[0051] According to the technical solution provided by the embodiment of the present disclosure, the training sample obtained by the sample construction method includes a target reasoning chain, and the target reasoning chain includes an scrambled reasoning chain of predetermined abnormal conditions. The scrambled reasoning chain may, for example, contain some errors, information that is inconsistent with the image or text, or incomplete information. At the same time, the target reasoning chain also includes a reference text. The reference text may, for example, explain the abnormal information in the scrambled reasoning chain. In this way, the target thinking chain in the training sample is a long thinking chain with the characteristics of thinking, error recognition, reflection, and correction. The deeper reasoning steps in the long thinking chain can improve the comprehensive reasoning ability of the multimodal large model. Therefore, using the training sample to train the multimodal large model can improve the reasoning ability of the multimodal large model in multimodal understanding, such as improving the error recognition, reflection, correction and other capabilities of the multimodal large model.

[0052] Next, the process of determining the initial reasoning chain is described.

[0053] In one example, instead of classifying the question text, general prompt information can be used to guide the multimodal model to output an initial reasoning chain. For example, the image, question text, and general prompt information can be used to generate an initial reasoning chain based on the multimodal model.

[0054] In another example, the question text can be classified to obtain its category. A target prompt matching the category is then determined from multiple candidate prompts. Based on the image, question text, and target prompt, an initial reasoning chain is generated based on the multimodal macro model. For example, a pre-trained classification model can be used to classify the question text and image into categories such as reasoning and problem-solving. The correspondence between categories and prompts can also be pre-configured, with different prompts corresponding to question texts of different categories. The image, question text, and target prompt corresponding to the question text category are input into the multimodal macro model, which then generates an initial reasoning chain. The prompt guides the multimodal macro model to consider different perspectives, such as perception, cognition, and reasoning, and outputs the reasoning steps it deems reasonable, constructing a linear reasoning path from question to answer, and then outputting the initial reasoning chain. This embodiment, by performing classification and selecting target prompts matching the category, can specifically guide the multimodal macro model to output the initial reasoning chain. By selecting more suitable target prompts, the reasoning accuracy of the multimodal macro model can be improved, thereby improving the quality of the initial reasoning chain.

[0055] In another example, multiple initial reasoning chains can be generated for the same data pair consisting of an image and question text, and then some high-quality initial reasoning chains can be screened from the multiple initial reasoning chains. In practical applications, multiple initial reasoning chains can be generated in a variety of ways. For example, a single multimodal large model can be used and required to generate multiple initial reasoning chains with different reasoning logic using different methods. For another example, multiple multimodal large models can be used to process the image, question text, and target prompt information respectively, thereby obtaining multiple initial reasoning chains output by the multiple multimodal large models. After generating multiple initial reasoning chains, some high-quality initial reasoning chains can be screened from the multiple initial reasoning chains, and then only the screened initial reasoning chains are scrambled to obtain the scrambled reasoning chains. In this embodiment, multiple initial reasoning chains are first generated, then screened, and then only the high-quality initial reasoning chains are scrambled. This can improve the quality of the scrambled reasoning chains by selecting high-quality initial reasoning chains.

[0056] The above describes the process of generating an initial reasoning chain.

[0057] Next, the process of screening high-quality initial reasoning chains from multiple initial reasoning chains is described.

[0058] In the disclosed embodiment, multiple initial reasoning chains can be pre-generated. Then, for each of the multiple initial reasoning chains, the initial reasoning chain is evaluated based on at least one predetermined dimension, based on the image, the question text, and the content of the initial reasoning chain, to obtain an evaluation result representing the quality of the initial reasoning chain. Then, based on the evaluation results of each of the multiple initial reasoning chains, a portion of the initial reasoning chains is filtered from the multiple initial reasoning chains to obtain filtered initial reasoning chains.

[0059] For example, the evaluation dimensions may include content consistency, logical rationality, depth of the thinking chain, etc. Content consistency indicates a degree of consistency between the initial reasoning chain and the content of the image and question text. For example, the image, question text, initial reasoning chain, and prompt information for evaluation may be input into a multimodal large model. The prompt information may be used to guide the multimodal large model to evaluate the initial reasoning chain from the above-mentioned dimensions and output an evaluation result. The evaluation result may include, for example, evaluation values ​​of each dimension, weighted average calculations such as those performed on the evaluation values ​​of each dimension, and the budget result may be used as the evaluation result. After obtaining the evaluation results of each initial reasoning chain, several initial reasoning chains with the highest evaluation results may be retained, or initial reasoning chains with calculated operation results higher than a threshold may be retained, while other initial reasoning chains are filtered out and no longer processed.

[0060] This embodiment evaluates the initial reasoning chain from several predetermined dimensions based on the image, question text and the content of the initial reasoning chain, and screens the initial reasoning chain based on the evaluation results, so as to accurately screen out high-quality initial reasoning chains.

[0061] In one example, the evaluation results include a first evaluation result. During the process of determining the first evaluation result, the initial reasoning chain can be evaluated based on at least one predetermined dimension based on the image, the question text, and the initial reasoning chain, respectively, based on multiple first evaluation models, to obtain evaluation sub-results determined by each of the multiple first evaluation models. The first evaluation result of the initial reasoning chain is then determined based on the evaluation sub-results of each of the multiple first evaluation models.

[0062] For example, the first evaluation model can be a multimodal model, and the process of evaluating the initial reasoning chain by a single first evaluation model can refer to the above. For example, the image, question text, initial reasoning chain and prompt information for guiding the multimodal large model to perform evaluation can be input into the first evaluation model to obtain an evaluation sub-result. In this embodiment, for the same initial reasoning chain, multiple first evaluation models can be used for evaluation, and each first evaluation model can output an evaluation sub-result. Then, the first evaluation result is determined based on the multiple evaluation sub-results output by the multiple first evaluation models. In this way, multiple first evaluation models can more accurately determine the first evaluation result through voting, weighting, cross-validation or other processing.

[0063] In another example, the evaluation result includes a second evaluation result. In the process of determining the second evaluation result, the image can be identified to obtain the object in the image, and then the initial reasoning chain is evaluated based on the second evaluation model according to the object text, question text and initial reasoning chain associated with the object to obtain the second evaluation result; the evaluation result includes the second evaluation result.

[0064] For example, high-precision visual operators can be used to recognize images. Visual operators can include optical character recognition models, object detection models, etc., and objects in the image can be extracted through visual operators. The object text associated with the object can be used to describe the object's category, characteristics, etc. For example, the second evaluation model can be a large language model. For example, the object text, question text, initial reasoning chain, and prompt information used to guide the large language model in evaluation can be input into the second evaluation model. The second evaluation model will evaluate the rationality of the initial reasoning path, the consistency of the initial reasoning chain with the visual information in the image, the rationality of the thinking logic, etc. If the second evaluation result indicates that the initial reasoning chain is of poor quality, the initial reasoning chain can be filtered out. In this embodiment, the object in the image is first extracted, then the object is converted into object text, and then the evaluation is performed based on the object text, question text, and initial reasoning chain. In this way, a large language model can be used as the second evaluation model. Compared with a multimodal large model, the large language model has stronger language logic and computing capabilities, resulting in more accurate evaluation.

[0065] It should be noted that, in actual applications, either the first evaluation model or the second evaluation model can be used for evaluation, or both can be used for evaluation. For example, the first evaluation model can be used for evaluation and to screen some initial reasoning chains, and then the second evaluation model can be used to evaluate the remaining initial reasoning chains again, thereby screening some of the reasoning chains again. Alternatively, for the same initial reasoning chain, the first evaluation model and the second evaluation model can be used to determine the first evaluation result and the second evaluation result, respectively, and then the first evaluation result and the second evaluation result can be weighted or subjected to other operations to determine the evaluation result of the initial reasoning chain.

[0066] The above describes the process of screening high-quality initial reasoning chains from multiple initial reasoning chains.

[0067] After obtaining the initial inference chain, the initial inference chain may be scrambled. Next, the scrambling process will be described.

[0068] In one example, entities in an initial reasoning chain can be replaced with other entities. For example, a text in the initial reasoning chain may be "There is a cat in the picture," but after replacement, the text becomes "There is a dog in the picture." Entity replacement can be achieved in a variety of ways. For example, the reasoning chain can be processed by named entity recognition or keyword extraction to determine the entity to be replaced. It is also possible to pre-determine the correspondence between entities and replace the entity to be replaced based on this correspondence. For another example, a multimodal model can be used for entity replacement. For example, an image, question text, an initial reasoning chain, and prompt information used to guide the multimodal model in entity replacement are input into the multimodal model, which then outputs a scrambled reasoning chain. This embodiment misinterprets or fuzzily describes the text in the reasoning chain, thereby simulating an illusion at the perceptual level, causing the scrambled reasoning chain to meet abnormal conditions, such as causing the scrambled reasoning chain to be erroneous or inconsistent with the image or question text.

[0069] In another example, contextual text that satisfies a logical relationship in an initial inference chain can be modified so that the modified contextual text no longer satisfies the logical relationship. For example, a section of text in the initial inference chain might read "Because two straight lines are parallel, they do not intersect"; after modification, the text might read "Because two straight lines are perpendicular, they do not intersect" or "Because two straight lines are not parallel, they intersect." Modifying logical relationships can be achieved in a variety of ways. For example, specific keywords in the inference chain can be modified to other words, such as by adding or removing negations, or by replacing keywords. For another example, a multimodal model can be used. An image, question text, an initial inference chain, and prompt information used to guide the multimodal model in generating erroneous logic can be input into the multimodal model, which then outputs a scrambled inference chain. This embodiment embeds assumptions or misrepresentations based on biased logic into the inference chain, thereby creating a hallucination simulation at the cognitive level, causing the scrambled inference chain to meet abnormal conditions, such as causing the scrambled inference chain to be erroneous or inconsistent with the image or question text.

[0070] In another example, some of the reasoning steps in the initial reasoning chain can be deleted. For example, if the initial reasoning chain includes several reasoning steps, some of the steps can be randomly deleted from the reasoning steps. For example, if the initial reasoning chain includes steps 1 to 5, steps 1 and 2 can be deleted, or step 3 can be deleted.

[0071] It should be noted that the above provides multiple means for scrambling the initial reasoning chain. In practical applications, the various means can be used in combination or individually. This embodiment does not limit the combination of the various scrambling means.

[0072] The above describes the scrambling process of the initial reasoning chain. Through the scrambling process, the thinking bias in the initial reasoning chain can be disturbed, thereby simulating the potential bias in human thinking.

[0073] It should be noted that the purpose of the scrambling process is to change the initial reasoning chain so that the scrambled reasoning chain meets the abnormal condition.

[0074] In one example, the abnormal condition includes that the scrambled inference chain is inconsistent with the image, for example, the image is a cat, but the inference chain describes that the image is a dog.

[0075] In another example, an abnormal condition includes the scrambled inference chain being inconsistent with the question text. For example, the question text asks for clothing recommendations, but the inference chain answers with food recommendations.

[0076] In another example, an abnormal condition includes: the scrambled reasoning chain contains erroneous content, which may include common sense errors, logical errors, etc. For example, the initial reasoning chain states "Because tomorrow's temperature is 35 degrees Celsius, it is recommended to wear light clothing," while the scrambled reasoning chain states "Because tomorrow's temperature is -5 degrees Celsius, it is recommended to wear light clothing."

[0077] In another example, the abnormal condition includes: the scrambled reasoning chain lacks some reasoning steps relative to the initial reasoning chain. For example, the initial reasoning chain includes steps 1 to 5, while the scrambled reasoning chain only includes steps 3 to 5.

[0078] It should be noted that this embodiment does not limit the specific means of scrambling processing, as long as the scrambled reasoning chain can meet the abnormal conditions.

[0079] According to another embodiment of the present disclosure, after obtaining the scrambled reasoning chain, a target reasoning chain can be determined, which contains a reference text. In the actual processing process, the target reasoning chain can be generated based on multimodal guidance. For example, the image, the question text, the initial reasoning chain before scrambling, the reasoning chain after scrambling and related prompt information are input into the multimodal large model, and the multimodal large model outputs the target reasoning chain, wherein the prompt information can provide clear reflection and thinking methods, and can also provide some samples to guide the multimodal large model to identify errors through logical deduction and generate reference text. In this embodiment, the initial reasoning chain before scrambling is equivalent to the correct answer. By introducing the correct answer and the reflection and error correction mechanism, the multimodal large model is guided to gradually return from the erroneous thinking path to the correct reasoning direction.

[0080] It should be noted that the multimodal large model does not correct the abnormal content in the scrambled reasoning chain, that is, it does not modify the wrong content to be correct, but describes the wrong information. In one example, the reference information can describe the location of the abnormal step in the scrambled reasoning chain, such as describing that the third reasoning step is wrong. In another example, the reference information can describe the abnormal reason for the abnormal step in the scrambled reasoning chain, for example, "Because the temperature is -5 degrees Celsius, the temperature is low, and light clothing has poor warmth retention, so the clothing recommendation is unreasonable." In another example, the reference information can describe the correction result for the abnormal step in the scrambled reasoning chain, such as "the calculation process 1+1=3, the correct calculation result should be 1+1=2."

[0081] In this embodiment, the initial reasoning chain before scrambling only provides forward reasoning steps and lacks features such as reflection and error recognition. However, this embodiment uses reference text to describe the errors after scrambling. This simulates the human ability to reconstruct thoughts after interference, giving the target reasoning chain containing the reference text a reflective nature. Furthermore, based on practical findings, reasoning chains with reflective properties can slow down the thinking of large models, thereby enhancing their reasoning capabilities.

[0082] According to another embodiment of the present disclosure, after obtaining the target reasoning chain, the image, question text and target reasoning chain can be directly combined into a training sample.

[0083] In another embodiment of the present disclosure, after obtaining the target inference chain, the target inference chain can be post-processed to obtain an adjusted target inference chain. The image, question text, and the adjusted target inference chain are then combined into a training sample. During the post-processing of the target inference chain, the target inference chain can be adjusted based on the large language model according to the question text and the target inference chain to obtain an adjusted inference chain.

[0084] In one example, the adjustment process for the target reasoning chain is used to adjust the number of reasoning steps in the target reasoning chain. For example, a large language model can be used. The input to the large language model includes the question text, the target reasoning chain, and prompt information for post-processing. The prompt information may include examples, including the question, the reasoning path (the reasoning path can provide the reasoning result), and the adjusted reasoning chain. The large language model is guided by the examples in the prompt information, so that the adjusted reasoning chain output by the large language model is similar to the adjusted reasoning chain in the example. In addition, the prompt information can gradually optimize the reasoning chain from aspects such as perception (such as extraction of basic features), cognition (such as modeling relationships between features), and reasoning (such as derivation of causal chains), thereby refining the target reasoning chain and controlling the depth of the reasoning path. In this way, the adjusted reasoning chain can both cover the core reasoning process of complex problems and maintain a reasonable length, avoiding excessive length or one-sidedness.

[0085] In another example, the adjustment process for the target reasoning chain is used to insert a description text that is semantically consistent with a predetermined text into the target reasoning chain to adjust the expression of the target reasoning chain.

[0086] For example, by using anthropomorphic thinking patterns, we can generate multiple reasoning paths that are closer to human thinking. For example, we can input the question text, the target reasoning chain, and prompt information into the large language model, which will then output an adjusted reasoning chain. The prompt information can guide the large language model to simulate the common human "hypothesis-verification-correction" model during the chain of thought. The prompt information can also provide examples. For example, an example in the prompt information is as follows:

[0087] Initial observation and hypothesis: "Wait, the flower in the picture doesn't seem to be a camellia. The shape of the flower above is very similar to a camellia."; Further reasoning based on knowledge: "Let me think about it, the characteristics of a camellia include feature 1 and feature 2, and the flower in the picture is very close in these features."; Discover contradictions and make corrections: "However, the patterns on the leaves of the flower in the picture don't seem to be consistent, which may indicate that it is not a camellia."; Final verification and conclusion: "Combined with observation and feature comparison, the flower in the picture may be an azalea because its leaf patterns are more consistent with the characteristics."

[0088] In this embodiment, some descriptive texts similar to "wait", "let me think about it", "seems inconsistent", etc. can be inserted into the reasoning chain. These descriptive texts are consistent with the reflection and correction mechanism of human logic, thereby optimizing the naturalness and logic of multiple reasoning paths and enhancing the intuitiveness and persuasiveness of the reasoning chain.

[0089] Figure 3 is a schematic diagram of a sample construction method according to an embodiment of the present disclosure.

[0090] In this embodiment, multiple initial reasoning chains 331, 332, and 333 can be generated based on the image 310 and the question text 320 through multiple generation models M1.1, M1.2, and M1.3. Taking the generation model M1.1 as an example, the generation model M1.1 can be a multimodal large model. The input of the generation model M1.1 includes the image 310, the question text 320, and the first prompt information prompt1. The first prompt information prompt1 can be determined based on the category of the question text 320, or it can be a general prompt information unrelated to the category. The output of the multimodal large model is the initial reasoning chain.

[0091] Next, the multiple initial reasoning chains 331 , 332 , and 333 may be evaluated using the evaluation model M2 . The evaluation model M2 , for example, includes multiple first evaluation models M2 . 1 and second evaluation models M2 . 2 .

[0092] The first evaluation model M2.1 can be a multimodal large model. The input of each first evaluation model M2.1 includes an image 310, a question text 320, an initial reasoning chain and a second prompt information prompt2. The evaluation sub-result can be determined based on the output of the first evaluation model M2.1. The multiple evaluation substructures of multiple first evaluation models M2.1 can calculate the first evaluation result, and then preliminary screening can be performed based on the first evaluation result to filter out some initial reasoning chains, such as filtering out the initial reasoning chain 332.

[0093] For the remaining initial reasoning chains, visual operators can be used to detect objects in image 310. These chains are then evaluated again using a second evaluation model M2.2, which can be a large language model. The inputs to the second evaluation model M2.2 include the object text associated with the object, question text 320, the initial reasoning chains, and the third prompt message prompt3. The second evaluation model M2.2 evaluates the rationality of the initial reasoning chains and screens for high-quality initial reasoning chains. For example, after screening, initial reasoning chain 331 is retained.

[0094] The initial reasoning chain 331 after screening can be scrambled to obtain a scrambled reasoning chain 340. During the scrambling process, hallucination simulation can be performed at the perceptual level, such as by misreading or vaguely describing the text in the reasoning chain. Hallucination simulation can also be performed at the cognitive level, for example, by embedding assumptions or misunderstandings based on biased logic in the thinking chain. It is also possible to randomly delete some reasoning steps in the reasoning chain based on rules. The scrambling process can be implemented based on a scrambling model, such as a large multimodal model. The input of the scrambling model includes the image 310, the question text 320, the initial reasoning chain, and the fourth prompt information prompt4. The output of the scrambling model is the scrambled reasoning chain 340. The scrambling process ensures that the scrambled reasoning chain 340 meets the predetermined abnormality condition.

[0095] After obtaining the scrambled reasoning chain 340, a target reasoning chain 360 can be determined based on the large model, the initial reasoning chain, and the scrambled reasoning chain 340. The target reasoning chain 360 includes reference text 350 and the scrambled reasoning chain 340. Reference text 350 is used to describe the abnormal information in the scrambled reasoning chain 340. The large model used in this process can be a multimodal large model. The input of the multimodal large model includes the image 310, the question text 320, the pre-scrambled reasoning chain, the scrambled reasoning chain 340, and the fifth prompt message prompt5. The multimodal large model generates the target reasoning chain 360. The target reasoning chain 360 is equivalent to inserting reference information into the scrambled reasoning chain 340. The reference information is used to describe the abnormal information in the scrambled reasoning chain 340, such as the location of the abnormal step in the scrambled reasoning chain 340, the cause of the abnormality, and the correction result.

[0096] After obtaining the target reasoning chain 360, post-processing can be performed on the target reasoning chain 360. Post-processing mainly involves adjusting the number of reasoning steps in the target reasoning chain 360 or inserting descriptive text that is semantically consistent with predetermined text into the target reasoning chain 360 to adjust the expression of the target reasoning chain 360.

[0097] For example, the post-processing process can be implemented based on a large language model. The input of the large language model includes the question text 320, the target reasoning chain 360, and the sixth prompt information prompt 6. The sixth prompt information prompt 6 may include an example. It is understood that the sixth prompt information prompt 6 used to guide the large language model to adjust the reasoning depth and the sixth prompt information prompt 6 used to guide the large language model to adjust the expression method can be different.

[0098] By post-processing the target reasoning chain 360 , an adjusted reasoning chain 370 may be obtained. Next, the image 310 , the question text 320 , and the adjusted reasoning chain 370 may be combined into a training sample 380 .

[0099] It should be noted that the above-mentioned multiple processes use multiple models, such as multimodal guidance, large language models, etc., wherein the large models used in different processes can be the same or different. For example, the first evaluation model and the large model used for scrambling can use the same multimodal large model, or different multimodal large models. In addition, the prompt information used by each large model is flexibly configured according to actual needs. This embodiment does not limit the format, content, and expression of the prompt information, and it is sufficient to guide the large model to achieve the corresponding function.

[0100] By adopting the sample construction method provided in this embodiment, a training sample 380 for training a multimodal large model can be obtained, and the training sample 380 includes long thought chain data with characteristics such as thinking, error recognition, reflection, and correction.

[0101] Figure 4 It is a schematic flowchart of a training method for a multimodal large model according to an embodiment of the present disclosure.

[0102] like Figure 4 As shown, the training method 400 of the multimodal large model may include operations S410 to S420.

[0103] In operation S410, a training sample is obtained, where the training sample includes an image, a question text, and a target reasoning chain. For example, the training sample is obtained using the above-mentioned sample construction method.

[0104] In operation S420 , a multimodal large model is trained using the training samples.

[0105] This embodiment uses training samples of long thinking chains with characteristics such as thinking, error recognition, reflection, and correction to train a large multimodal model, which can improve the reasoning ability of the large multimodal model, thereby understanding and responding to complex reasoning scenarios. It can also enhance the slow thinking ability of the large multimodal model. For example, the large multimodal model can simulate the "slow thinking" process in reasoning, gradually generate hypotheses, verify hypotheses, identify errors and correct them, thereby improving its multi-step reasoning performance in complex problems. In addition, compared to manually annotating reasoning steps, since the training samples are automatically generated, the cost of sample generation is reduced, thereby reducing the model cost.

[0106] Figure 5 is a schematic structural block diagram of a sample construction device according to an embodiment of the present disclosure.

[0107] like Figure 5 As shown, the sample construction device 500 may include an initial reasoning chain determination module 510 , a scrambling module 520 , a target reasoning chain determination module 530 , and a construction module 540 .

[0108] The initial reasoning chain determination module 510 is used to determine an initial reasoning chain for answering the question text based on the associated image and question text.

[0109] The scrambling module 520 is configured to perform scrambling processing on the initial inference chain to obtain a scrambled inference chain that meets a predetermined abnormality condition.

[0110] The target reasoning chain determination module 530 is used to determine the target reasoning chain based on the large model, according to the initial reasoning chain and the scrambled reasoning chain. The target reasoning chain includes a reference text and the scrambled reasoning chain. The reference text is used to describe abnormal information in the scrambled reasoning chain.

[0111] The construction module 540 is used to construct training samples based on images, question texts and target reasoning chains.

[0112] According to another embodiment of the present disclosure, a scrambling module includes at least one of the following: a first scrambling submodule, a second scrambling submodule, and a third scrambling submodule. The first scrambling submodule is configured to replace entities in an initial inference chain with other entities. The second scrambling submodule is configured to modify context text that satisfies a logical relationship in the initial inference chain so that the modified context text no longer satisfies the logical relationship. The third scrambling submodule is configured to delete some inference steps in the initial inference chain.

[0113] According to another embodiment of the present disclosure, the abnormal condition includes at least one of the following: the scrambled reasoning chain is inconsistent with the image; the scrambled reasoning chain is inconsistent with the question text; the scrambled reasoning chain contains erroneous content; and the scrambled reasoning chain lacks some reasoning steps relative to the initial reasoning chain.

[0114] According to another embodiment of the present disclosure, the reference text is used to describe at least one of the following information: the location of the abnormal step in the scrambled reasoning chain; the abnormal cause of the abnormal step in the scrambled reasoning chain; and the correction result for the abnormal step in the scrambled reasoning chain.

[0115] According to another embodiment of the present disclosure, the construction module includes: an adjustment submodule and a construction submodule. The adjustment submodule is used to adjust the target reasoning chain based on the large language model according to the question text and the target reasoning chain to obtain an adjusted reasoning chain. The construction submodule is used to construct a training sample for training the multimodal large model according to the image, the question text and the adjusted reasoning chain. The adjustment process is used to adjust at least one of the following: the number of reasoning steps in the target reasoning chain, and inserting a descriptive text that is consistent with the predetermined text semantics in the target reasoning chain to adjust the expression of the target reasoning chain.

[0116] According to another embodiment of the present disclosure, there are multiple initial reasoning chains; the scrambling module includes: an evaluation submodule, a screening submodule, and a scrambling submodule. The evaluation submodule is used to evaluate each of the multiple initial reasoning chains from at least one predetermined dimension based on the image, the question text, and the content of the initial reasoning chain, and obtain an evaluation result that characterizes the quality of the initial reasoning chain. The screening submodule is used to screen some of the initial reasoning chains from the multiple initial reasoning chains based on the respective evaluation results of the multiple initial reasoning chains to obtain screened initial reasoning chains. The scrambling submodule is used to scramble the screened initial reasoning chains to obtain scrambled reasoning chains.

[0117] According to another embodiment of the present disclosure, the evaluation submodule includes an evaluation unit and a first evaluation result determination unit. The evaluation unit is configured to evaluate the initial reasoning chain from at least one predetermined dimension based on the image, the question text, and the initial reasoning chain based on multiple first evaluation models, thereby obtaining an evaluation sub-result determined by each of the multiple first evaluation models. The first evaluation result determination unit is configured to determine a first evaluation result of the initial reasoning chain based on the evaluation sub-results of each of the multiple first evaluation models; the evaluation results include the first evaluation result.

[0118] According to another embodiment of the present disclosure, the evaluation submodule includes: a recognition unit and a second evaluation result determination unit. The recognition unit is configured to recognize an image and obtain an object in the image. The second evaluation result determination unit is configured to evaluate the initial reasoning chain based on a second evaluation model, based on object text associated with the object, question text, and the initial reasoning chain, to obtain a second evaluation result. The evaluation result includes the second evaluation result.

[0119] According to another embodiment of the present disclosure, the initial reasoning chain determination module includes a classification submodule, a prompt determination submodule, and a generation submodule. The classification submodule is used to classify the question text and determine the question text category. The prompt determination submodule is used to determine target prompt information that matches the category from multiple candidate prompt information. The generation submodule is used to generate an initial reasoning chain based on the image, question text, and target prompt information based on the multimodal large model.

[0120] Figure 6 It is a schematic structural block diagram of a training device for a multimodal large model according to an embodiment of the present disclosure.

[0121] like Figure 6 As shown, the multimodal large model training device 600 may include an acquisition module 610 and a training module 620.

[0122] The acquisition module 610 is used to acquire training samples. The training samples include images, question texts, and target reasoning chains. The training samples are obtained using the above-mentioned sample construction device.

[0123] The training module 620 is used to train a large multimodal model using training samples.

[0124] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, comprising at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute at least one of the above-mentioned sample construction method and multimodal large model training method.

[0125] According to an embodiment of the present disclosure, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute at least one of the above-mentioned sample construction method and multimodal large model training method.

[0126] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program, which implements at least one of the above-mentioned sample construction method and multimodal large model training method when executed by a processor.

[0127] Figure 7 1 is a block diagram of an electronic device for implementing the sample construction method and the training method of the multimodal large model of the embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0128] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. Computing unit 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0129] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0130] The computing unit 701 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as at least one of the aforementioned sample construction method and the multimodal large model training method. For example, in some embodiments, at least one of the aforementioned sample construction method and the multimodal large model training method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of at least one of the aforementioned sample construction method and the multimodal large model training method can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured in any other appropriate manner (for example, by means of firmware) to execute at least one of the above-mentioned sample construction method and multimodal large model training method.

[0131] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0132] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0133] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0134] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0135] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0136] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0137] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0138] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A sample construction method, comprising: determining an initial reasoning chain for answering the question text based on the associated image and question text; Scrambling the initial reasoning chain to obtain a scrambled reasoning chain that meets a predetermined abnormality condition; Determining a target reasoning chain based on the large model and the initial reasoning chain and the scrambled reasoning chain, wherein the target reasoning chain includes a reference text and the scrambled reasoning chain, and the reference text is used to describe abnormal information in the scrambled reasoning chain; as well as A training sample is constructed based on the image, the question text and the target reasoning chain, including: adjusting the target reasoning chain based on a large language model according to the question text and the target reasoning chain to obtain an adjusted reasoning chain; constructing a training sample for training a multimodal large model according to the image, the question text and the adjusted reasoning chain; wherein the adjustment processing is used to adjust at least one of the following: the number of reasoning steps in the target reasoning chain, and inserting a descriptive text that is semantically consistent with a predetermined text into the target reasoning chain to adjust the expression of the target reasoning chain.

2. The method according to claim 1, wherein The scrambling process includes at least one of the following: replacing entities in the initial chain of reasoning with other entities; Modifying the context text that satisfies the logical relationship in the initial reasoning chain so that the modified context text does not satisfy the logical relationship; as well as Delete some of the reasoning steps in the initial reasoning chain.

3. The method according to claim 1, wherein The abnormal condition includes at least one of the following: The scrambled chain of reasoning is inconsistent with the image; The scrambled reasoning chain is inconsistent with the question text; The scrambled inference chain contains erroneous content; as well as The scrambled reasoning chain lacks some reasoning steps compared to the initial reasoning chain.

4. The method according to claim 1, wherein The reference text is used to describe at least one of the following information: the location of the abnormal step in the scrambled chain of reasoning; The abnormal reason of the abnormal step in the scrambled reasoning chain; and Correction results for abnormal steps in the scrambled reasoning chain.

5. The method according to claim 1, wherein The number of the initial reasoning chains is multiple; and the scrambling of the initial reasoning chains to obtain scrambled reasoning chains that meet a predetermined abnormal condition comprises: For each of the plurality of initial reasoning chains, evaluating the initial reasoning chain from at least one predetermined dimension based on the image, the question text, and the content of the initial reasoning chain to obtain an evaluation result representing the quality of the initial reasoning chain; screening some initial reasoning chains from the plurality of initial reasoning chains according to the respective evaluation results of the plurality of initial reasoning chains to obtain screened initial reasoning chains; The scrambling process is performed on the screened initial inference chain to obtain the scrambled inference chain.

6. The method according to claim 5, wherein: The step of evaluating the initial reasoning chain from at least one predetermined dimension based on the image, the question text, and the content of the initial reasoning chain to obtain an evaluation result representing the quality of the initial reasoning chain includes: Based on the plurality of first evaluation models, respectively evaluating the initial reasoning chain from at least one predetermined dimension according to the image, the question text, and the initial reasoning chain, to obtain an evaluation sub-result determined by each of the plurality of first evaluation models; and A first evaluation result of the initial reasoning chain is determined according to the evaluation sub-results of each of the plurality of first evaluation models; the evaluation result includes the first evaluation result.

7. The method according to claim 5, wherein: The step of evaluating the initial reasoning chain from at least one predetermined dimension based on the image, the question text, and the content of the initial reasoning chain to obtain an evaluation result representing the quality of the initial reasoning chain includes: Recognizing the image to obtain an object in the image; and According to the object text associated with the object, the question text and the initial reasoning chain, the initial reasoning chain is evaluated based on a second evaluation model to obtain a second evaluation result; the evaluation result includes the second evaluation result.

8. The method according to claim 1, wherein Determining an initial reasoning chain for answering the question text based on the associated image and question text includes: Classifying the question text to obtain a category of the question text; Determining target prompt information matching the category from a plurality of candidate prompt information; and The initial reasoning chain is generated based on the multimodal large model according to the image, the question text and the target prompt information.

9. A method for training a large multimodal model, comprising: Obtaining a training sample, wherein the training sample includes an image, a question text, and a target reasoning chain; as well as Using the training samples, training a large multimodal model; The training samples are constructed using the method described in any one of claims 1 to 8.

10. A sample construction device, comprising: an initial reasoning chain determining module, configured to determine an initial reasoning chain for answering the question text based on the associated image and question text; a scrambling module, configured to scramble the initial inference chain to obtain a scrambled inference chain that satisfies a predetermined abnormality condition; a target reasoning chain determination module, configured to determine a target reasoning chain based on the large model and according to the initial reasoning chain and the scrambled reasoning chain, wherein the target reasoning chain includes a reference text and the scrambled reasoning chain, and the reference text is used to describe abnormal information in the scrambled reasoning chain; as well as A construction module, configured to construct a training sample based on the image, the question text, and the target reasoning chain; Wherein, the building blocks include: an adjustment submodule, configured to adjust the target reasoning chain based on the large language model according to the question text and the target reasoning chain to obtain an adjusted reasoning chain; and A construction submodule, configured to construct a training sample for training a multimodal large model based on the image, the question text, and the adjusted reasoning chain; The adjustment process is used to adjust at least one of the following: the number of reasoning steps in the target reasoning chain, and inserting a description text that is semantically consistent with a predetermined text into the target reasoning chain to adjust the expression of the target reasoning chain.

11. The device according to claim 10, wherein The scrambling module includes at least one of the following: a first scrambling submodule, configured to replace entities in the initial reasoning chain with other entities; a second scrambling submodule, configured to modify the context text satisfying the logical relationship in the initial reasoning chain so that the modified context text does not satisfy the logical relationship; as well as The third scrambling submodule is configured to delete some reasoning steps in the initial reasoning chain.

12. The device according to claim 10, wherein The abnormal condition includes at least one of the following: The scrambled chain of reasoning is inconsistent with the image; The scrambled reasoning chain is inconsistent with the question text; The scrambled inference chain contains erroneous content; as well as The scrambled reasoning chain lacks some reasoning steps compared to the initial reasoning chain.

13. The device according to claim 10, wherein The reference text is used to describe at least one of the following information: the location of the abnormal step in the scrambled chain of reasoning; The abnormal reason of the abnormal step in the scrambled reasoning chain; and Correction results for abnormal steps in the scrambled reasoning chain.

14. The device according to claim 10, wherein The number of the initial inference chains is multiple; the scrambling module includes: an evaluation submodule, configured to evaluate each of the plurality of initial reasoning chains from at least one predetermined dimension based on the image, the question text, and the content of the initial reasoning chain, to obtain an evaluation result representing the quality of the initial reasoning chain; a screening submodule, configured to screen some initial reasoning chains from the plurality of initial reasoning chains according to the respective evaluation results of the plurality of initial reasoning chains, to obtain screened initial reasoning chains; The scrambling submodule is configured to perform the scrambling process on the filtered initial inference chain to obtain the scrambled inference chain.

15. The device according to claim 14, wherein The evaluation submodule includes: an evaluation unit, configured to evaluate the initial reasoning chain from at least one predetermined dimension based on the image, the question text, and the initial reasoning chain based on a plurality of first evaluation models, and obtain an evaluation sub-result determined by each of the plurality of first evaluation models; and A first evaluation result determining unit is configured to determine a first evaluation result of the initial reasoning chain according to the evaluation sub-results of each of the plurality of first evaluation models; the evaluation result includes the first evaluation result.

16. The device according to claim 14, wherein The evaluation submodule includes: a recognition unit, configured to recognize the image and obtain an object in the image; and The second evaluation result determining unit is used to evaluate the initial reasoning chain based on the second evaluation model according to the object text associated with the object, the question text and the initial reasoning chain to obtain a second evaluation result; the evaluation result includes the second evaluation result.

17. The device according to claim 10, wherein The initial reasoning chain determination module includes: A classification submodule, configured to classify the question text to obtain a category of the question text; a prompt determination submodule, configured to determine target prompt information matching the category from a plurality of candidate prompt information; and A generation submodule is used to generate the initial reasoning chain based on the multimodal large model according to the image, the question text and the target prompt information.

18. A multimodal large model training device, comprising: An acquisition module, configured to acquire training samples, wherein the training samples include images, question texts, and target reasoning chains; as well as A training module, configured to train a large multimodal model using the training samples; Wherein, the training samples are constructed using the device described in any one of claims 10 to 17.

19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

21. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Image text double-end migration attack method and device and medium

    CN116523032A

  • Multi-modal thinking chain reasoning method and device based on knowledge distillation

    CN118014077A