A multi-modal joint-based large language model optimization method
By constructing multimodal data preprocessing rule templates and task graphs, the task processing flow of large language models is optimized, solving the problem of multimodal data fusion processing and improving the multimodal data processing capability of large language models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAVAL UNIV OF ENG PLA
- Filing Date
- 2026-03-31
- Publication Date
- 2026-05-29
AI Technical Summary
Existing large language models lack effective fusion processing capabilities when dealing with multimodal data, making it difficult to adapt to the future development needs of multimodal data.
By constructing multimodal data preprocessing rule templates, defining multimodal task type databases and graphs, establishing unit task candidate sets, and performing task reasoning optimization, the large language model is optimized to achieve multimodal data fusion processing.
It simplifies the fusion and processing of multimodal data, improves the compatibility and processing efficiency of large language models for multimodal data, and enables compatible processing of multimodal data by different types of large language models.
Smart Images

Figure CN122114181A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of large language model optimization technology, and particularly relates to a large language model optimization method based on multimodal collaboration. Background Technology
[0002] Large language models are automated processing technologies that integrate natural language processing, text recognition, and other technologies to process and analyze complex types of data. With the continuous development of machine recognition and analysis methods, large language model technology is no longer limited to the comprehensive processing of simple symbol entities such as characters and text. Instead, it has formed a joint processing of multiple types and modalities of data, including text, images, videos, and even audio-visual materials. Its ultimate goal is to use large language models to simulate the human brain's ability to jointly process and analyze multiple types of multimodal data, replacing manual labor in completing massive data analysis and processing tasks. It is an important direction for the future development of digital processing technology. However, currently, the commonly used large language models are mainly applied in specific fields for training and analysis of specific types of data, and there is a lack of specific solutions that can well adapt to the future development needs of multimodal data. Summary of the Invention
[0003] The purpose of this invention is to propose an optimized method based on multimodal data and its features, which is easy to implement and can help improve the multimodal data fusion processing capability and efficiency of various large language models, based on practical needs and the current basic principles and training methods of large language models.
[0004] To achieve the above objectives, the present invention adopts the following technical solution.
[0005] A method for optimizing large language models based on multimodal collaboration includes the following steps:
[0006] S1. Multimodal data preprocessing: construct feature rule templates for multimodal large language models, identify and extract feature structural elements; complete feature and element classification and corresponding semantic descriptions, and generate data structure labels;
[0007] S2. Define a multimodal task type database and build a multimodal task graph.
[0008] S3. Establish and update the candidate set of unit tasks.
[0009] Based on the generated multimodal task types and multimodal task graph features, unit tasks are extracted and segmented. The unit task is defined as: a modal data analysis and processing task that can be driven by a single functional unit within a large language model, belongs to one or a specific step in a multimodal task, and can be directly structurally mapped by the multimodal task graph; based on the unit task extraction and segmentation process, unit task modal query tags and corresponding context data are generated, and a candidate set of unit tasks is established;
[0010] S4, Task Reasoning Optimization
[0011] Construct a unit task database, perform historical task embedding vector similarity analysis on each unit task in the unit task database, set a similarity threshold, merge unit tasks with embedding vector similarity higher than the threshold, and add unmatched unit tasks to the database;
[0012] During each push process of the large language model, a joint screening is performed based on the current context and the unit task database to determine the unit task with the highest matching degree from the unit task database.
[0013] Based on the defined unit task, the modal element type and structural element in the original input are obtained and added to the context information of the current thrust result;
[0014] In each inference process, the task analysis is evaluated and processed based on the inference analysis results of the large language model, the inference results of the large language model are output, and the corresponding related data are output in combination with the multimodal task graph.
[0015] In a further improved or preferred implementation of the aforementioned large language model optimization method based on multimodal collaboration, step S1 specifically includes:
[0016] 1a. Construct feature rule templates for a multimodal large language model, and identify and extract feature structural elements based on the feature rule templates; the feature rule templates include the definitions of modal feature types and structural elements;
[0017] 1b. Analyze multimodal content, complete the classification of elements and corresponding semantic descriptions.
[0018] Acquire multimodal content input by the user, classify and recognize the multimodal data content based on machine recognition technology or classification recognition methods, and call the corresponding recognition methods according to different modal element types to extract elements within the modal elements;
[0019] 1c. Extract structural elements based on different modal feature types and generate data structure labels;
[0020] Based on the model feature rule template, feature labels and structural elements are extracted for different modal data to obtain the structural labels of the modal data.
[0021] A further improvement or preferred implementation of the aforementioned large language model optimization method based on multimodal synergy includes, in the case of multimodal element rule templates, the element types involved at least include text elements, image elements, data elements, and tag element types; the structural elements at least include: tag elements, description elements, feature elements, and state elements;
[0022] Tag elements refer to data such as characters, image blocks, images, voice, and codes used to determine the type and attributes of element objects; for example, names (person names, product names, company names, etc.), codes (serial numbers, IDs, etc.), and logo graphics (product identifiers, trademarks, symbols, etc.).
[0023] Description elements refer to data such as strings, images, videos, audio, and codes used to describe the basic information of an element; for example, descriptive text, promotional images, or videos of a unit or entity.
[0024] Feature elements refer to character and coded data used to define the basic characteristics of elements; such as function names, requirement names, entity definitions, parameters, etc.
[0025] A state element refers to the character or encoded data used to define the current state or state change of an element.
[0026] In a further improvement or preferred implementation of the aforementioned large language model optimization method based on multimodal collaboration, in step S2, the multimodal task type database includes at least three types of task types: information confirmation, behavior recognition, and prediction analysis.
[0027] Information confirmation tasks refer to matching and retrieval tasks targeting specific objectives.
[0028] Behavior recognition refers to tasks used to identify potential connections between targets and determine the relationships between them.
[0029] Predictive analytics refers to task types used to obtain information related to changes in a target.
[0030] The specific steps also include:
[0031] 2a. Based on the entities involved in the classification results of information confirmation tasks, complete the modal data entity alignment and add the entities to the nodes or edge sets in the multimodal task graph; construct graph anchor points based on the entities corresponding to the goals or results of information confirmation tasks and initialize graph triples; dynamically update the multimodal task graph relationships based on the entity correspondence relationships in information confirmation tasks.
[0032] 2b. Establish the logical path of modal elements according to the logical reasoning order of behavior recognition tasks, determine the execution chain of multimodal tasks, and perform verification and enhancement on the multimodal task graph nodes and edge sets based on the execution chain to complete the multimodal task graph enhancement;
[0033] 2c. Based on the task content involved in the predictive analysis task, extract data timestamps, generate a multimodal task map or time-series change map based on data entity relationships and task behavior connections, and establish a multimodal task map trend model based on entity relationships.
[0034] A further improvement or preferred implementation of the aforementioned large language model optimization method based on multimodal collaboration, wherein the modal data The structural label can be represented as:
[0035]
[0036] in It refers to the nth element label. This refers to the tag type and tag value of the k-th structure element.
[0037] Further improvements or preferred implementations of the aforementioned large language model optimization method based on multimodal joint methods include: establishing an attribute embedding module, using the attribute embedding module to encode and concatenate unique label elements of entities to form embedded labels, embedding and adjusting the dimensions of unique label elements by establishing a lightweight pre-trained model, finally outputting the corresponding embedding vector by pooling compression, and replacing the corresponding label element sequence in the data processing process with the embedding vector that has a simpler data form and less data volume.
[0038] In a further improvement or preferred implementation of the aforementioned large language model optimization method based on multimodal collaboration, in step S3, a trainable low-rank matrix is added as an additional parameter to the context data to control the vector matching process through the trainable low-rank matrix; the low-rank matrix is added as part of the vector matching projection matrix; it is used to control the output format of the control task context information and to adjust the control data output in a directional manner.
[0039] In a further improvement or preferred embodiment of the aforementioned large language model optimization method based on multimodal collaboration, in step S3, the unit task modal query label can be represented as: Context data can be represented as The established candidate set of unit tasks is ,in Indicates the unit task number. Indicates the time when the unit task was generated.
[0040] Further improvements or preferred implementations of the aforementioned large language model optimization method based on multimodal collaboration include evaluating and processing the task analysis based on the large language model inference analysis results, and outputting the large language model inference results, specifically including:
[0041] 1) Analyze the confidence of a unit task based on the semantic validity and semantic distinguishability scores of the element types and structural elements in the context information corresponding to the unit task. If the confidence of a unit task is lower than a given threshold for two consecutive times, or if the confidence score of a unit task decreases by a rate lower than a given threshold for several consecutive times, then stop the inference analysis.
[0042] 2) Based on the coverage of the unit tasks in the current unit task database with the expected task types of the large language model design, determine if the coverage meets the design requirements, then stop the inference;
[0043] 3) Compare the consistency scores of the convergence of the semantic features of the data structure labels. When the semantic consistency of the semantic features of the data structure labels of the unit task obtained by multiple rounds of reasoning approaches saturation, stop reasoning.
[0044] 4) Output the model task inference results based on the unit task data when the confidence level, coverage, or convergence is stable.
[0045] Its beneficial effects are as follows:
[0046] This application presents a multimodal joint large language model optimization method. In the analysis and processing of large language input and output data, it simplifies the multimodal data fusion process by establishing a system based on modal elements and specific element definitions and updates. It also performs data-driven analysis by extracting and segmenting complex multimodal processing tasks into unit tasks, thus avoiding the problem that a single large language model generally cannot directly perform joint processing on multimodal data. This provides a new processing scheme for compatible processing of multimodal data using different types of large language models. The scheme exhibits good compatibility with various large language models of different structures and training methods, and is easy to implement. Attached Figure Description
[0047] Figure 1 This is a flowchart illustrating a large language model optimization method based on multimodal collaboration. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0049] This application relates to a method for optimizing large language models based on multimodal collaboration. It is mainly used to provide a way to achieve multimodal data fusion processing and simplify complex training and inference tasks through fragmented unit processing. This reduces the problem that various large language models cannot directly handle complex multimodal data, and provides a way for various large language models to complete complex training and inference tasks.
[0050] Since different large language models may use different data formats and processing specifications, in order to ensure that multimodal data can be effectively utilized, necessary preprocessing should be performed first, namely step S1 multimodal data preprocessing.
[0051] S1. Multimodal data preprocessing
[0052] a. Construct feature rule templates for a multimodal large language model, and identify and extract feature structural elements based on the feature rule templates; the feature rule templates include the definitions of modal feature types and structural elements;
[0053] Multimodal large language model element rule template refers to one of the basic specifications for subsequent data analysis and processing of large language models. Among them, modal elements determine the basic types and representations of modal data that the large language model can process, while structural elements are the basis for analyzing effective information and related features in the corresponding multimodal data.
[0054] To ensure the rational analysis and utilization of multimodal data, the multimodal element rule template in this application includes at least text elements, image elements, data elements, and label element types.
[0055] The structural elements include at least: label elements, description elements, feature elements, and state elements;
[0056] Tag elements refer to data such as characters, image blocks, images, voice, and codes used to determine the type and attributes of element objects; for example, names (person names, product names, company names, etc.), codes (serial numbers, IDs, etc.), and logo graphics (product identifiers, trademarks, symbols, etc.).
[0057] Description elements refer to data such as strings, images, videos, audio, and codes used to describe the basic information of an element; for example, descriptive text, promotional images, or videos of a unit or entity.
[0058] Feature elements refer to character and coded data used to define the basic characteristics of elements; such as function names, requirement names, entity definitions, parameters, etc.
[0059] A state element refers to the character or encoded data used to define the current state or state change of an element;
[0060] It should be noted that the classification and judgment of the above structural elements are not based on the classification and judgment of the original multimodal input data directly. Instead, they should be obtained by extracting and separating the original multimodal input data using various element recognition technologies and then analyzing and judging it. For example, text, images or symbols contained in voice and video materials may also be tag elements, and some text elements may also belong to tag or data elements at the same time. For example, for a voice segment and a video segment in a video promotional material for a certain brand of facial tissue, the elements and element types shown in Table 1 can be obtained.
[0061] Table 1. Multimodal element rule analysis of a certain flat-screen facial tissue promotional video data
[0062]
[0063] b. Analyze multimodal content, complete the classification of elements and corresponding semantic descriptions.
[0064] Acquire multimodal content input by the user, classify and recognize the multimodal data content based on machine recognition technology or classification recognition methods, and call the corresponding recognition methods according to different modal element types to extract elements within the modal elements;
[0065] Generally, apart from text and images belonging to labels and numeric characters belonging to features, most other types of elements and structural elements need to be identified and processed by various manual or machine recognition technologies, including text recognition technology and image recognition and classification technology. When necessary, manual review or annotation is also required to actively annotate specific types of modal data classification and structural elements.
[0066] c. Extract structural elements based on different modal feature types and generate data structure labels;
[0067] Based on model feature rule templates, for different modal data Extract its feature labels and structural elements to obtain modal data. structural tags ,in It refers to the nth element label. This refers to the tag type and tag value of the k-th structural element;
[0068] Specifically, for entity elements that act as executors or requesters in various tasks, an attribute embedding module is established to build their entity attribute features. This ensures improved data processing efficiency and optimized performance of the large language model backend during entity alignment and logical association. The attribute embedding module encodes and concatenates the unique label elements of the entity to form embedded labels. Since the attributes and state types of the label elements are relatively simple, a lightweight pre-trained model can be built to embed and adjust the dimensions of the unique label elements. Finally, the corresponding embedding vector is output through pooling compression. The corresponding label element sequence in the data processing process is replaced with the embedding vector, which has a simpler data form and less data volume. This improves the processing efficiency of the large language model backend and reduces unnecessary analysis processes.
[0069] S2. Define a multimodal task type database and build a multimodal task graph based on the multimodal task type database;
[0070] The multimodal task type database includes at least three types of task types: information confirmation, behavior recognition, and predictive analysis.
[0071] Information confirmation tasks refer to matching and retrieval tasks targeting specific objectives.
[0072] Behavior recognition refers to tasks used to identify potential connections between targets and determine the relationships between them.
[0073] Predictive analytics refers to task types used to obtain information related to changes in a target.
[0074] The functions or uses performed by multimodal large language models involving the above-mentioned task categories can be decomposed into a series of multimodal task graphs composed of multimodal task instances and features such as the association attributes between task instances. As a type of knowledge graph, the multimodal task graph, in addition to containing feature data such as text, images, videos, and audio involved in multimodal tasks, also includes entity connection relationships or task relationships composed of ternary relation groups. Its most basic processing principles include:
[0075] 1) Based on the entities involved in the classification results of information confirmation tasks, complete the modal data entity alignment and add the entities to the nodes or edge sets in the multimodal task graph; construct graph anchor points based on the entities corresponding to the goals or results of information confirmation tasks and initialize graph triples; dynamically update the multimodal task graph relationships based on the entity correspondence in information confirmation tasks.
[0076] 2) Establish the logical path of modal elements according to the logical reasoning order of behavior recognition tasks, determine the execution chain of multimodal tasks, and verify and enhance the nodes and edge sets of the multimodal task graph based on the execution chain to complete the multimodal task graph enhancement;
[0077] 3) Based on the task content involved in the predictive analysis task, extract data timestamps, generate a heat multimodal task map or time-series change map based on data entity relationships and task behavior connections, and establish a multimodal task map trend model based on entity relationships. The multimodal task map or time-series change map and the multimodal task map trend model are used to provide a reference for unit task execution priority decision in the reasoning process of the large language model during the connection reasoning task.
[0078] Systematic learning and iterative upgrades of large language models during task processing are the foundation for them to complete complex tasks and adapt to new data. To meet the update requirements of multimodal data processing tasks, a multimodal task map is established. After decomposing complex multimodal tasks into unit tasks, the database labels involved in the unit tasks can be used to quickly locate and analyze subsequent or related unit tasks, and quickly establish a complete task chain and update process for complex multimodal analysis tasks, i.e., step S3.
[0079] S3. Establish and update the candidate set of unit tasks.
[0080] Based on the generated multimodal task type and the multimodal task map feature extraction and segmentation unit task, the unit task is defined as: a modal data analysis and processing task that can be driven by a single functional unit within a large language model, belongs to one of the multimodal tasks or a specific link within it, and can be directly structurally mapped by the multimodal task map.
[0081] Generate unit task modal query tags based on the unit task extraction and segmentation process. and the corresponding context data Establish a candidate set of unit tasks ,in Indicates the unit task number. Indicates the generation time of the unit task;
[0082] In particular, to facilitate the optimization and control of the context matching process of the large language model according to the preferences and weights of specific tasks, a trainable low-rank matrix is added as an additional parameter to the context data during the context modality data vector matching process, so as to control the vector matching process through the trainable low-rank matrix.
[0083] In particular, to facilitate the output format of the control unit's task context information, and to enable control data output and orientation adjustment, a low-rank matrix is added as part of the vector matching projection matrix;
[0084] S4, Task Reasoning Optimization
[0085] Construct a unit task database, perform historical task embedding vector similarity analysis on each unit task in the unit task database, set a similarity threshold, merge unit tasks with embedding vector similarity higher than the threshold, and add unmatched unit tasks to the database;
[0086] During each push process of the large language model, a joint screening is performed based on the current context and the unit task database to determine the unit task with the highest matching degree from the unit task database.
[0087] Based on the defined unit task, the modal element type and structural element in the original input are obtained and added to the context information of the current thrust result;
[0088] In each inference process, the task analysis is evaluated and processed based on the inference analysis results of the large language model, specifically referring to:
[0089] 1) Analyze the confidence of a unit task based on the semantic validity and semantic distinguishability scores of the element types and structural elements in the context information corresponding to the unit task. If the confidence of a unit task is lower than a given threshold for two consecutive times, or if the confidence score of a unit task decreases by a rate lower than a given threshold for several consecutive times, then stop the inference analysis.
[0090] 2) Based on the coverage of the unit tasks in the current unit task database with the expected task types of the large language model design, determine if the coverage meets the design requirements, then stop the inference;
[0091] 3) Compare the consistency scores of the convergence of the semantic features of the data structure labels. When the semantic consistency of the semantic features of the data structure labels of the unit task obtained by multiple rounds of reasoning approaches saturation, stop reasoning.
[0092] 4) Output the model task inference results based on the unit task data when the confidence, coverage or convergence is stable, and output the corresponding related data in combination with the multimodal task map.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A method for optimizing a large language model based on multimodal collaboration, characterized in that, Includes the following steps: S1. Multimodal data preprocessing: construct feature rule templates for multimodal large language models, identify and extract feature structural elements; complete feature and element classification and corresponding semantic descriptions, and generate data structure labels; S2. Define a multimodal task type database and build a multimodal task graph. S3. Establish and update the candidate set of unit tasks; Based on the generated multimodal task types and multimodal task graph features, unit tasks are extracted and segmented. The unit task is defined as: a modal data analysis and processing task that can be driven by a single functional unit within a large language model, belongs to one or a specific step in a multimodal task, and can be directly structurally mapped by the multimodal task graph; based on the unit task extraction and segmentation process, unit task modal query tags and corresponding context data are generated, and a candidate set of unit tasks is established; S4, Task Reasoning Optimization Construct a unit task database, perform historical task embedding vector similarity analysis on each unit task in the unit task database, set a similarity threshold, merge unit tasks with embedding vector similarity higher than the threshold, and add unmatched unit tasks to the database; During each push process of the large language model, a joint screening is performed based on the current context and the unit task database to determine the unit task with the highest matching degree from the unit task database. Based on the defined unit task, the modal element type and structural element in the original input are obtained and added to the context information of the current thrust result; In each inference process, the task analysis is evaluated and processed based on the inference analysis results of the large language model, the inference results of the large language model are output, and the corresponding related data are output in combination with the multimodal task graph.
2. The method for optimizing a large language model based on multimodal collaboration as described in claim 1, characterized in that, Step S1 specifically includes: 1a. Construct feature rule templates for a multimodal large language model, and identify and extract feature structural elements based on the feature rule templates; the feature rule templates include the definitions of modal feature types and structural elements; 1b. Analyze multimodal content, complete the classification of elements and corresponding semantic descriptions. Acquire multimodal content input by the user, classify and recognize the multimodal data content based on machine recognition technology or classification recognition methods, and call the corresponding recognition methods according to different modal element types to extract elements within the modal elements; 1c. Extract structural elements based on different modal feature types and generate data structure labels; Based on the model feature rule template, feature labels and structural elements are extracted for different modal data to obtain the structural labels of the modal data.
3. The method for optimizing a large language model based on multimodal collaboration according to claim 2, characterized in that, The multimodal feature rule template involves feature types that include at least text features, image features, data features, and tag features; the structural elements include at least: tag elements, description elements, feature elements, and state elements. Tag elements refer to data such as characters, image blocks, images, voice, and codes used to determine the type and attributes of element objects; for example, names (person names, product names, company names, etc.), codes (serial numbers, IDs, etc.), and logo graphics (product identifiers, trademarks, symbols, etc.). Description elements refer to data such as strings, images, videos, audio, and codes used to describe the basic information of an element; for example, descriptive text, promotional images, or videos of a unit or entity. Feature elements refer to character and coded data used to define the basic characteristics of elements; such as function names, requirement names, entity definitions, parameters, etc. A state element is a character or encoded data used to define the current state or state change of an element.
4. The method for optimizing a large language model based on multimodal collaboration according to claim 1, characterized in that, In step S2, the multimodal task type database includes at least three types of task types: information confirmation, behavior recognition, and prediction analysis. Information confirmation tasks refer to matching and retrieval tasks targeting specific objectives. Behavior recognition refers to tasks used to identify potential connections between targets and determine the relationships between them. Predictive analytics refers to task types used to obtain information related to changes in a target. The specific steps also include: 2a. Based on the entities involved in the classification results of information confirmation tasks, complete the modal data entity alignment and add the entities to the nodes or edge sets in the multimodal task graph; construct graph anchor points based on the entities corresponding to the goals or results of information confirmation tasks and initialize graph triples; dynamically update the multimodal task graph relationships based on the entity correspondence relationships in information confirmation tasks. 2b. Establish the logical path of modal elements according to the logical reasoning order of behavior recognition tasks, determine the execution chain of multimodal tasks, and perform verification and enhancement on the multimodal task graph nodes and edge sets based on the execution chain to complete the multimodal task graph enhancement; 2c. Based on the task content involved in the predictive analysis task, extract data timestamps, generate a multimodal task map or time-series change map based on data entity relationships and task behavior connections, and establish a multimodal task map trend model based on entity relationships.
5. The method for optimizing a large language model based on multimodal collaboration according to claim 4, characterized in that, The modal data The structural label can be represented as: in It refers to the nth element label. This refers to the tag type and tag value of the k-th structure element.
6. The method for optimizing a large language model based on multimodal collaboration according to claim 4, characterized in that, It also includes: establishing an attribute embedding module, using the attribute embedding module to encode and concatenate the unique label elements of the entity to form embedded labels, establishing a lightweight pre-trained model to embed and adjust the dimensions of the unique label elements, and finally outputting the corresponding embedding vector through pooling compression, and replacing the corresponding label element sequence in the data processing process with the embedding vector with a simpler data form and less data volume.
7. The method for optimizing a large language model based on multimodal collaboration according to claim 1, characterized in that, In step S3, a trainable low-rank matrix is added as an additional parameter to the context data to control the vector matching process through the trainable low-rank matrix. Add a low-rank matrix as part of the vector matching projection matrix; this is used to control the output format of the control task context information and to orient the control data output.
8. The method for optimizing a large language model based on multimodal collaboration according to claim 1, characterized in that, In step S3, the unit task modal query tag can be represented as: Context data can be represented as The established candidate set of unit tasks is ,in Indicates the unit task number. Indicates the time when the unit task was generated.
9. The method for optimizing a large language model based on multimodal collaboration according to claim 1, characterized in that, Based on the reasoning analysis results of the large language model, the task analysis is evaluated and processed, and the reasoning results of the large language model are output, specifically including: 1) Analyze the confidence of a unit task based on the semantic validity and semantic distinguishability scores of the element types and structural elements in the context information corresponding to the unit task. If the confidence of a unit task is lower than a given threshold for two consecutive times, or if the confidence score of a unit task decreases by a rate lower than a given threshold for several consecutive times, then stop the inference analysis. 2) Based on the coverage of the unit tasks in the current unit task database with the expected task types of the large language model design, determine if the coverage meets the design requirements, then stop the inference; 3) Compare the consistency scores of the convergence of the semantic features of the data structure labels. When the semantic consistency of the semantic features of the data structure labels of the unit task obtained by multiple rounds of reasoning approaches saturation, stop reasoning. 4) Output the model task inference results based on the unit task data when the confidence level, coverage, or convergence is stable.