Multi-modal large model reasoning method and system based on self-driven feedback and symbol collaboration

By employing a multimodal large-scale model reasoning framework based on self-driven feedback and symbolic collaboration, and utilizing knowledge graphs and symbolic logic rules, combined with human feedback to optimize the reasoning process of the multimodal large-scale model, the problem of insufficient interpretability and self-driven feedback in complex tasks of the multimodal large-scale model is solved, and the autonomous evolution and logical improvement of the model are realized.

CN121480705AActive Publication Date: 2026-02-06XI AN JIAOTONG UNIV

Patent Information

Application Number
CN202511552385.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-02-06
Estimated Expiration
2045-10-28

AI Technical Summary

Technical Problem

When dealing with fine-grained reasoning or complex tasks involving multi-attribute logical relationships, multimodal large models suffer from a lack of interpretability of reasoning results and insufficient self-driven feedback, making it difficult to adapt to the needs of dynamically changing complex scenarios.

Method used

By constructing a multimodal large-scale model reasoning framework based on self-driven feedback and symbolic collaboration, utilizing knowledge graphs to structurally represent multimodal information, combining symbolic logic rules and human feedback, and designing a hybrid reward function, autonomous evolutionary cyclic training is achieved, thereby enhancing the model's logicality and interpretability.

Benefits of technology

It enhances the logic, interpretability, and autonomous evolution capabilities of multimodal large models in complex reasoning tasks, and achieves self-driven feedback and continuous optimization without the need for a large amount of manually labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121480705A_ABST
    Figure CN121480705A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal large model reasoning method and system based on self-driven feedback and symbol collaboration, and the method comprises the steps: carrying out the structural representation of multi-modal information through a knowledge graph, carrying out the entity recognition and relation extraction in combination with a large model, constructing a unified knowledge graph, and generating a knowledge ternary set; defining a symbol logic expression, constructing a diversified symbol logic rule by using the knowledge ternary set, and calculating a symbol consistency award of a reasoning path; constructing a symbol-human feedback collaborative reward mechanism to obtain a mixed reward function; in the process of interacting with the multi-modal environment, sampling a group of outputs for specific tasks, and constructing an intra-group relative reward optimization strategy network objective function in a multi-task scene in combination with a mixed reward function; the system interacts with the environment to realize autonomous evolution cycle to generate a training sample, and iterative cycle realizes self-driven feedback without a large amount of manual annotation data; and the logicality, the interpretability and the autonomous evolution ability of the multi-modal large model in a complex reasoning task are promoted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of multi-modal artificial intelligence reasoning system, and particularly relates to a multi-modal large model reasoning framework based on self-driven feedback and neural-symbol cooperation. BACKGROUND

[0002] As a cutting-edge technology paradigm in the field of artificial intelligence, Multimodal Large Language Models (MLLM) surpasses the limitations of traditional single modal, and can understand, process and generate information from multiple different sources at the same time, including text, images, audio, video and even depth, infrared and other sensor data. The core of such models is their huge number of parameters and pre-training on massive cross-modal datasets, which enables them to learn deep relationships and alignment between complex modalities. Through a shared representation space, it encodes and fuses information from different modalities, achieving cross-modal semantic understanding and creative output.

[0003] However, multimodal large models rely on a large amount of training data to learn patterns and relationships, but they perform poorly when dealing with fine-grained reasoning or complex tasks involving multi-attribute logical relationships. The reasoning process generally has the following shortcomings: (1) The reasoning mechanism based on autoregressive generation essentially captures large-scale correlation patterns rather than strict logical deduction; (2) It is difficult for the model to incorporate structured logical rules into the reasoning process, resulting in a lack of explainability in the reasoning results; (3) Traditional frameworks rely on human-labeled data for optimization and cannot achieve continuous iteration of reasoning strategies through self-driven feedback, making it difficult to adapt to dynamic and complex scenario requirements.

[0004] These technical bottlenecks severely restrict the in-depth application of multimodal large models in key fields. To address these issues, numerous related works have emerged in recent years. Existing technology, patent CN119670878A, provides a self-reinforcement learning method for multimodal large models based on thought chain guidance. Its core lies in iteratively fine-tuning the model through thought chains generated by the model itself, thus solving the problem of poor complex reasoning ability in multimodal models at low cost. This method is the first to propose a self-reinforcement learning approach for multimodal large models based on thought chains. Using existing visual reasoning question-answering datasets, it guides the multimodal model to generate accurate thought chains, thereby constructing a high-quality complex reasoning fine-tuning dataset and iteratively enhancing the complex reasoning ability of the multimodal model. Although iteratively fine-tuning the model through self-generated thought chains, the thought chain generation process does not explicitly integrate symbolic knowledge, relying mainly on the model's own learning from the data. This lack of deep collaboration between symbolic logic and neural networks limits the interpretability and logical rigor of the reasoning process. Furthermore, the self-reinforcement process does not introduce external symbolic rules for guidance. When handling complex reasoning tasks requiring strict logical constraints, the lack of explicit verification at the symbolic level may lead to reasoning biases. The publication number CN120046711A provides a logic reasoning-driven knowledge graph-enhanced large model generation method. Its core lies in constructing a knowledge graph, including dataset preparation; triple extraction and filtering; triple transformation and knowledge graph construction; original query question acquisition; user question keyword extraction; graph query engine construction and query execution; knowledge integration and generation. It enhances the large model generation capability by leveraging knowledge graphs and logical reasoning, but it mainly relies on pre-constructed triples and static knowledge graphs, lacking a dynamic self-driven feedback mechanism. It is difficult to optimize the knowledge graph construction and query strategy based on real-time feedback during the reasoning process, and the injection of symbolic logic relies heavily on manually selected triples, making it insufficiently adaptable to dynamically changing logical relationships in complex scenarios. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention provides a multimodal large model reasoning framework based on self-driven feedback and symbolic collaboration. This method addresses the insufficient logical reasoning ability of multimodal large models by realizing a reasoning framework that combines interpretability and autonomous evolution capabilities, thereby promoting the logicality, interpretability, and autonomous evolution capabilities of multimodal large models in complex reasoning tasks.

[0006] To achieve the above objectives, in a first aspect, the present invention provides a multimodal large model inference method based on self-driven feedback and symbolic collaboration, comprising the following steps: S1. Acquire multimodal data, represent multimodal information in a structured way through knowledge graphs, and combine with large models to perform entity recognition and relationship extraction, construct a unified knowledge graph, and generate a knowledge ternary set; S2, define symbolic logic expressions, construct diverse symbolic logic rules using knowledge ternary sets, and calculate the symbolic consistency reward of the reasoning path. ; S3, combining symbolic consistency rewards, constructs a symbolic-human feedback collaborative reward mechanism to obtain a hybrid reward function; S4, during the interaction with the multimodal environment, samples a set of outputs for a specific task and combines them with a hybrid reward function to construct the objective function of the intra-group relative reward optimization strategy network in the multi-task scenario; S5 interacts with the environment to achieve autonomous evolution and cyclical generation of training samples, iterating through S2-S4 to achieve self-driven feedback without the need for a large amount of manually labeled data.

[0007] Furthermore, in S1, the symbolic knowledge representation module uses cosine similarity to align multimodal information and employs a large model to perform structured knowledge extraction on the fused multimodal data, including entity recognition and relation extraction, to generate a knowledge ternary set.

[0008] Furthermore, in S2, when constructing diverse symbolic logic rules using knowledge ternary sets, cross-modal consistency rules are used. Attribute matching rules causal relationship rules And logical transitivity rules, calculating symbolic consistency rewards for multiple multimodal large model inference paths: .

[0009] Furthermore, the cross-modal consistency rule requires that entities appearing in modality A should also appear in modality B. Let there be two modalities: text and image. A pointer function that indicates if the text contains entities. The value is 1 if the image is positive and 0 otherwise; define the image detection function. Define cross-modal consistency score for:

[0010] in, For entities The weight, Candidate inference trajectories; Attribute matching rules, for the same entity Let the attributes extracted from the text be... Define the attribute distance function Cosine similarity, attribute matching score Defined as:

[0011] Causal relationship rules are established by predefining a set of causal relationships in the knowledge graph. , where each relationship Indicates an event Causally leading to the event ,set up Indicates an event In the reasoning path If the time sequence or position index is used, then the causal relationship rule requires: ,if Then it satisfies The causal consistency score is defined as follows:

[0012] in, For relationship The weights; The logical transitivity rule reflects the classic syllogistic relation, that is, if... and If it exists, then it should be deduced. Logical transit score Defined as:

[0013] in, For candidate inference trajectories, This refers to a set of propositions or logical relationships that exist within the trajectory. The indicator function is defined as follows:

[0014] This represents all event triples that satisfy the preconditions. The weights of the event triples; Finally, based on the above logical rules, the symbolic consistency reward for multiple multimodal large model inference paths is calculated:

[0015] in, The weights for each part of the score, For candidate inference trajectories, this rule base supports extensions, and the general expression is: .

[0016] Furthermore, in S3, a reward mechanism based on symbolic-human feedback collaboration is constructed, including: Symbolic reward design, the symbolic reward design directly inherits the symbolic consistency reward obtained from S2. ; Human feedback rewards are obtained by comparing and ranking candidate path pairs according to human preferences, and a lightweight model is learned to predict the probability of human preference, thus obtaining human feedback rewards. ; Environmental execution rewards are given based on the outcome of environmental validation, which can be either failure or success. ; The hybrid reward function is a weighted sum of symbolic rewards, human feedback rewards, and environmental performance rewards. The weights vary depending on the task type; logical tasks emphasize symbolic rewards, while open tasks emphasize human preferences.

[0017] Furthermore, in S4, during interaction with the multimodal environment, for tasks... Sample a set of outputs ,in The group size, combined with a hybrid reward function, forms the objective function of the intra-group relative reward optimization strategy network in a multi-task scenario, including: First, combine mixed rewards Calculate the standardized dominance function ; Then, we introduce the ranking function. This ensures an increased probability of paths with high sign consistency, generating a ranking advantage function. :

[0018] in, For each task From the current strategy A set of outputs sampled in the middle , To be according to Rank the paths within the group from highest to lowest; Obtain the mixed advantage function ,in These are dynamic weighting coefficients that vary depending on the task type. Subsequently, based on the hybrid advantage function, by maximizing Objective function update strategy model :

[0019]

[0020] in, and It's a hyperparameter. It is the dominant function, and This represents the strategy from the previous iteration. This represents the strategy that needs to be updated. For reference strategy, For the distribution of tasks, For the inference path within the group, For the clipping function, The KL distance between the old and new strategies is calculated using the following formula: .

[0021] Furthermore, symbolic constraint regularization is introduced during the model reinforcement learning process to force the policy not to deviate from the symbolic rule base:

[0022] in, This is the scoring threshold; when the symbol consistency score is below the threshold... When this happens, the probability corresponding to that path will be penalized. Minimizing this penalty during training makes the finally learned policy more inclined to choose actions with high symbolic consistency, resulting in the loss function:

[0023] in, This is a hyperparameter.

[0024] Furthermore, in S5, new candidate inference paths are generated in the environment as training samples, and symbolic consistency rewards are calculated based on the training samples. The mixed reward function is obtained by combining the symbolic consistency rewards, forming a self-driven feedback loop to generate a set of new candidate inference paths. The symbolic consistency rewards of the new candidate inference paths are calculated, and the mixed reward is generated using the symbolic consistency rewards. The model policy is updated by maximizing the objective function using the mixed reward.

[0025] Secondly, this invention provides a multimodal large-model inference system based on self-driven feedback and symbolic collaboration, comprising the following steps: The symbolic knowledge representation module is based on acquiring multimodal data, using knowledge graphs to structurally represent multimodal information, and combining large models to perform entity recognition and relation extraction, thereby constructing a unified knowledge graph. The symbolic knowledge reasoning module is used to define symbolic logic expressions, construct diverse symbolic logic rules, and calculate the symbolic consistency reward of the reasoning path; The reward signal design and fusion module is used to construct a reward mechanism that coordinates symbolic and human feedback, resulting in a hybrid reward function; The strategy optimization module is used to sample a set of outputs for a specific task during the interaction with the multimodal environment, take the mixed reward as input, and construct the objective function of the intra-group relative reward optimization strategy network in the multi-task scenario. The autonomous evolution loop module interacts with the environment to autonomously generate training samples, which are then fed into the symbolic knowledge processing and reasoning unit to obtain symbolic consistency rewards. These rewards are then injected into the reward signal design and fusion module and the policy optimization module, achieving self-driven feedback without the need for a large amount of manually labeled data.

[0026] Thirdly, the present invention provides a computer device, including a processor and a memory, wherein the memory is used to store a computer executable program, the processor reads part or all of the computer executable program from the memory and executes it, and the processor can realize the above-mentioned multimodal large model inference method based on self-driven feedback and symbolic collaboration when executing part or all of the computer executable program.

[0027] Finally, a computer-readable storage medium may be provided, in which a computer program is stored, which, when executed by a processor, can realize the above-mentioned multimodal large model inference method based on self-driven feedback and symbolic collaboration.

[0028] Compared with existing technologies, this invention has at least the following beneficial effects: This invention constructs a symbolic neural collaborative reasoning method that integrates symbolic reasoning and reinforcement learning to build self-driven feedback. The innovation lies in the following: First, it constructs a self-driven feedback symbolic collaborative reasoning method, integrating symbolic reasoning and reinforcement learning to enhance feedback capabilities. Second, it designs a symbolic mechanism generation and injection mechanism, extracting a symbolic knowledge base and rule base from a multimodal large model based on a knowledge graph. Based on this, it calculates the symbolic consistency score of multiple reasoning paths and injects it into the reinforcement learning reward function and policy optimization loss, providing the multimodal large model with clearer and more interpretable reasoning paths, enhancing its reasoning consistency and accuracy in complex and diverse tasks. Finally, it proposes a reinforcement learning strategy that integrates symbolic reasoning and human feedback to dynamically optimize the reasoning strategy and achieve continuous autonomous evolution. In summary, this invention proposes an innovative reasoning method integrating symbolic reasoning, adaptive optimization, and reinforcement learning to improve the logic, interpretability, and autonomous evolution capabilities of multimodal large models in complex reasoning tasks.

[0029] To further understand the features and technical content of the present invention, please refer to the following detailed description and drawings of the present invention. However, the drawings provided are for reference and illustration only and are not intended to limit the present invention. Attached Figure Description

[0030] Figure 1 This is a schematic diagram of the framework of the present invention; Figure 2 This is a flowchart of the algorithm of this invention. Detailed Implementation

[0031] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, exemplary in nature, and intended to explain this invention, and should not be construed as limiting this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort should fall within the scope of protection of this invention.

[0032] This invention provides a multimodal large-scale model inference system based on self-driven feedback and symbolic collaboration. It couples a symbolic knowledge processing and inference unit and a symbolic reinforcement learning optimization unit, which work collaboratively to enhance the model's inference capabilities. The symbolic knowledge processing and inference unit, serving as the knowledge foundation, endows the model with logical reasoning capabilities and cross-modal consistency through its internal symbolic knowledge representation and inference modules, and provides symbolic reward signals for the optimization process. The symbolic reinforcement learning optimization unit, acting as the optimization engine, autonomously learns and evolves the optimal inference path through its internal reward signal design and fusion module, policy optimization module, and autonomous evolution loop module, combined with the symbolic signals provided by the aforementioned units and environmental feedback.

[0033] Specifically, the symbolic knowledge processing and reasoning unit includes a symbolic knowledge representation module and a symbolic knowledge reasoning module; The symbolic knowledge representation module uses knowledge graphs to structurally represent multimodal information and combines them with large models to perform entity recognition and relation extraction, thereby constructing a unified knowledge graph. The symbolic knowledge reasoning module defines symbolic logical expressions, constructs a symbolic reasoning rule base, supports symbolic reasoning paths for multimodal relationships, and finally calculates the symbolic consistency score of the reasoning path to provide symbolic reward signals for reinforcement learning and guide the model to optimize logical constraints.

[0034] The symbolic reinforcement learning optimization unit introduces an interactive evolutionary self-training paradigm with symbolic injection. The unit comprises a reward signal design and fusion module, a policy optimization module, and an autonomous evolutionary loop module. The reward signal design and fusion module constructs a reward mechanism that integrates symbolic and human feedback. Based on this, the policy optimization module utilizes these reward signals to adjust the model's intrinsic policy, making it more inclined to generate higher-scoring inference results. The autonomous evolutionary loop module achieves data-driven self-evolution without extensive manual annotation through continuous interaction with the environment.

[0035] like Figure 1 and Figure 2 As shown in the figure, an embodiment of the present invention provides a multimodal large model inference method based on self-driven feedback and symbolic collaboration, comprising the following steps: S1. Acquire multimodal data. The symbolic knowledge representation module uses a knowledge graph to structurally represent multimodal information and combines it with a large model for entity recognition and relation extraction to construct a unified knowledge graph. Specifically, cosine similarity is used to align multimodal information, given two modality vectors. and (e.g., text and images), cosine similarity is defined as:

[0036] When the similarity is greater than the preset threshold At that time, it was assumed that the two were semantically related, thus establishing a correspondence in subsequent knowledge extraction.

[0037] Subsequently, a large model was used to extract structured knowledge from the fused multimodal data. The structured knowledge extraction included entity recognition and relation extraction.

[0038] Entity recognition, for text data, utilizes a pre-trained BERT model. Perform entity recognition: ,in, This refers to the set of entities extracted from text. For data such as images and sensor data, modality-aligned text descriptions can be used to assist in entity recognition. Subsequently, the GPT-4 model can be utilized. Combining contextual information with the identified entities Extract the relationships between entities to generate a knowledge ternary set. :

[0039] Each triplet is in the form of ,in , An entity represents a specific object or concept. Describe the relationship between two entities.

[0040] S2 completes the definition of symbolic logic expressions in the symbolic knowledge reasoning module, constructs diverse symbolic logic rules, calculates the symbolic consistency reward of the reasoning path, and provides symbolic reward signals for the symbolic reinforcement learning optimization unit.

[0041] First, there's the cross-modal consistency rule. In a multimodal scenario, assuming there are two modalities: text and image, let... A pointer function that indicates if the text contains entities. The value is 1 if the function is active, and 0 otherwise; similarly, an image detection function can be defined. Cross-modal consistency requires that entities described in the text also appear in the image. To measure consistency, a cross-modal consistency score is defined as:

[0042] in For entities The weight, These are candidate inference trajectories.

[0043] Secondly, regarding attribute matching rules, for the same entity Assuming the attributes extracted from the text are Define the attribute distance function. If we use cosine similarity, then attribute matching requires the distance between the two attributes to be sufficiently small. The attribute matching score can be defined as:

[0044] in, The higher the score (the smaller the negative distance), the better the attribute matching.

[0045] The causal relationship rules are defined as follows: A predefined set of causal relationships is established in the knowledge graph. , where each relationship Indicates an event Causally leading to the event .make Indicates an event In the reasoning path If the time sequence or position index is used, then the causal relationship rule requires: ,if Then it satisfies Therefore, the causal consistency score is defined as:

[0046] in, For relationship The weight.

[0047] Finally, we define the logical transitivity rule, which reflects the classic syllogistic relation, namely, if... and If it exists, then it should be deduced. If the candidate reasoning trajectory contains a set of propositions or logical relationships, then the indicator function is defined as follows:

[0048] The logical transitivity score is:

[0049] in, This represents all event triples that satisfy the preconditions. Its weight.

[0050] Finally, based on the above logical rules, the symbolic consistency reward for multiple multimodal large model inference paths is calculated:

[0051] in, The weights for each part of the score are used to balance the importance of different rules. This rule base can be expanded in future research, especially for specific scenarios. Therefore, the general expression for the overall score is:

[0052] Inject the symbol consistency reward into the reward signal design and fusion module.

[0053] S3 uses a reward signal design and fusion module to construct a reward mechanism that facilitates symbolic-human feedback collaboration. The first step is the symbolic reward design, which directly inherits the symbolic consistency reward obtained from the symbolic knowledge processing and reasoning unit.

[0054] Secondly, human feedback rewards ( Modeling. Collect small sample preference annotations, i.e., human comparison and ranking of candidate path pairs. ,in express Better. Assuming human reasoning skills are better than candidate inference paths. and The preferences follow the following probability distribution:

[0055] in, It is a path The implicit utility value, i.e., a lightweight model, is learned through a neural network.

[0056] Candidate paths The multimodal information is encoded as a vector, and implicit utility values ​​are mapped through a fully connected layer to minimize the cross-entropy loss.

[0057] Obtain a lightweight model trained using a preference prediction model. Predicting the probability of human preferences. Mapping utility values ​​to... The interval is thus obtained. :

[0058] in, It represents all candidate paths in the current batch.

[0059] Adding environmental performance rewards, we obtain the hybrid reward function:

[0060] This is the environmental verification result (such as the correctness of program execution and the success rate of robot actions), with a value of 0 (failure) or 1 (success). These are dynamic weighting coefficients that vary depending on the task type. For example, logic tasks emphasize symbolic rewards, while open tasks emphasize human preferences. In the early stages of training, symbolic rewards are prioritized to ensure the correctness of basic logic; in the later stages of training, the weight of human feedback is gradually increased to enhance the interpretability of the results.

[0061] S4, in the process of interacting with the multimodal environment, is designed for specific tasks. Sample a set of outputs ,in, It is the size of the group. Mixed rewards will be applied. The input strategy optimization module constructs the objective function of the intra-group relative reward optimization strategy network in a multi-task scenario.

[0062] First, calculate the standardized dominance function. Calculate the mean of the paths within the group. Standard deviation Thus we get:

[0063] To ensure an increased probability of high-symmetric-consistency paths, a ranking function is introduced. Ranking can better reflect "relative good or bad" rather than "absolute value size," thus preventing outliers from dominating gradient updates. Therefore, by Rank the paths within the group from highest to lowest, and generate a ranking advantage function:

[0064] This yields a mixed dominance function, which combines numerical differences with ranking information:

[0065] These are dynamic weighting coefficients that vary depending on the task type.

[0066] Subsequently, based on the hybrid advantage function, by maximizing Objective function to update policy model :

[0067]

[0068] in and It's a hyperparameter. It is the dominant function, and This represents the strategy from the previous iteration. This represents the strategy that needs to be updated. For reference strategy, For the distribution of tasks, For the inference path within the group, For the clipping function, The KL distance between the old and new strategies is calculated using the following formula:

[0069] Its main function is to constrain the current policy from changing too quickly, that is, to require that its KL divergence with the reference policy not be too large.

[0070] To ensure that the model automatically tends to generate logically consistent and symbolically consistent inference results during reinforcement learning, symbolic constraint regularization is further introduced to force the policy not to deviate from the symbolic rule base:

[0071] in, This is the scoring threshold. When the symbol consistency score is below the threshold... When this happens, the probability corresponding to that path will be penalized. This regularization term aims to minimize this penalty during training, making the finally learned policy more inclined to choose actions with high symbolic consistency.

[0072] Finally, the loss function is obtained:

[0073] in This is a hyperparameter.

[0074] S5 achieves autonomous evolution loop through autonomous evolution loop module. This module generates training samples autonomously through environmental interaction, inputs them into symbolic knowledge processing and reasoning unit to obtain symbolic consistency reward, and injects the reward signal design and fusion module and policy optimization module to achieve self-driven feedback without the need for a large amount of manually labeled data.

[0075] To enable the model to continuously adapt to dynamically changing and complex scenarios and achieve unsupervised, self-driven improvement in reasoning capabilities, rather than continuously relying on large amounts of manually labeled data, the autonomous evolution loop module establishes continuous interaction between the model and the multimodal environment, enabling the autonomous generation, evaluation, and reuse of training samples. Specifically, the autonomous evolution loop module guides the model to generate new candidate inference paths in the environment as training samples. These samples are then sent back to the symbolic knowledge processing and inference unit for symbolic consistency reward calculation. The rewards are then injected into the reward signal design and fusion module and the policy optimization module, forming a self-driven feedback loop. This loop allows the model to iteratively optimize its inference strategy through trial and error and learning, achieving autonomous evolution without a large amount of manually labeled data. Its working mechanism includes: Generate candidate paths: The model generates a new set of candidate inference paths based on the interaction between the current policy and the environment.

[0076] Feedback Acquisition and Evaluation: These generated paths are sent to the symbolic knowledge processing and reasoning unit to calculate their symbolic consistency rewards. Simultaneously, these paths are also evaluated by the reward signal design and fusion module to generate hybrid rewards (including symbolic consistency rewards, human feedback rewards, and environmental performance rewards).

[0077] Injection optimization: The calculated symbolic consistency reward is injected into the reward signal design and fusion module, and the mixed reward is injected into the policy optimization module. The policy optimization module uses these feedback signals to update the model policy by maximizing the objective function.

[0078] Iterative loop: The optimized model then uses a new strategy to generate the next batch of candidate paths, repeating the process of generating candidate paths, obtaining and evaluating feedback, and injecting optimization, thus forming a continuous evolutionary loop of generation-evaluation-optimization.

[0079] When performing inference using the method described in this application, candidate paths are first generated for the input task. Model generation Candidate reasoning paths Grouped by task type Remove violations of security thresholds (e.g.) The algorithm first identifies the path to the target area, then calculates a hybrid reward function based on symbolic representation and human feedback, assesses relative advantage within the group, and incorporates ranking advantage. Then, it guides the policy network update through advantage, ensuring symbolic consistency and suppressing the probability of low-symbol-score paths, gradually completing gradient descent through interaction with the environment.

[0080] The overall process can be summarized as follows: model generates candidate paths → symbolic and human feedback evaluation → policy optimization. Through this iterative process, the model gains rich feedback information by continuously interacting with the environment. Each cycle of generation, evaluation, and optimization accumulates valuable experience for the model, enabling it to continuously correct and improve its policies. Ultimately, without a large amount of manually labeled data, the model achieves autonomous evolution through its own feedback mechanism, continuously moving towards higher performance levels.

[0081] In addition, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can realize the multimodal large model inference method based on self-driven feedback and symbolic collaboration described in the present invention.

[0082] The present invention can also provide a computer device, including a processor and a memory, wherein the memory is used to store a computer executable program, the processor reads the computer executable program from the memory and executes it, and the processor can realize the multimodal large model inference method based on self-driven feedback and symbolic collaboration described in the present invention when executing the computer executable program.

[0083] The computer device may be a laptop, a desktop computer, or a workstation.

[0084] The processor can be a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or an off-the-shelf programmable gate array (FPGA).

[0085] The memory described in this invention can be an internal storage unit of a laptop, desktop computer, or workstation, such as memory or hard disk; or it can be an external storage unit, such as a portable hard disk or flash memory card.

[0086] Computer-readable storage media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media can include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. Random access memory can include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).

[0087] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multimodal large model inference method based on self-driven feedback and symbolic collaboration, characterized in that, Includes the following steps: S1. Acquire multimodal data, represent multimodal information in a structured way through knowledge graphs, and combine with large models to perform entity recognition and relationship extraction, construct a unified knowledge graph, and generate a knowledge ternary set; S2, define symbolic logic expressions, construct diverse symbolic logic rules using knowledge ternary sets, and calculate the symbolic consistency reward of the reasoning path. ; S3, combining symbolic consistency rewards, constructs a symbolic-human feedback collaborative reward mechanism to obtain a hybrid reward function; S4, during the interaction with the multimodal environment, samples a set of outputs for a specific task and combines them with a hybrid reward function to construct the objective function of the intra-group relative reward optimization strategy network in the multi-task scenario; S5 interacts with the environment to achieve autonomous evolution and cyclical generation of training samples, iterating through S2-S4 to achieve self-driven feedback without the need for a large amount of manually labeled data.

2. The multimodal large model inference method based on self-driven feedback and symbolic collaboration according to claim 1, characterized in that, In S1, the symbolic knowledge representation module uses cosine similarity to align multimodal information and employs a large model to extract structured knowledge from the fused multimodal data, including entity recognition and relation extraction, to generate a knowledge ternary set.

3. The multimodal large model inference method based on self-driven feedback and symbolic collaboration according to claim 1, characterized in that, In S2, when constructing diverse symbolic logic rules using knowledge ternary sets, cross-modal consistency rules are used. Attribute matching rules causal relationship rules And logical transit rules, The cross-modal consistency rule requires that entities appearing in modality A should also appear in modality B. Let there be two modalities: text and image. A pointer function that indicates if the text contains entities. The value is 1 if the image is positive and 0 otherwise; define the image detection function. Define cross-modal consistency score for: in, For entities The weight, Candidate inference trajectories; Attribute matching rules, for the same entity Let the attributes extracted from the text be... Define the attribute distance function Cosine similarity, attribute matching score Defined as: Causal relationship rules are established by predefining a set of causal relationships in the knowledge graph. , where each relationship Indicates an event Causally leading to the event ,set up Indicates an event In the reasoning path If the time sequence or position index is used, then the causal relationship rule requires: ,if Then it satisfies The causal consistency score is defined as follows: in, For relationship The weights; The logical transitivity rule reflects the classic syllogistic relation, that is, if... and If it exists, then it should be deduced. Logical transit score Defined as: in, For candidate inference trajectories, This refers to a set of propositions or logical relationships that exist within the trajectory. The indicator function is defined as follows: This represents all event triples that satisfy the preconditions. The weights of the event triples; Finally, based on the above logical rules, the symbolic consistency reward for multiple multimodal large model inference paths is calculated: in, The weights for each part of the score, For candidate inference trajectories, this rule base supports extensions, and the general expression is: 。 4. The multimodal large model inference method based on self-driven feedback and symbolic collaboration according to claim 1, characterized in that, In S3, a reward mechanism that combines symbolic and human feedback is constructed, including: Symbolic reward design, the symbolic reward design directly inherits the symbolic consistency reward obtained from S2. ; Human feedback rewards are obtained by comparing and ranking candidate path pairs according to human preferences, and a lightweight model is learned to predict the probability of human preference, thus obtaining human feedback rewards. ; Environmental execution rewards are given based on the outcome of environmental validation, which can be either failure or success. ; The hybrid reward function is a weighted sum of symbolic rewards, human feedback rewards, and environmental performance rewards. The weights vary depending on the task type; logical tasks emphasize symbolic rewards, while open tasks emphasize human preferences.

5. The multimodal large model inference method based on self-driven feedback and symbolic collaboration according to claim 1, characterized in that, In S4, during interaction with the multimodal environment, for the task... Sample a set of outputs ,in The group size, combined with a hybrid reward function, forms the objective function of the intra-group relative reward optimization strategy network in a multi-task scenario, including: First, combine mixed rewards Calculate the standardized dominance function ; Then, we introduce the ranking function. This ensures an increased probability of paths with high sign consistency, generating a ranking advantage function. : in, For each task From the current strategy A set of outputs sampled in the middle , To be according to Rank the paths within the group from highest to lowest; Obtain the mixed advantage function ,in These are dynamic weighting coefficients that vary depending on the task type. Subsequently, based on the hybrid advantage function, by maximizing Objective function update strategy model : in, and It's a hyperparameter. It is the dominant function, and This represents the strategy from the previous iteration. This represents the strategy that needs to be updated. For reference strategy, For the distribution of tasks, For the inference path within the group, For the clipping function, The KL distance between the old and new strategies is calculated using the following formula: 。 6. The multimodal large model inference method based on self-driven feedback and symbolic collaboration according to claim 5, characterized in that, In the reinforcement learning process of the model, symbolic constraint regularization is introduced to force the policy not to deviate from the symbolic rule base: in, This is the scoring threshold; when the symbol consistency score is below the threshold... When this happens, the probability corresponding to that path will be penalized. Minimizing this penalty during training makes the finally learned policy more inclined to choose actions with high symbolic consistency, resulting in the loss function: in, This is a hyperparameter.

7. The multimodal large model inference method based on self-driven feedback and symbolic collaboration according to claim 1, characterized in that, In S5, new candidate inference paths are generated in the environment as training samples, and symbolic consistency rewards are calculated based on the training samples. The mixed reward function is obtained by combining the symbolic consistency rewards, forming a self-driven feedback loop. A new set of candidate inference paths is generated, and the symbolic consistency rewards are calculated for the new candidate inference paths. The mixed reward is then used to generate the mixed reward, and the model policy is updated by maximizing the objective function.

8. A multimodal large-model inference system based on self-driven feedback and symbolic collaboration, characterized in that, Includes the following steps: The symbolic knowledge representation module is based on acquiring multimodal data, using knowledge graphs to structurally represent multimodal information, and combining large models to perform entity recognition and relation extraction, thereby constructing a unified knowledge graph. The symbolic knowledge reasoning module is used to define symbolic logic expressions, construct diverse symbolic logic rules, and calculate the symbolic consistency reward of the reasoning path; The reward signal design and fusion module is used to construct a reward mechanism that coordinates symbolic and human feedback, resulting in a hybrid reward function; The strategy optimization module is used to sample a set of outputs for a specific task during the interaction with the multimodal environment, take the mixed reward as input, and construct the objective function of the intra-group relative reward optimization strategy network in the multi-task scenario. The autonomous evolution loop module interacts with the environment to autonomously generate training samples, which are then fed into the symbolic knowledge processing and reasoning unit to obtain symbolic consistency rewards. These rewards are then injected into the reward signal design and fusion module and the policy optimization module, achieving self-driven feedback without the need for a large amount of manually labeled data.

9. A computer device, characterized in that, It includes a processor and a memory, the memory being used to store a computer-executable program, the processor reading part or all of the computer-executable program from the memory and executing it, and the processor executing part or all of the computed executable program is able to implement the multimodal large model inference method based on self-driven feedback and symbolic collaboration as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program that, when executed by a processor, enables the implementation of the multimodal large model inference method based on self-driven feedback and symbolic collaboration as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-modal large model self-reinforcement learning method based on thinking chain guidance

    CN119670878A

  • Logical reasoning-driven knowledge graph enhanced large model generation method

    CN120046711A

  • Path planning method, application and device based on knowledge and data combination

    CN117808180A

  • Intelligent agent training method and device, equipment and storage medium

    CN119494383A

  • Self-supervised neural symbol fusion interpretable AI inference method and system

    CN119623635A

Cited By

  • Method, device and equipment for constructing large knowledge extraction model in military field

    CN121882037A

  • Water industry knowledge enhanced large model training and inference method and system

    CN122287849A