Large language model system based on fuzzy reasoning and training method

By integrating deep learning methods with fuzzy reasoning capabilities, the ability of large language models to perform multi-step logical reasoning and fuzzy concept processing has been improved, solving the logical reasoning defects and uninterpretability problems in existing technologies, and achieving more reliable and interpretable output.

CN121562784APending Publication Date: 2026-02-24YANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511610350.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing large language models have shortcomings in handling multi-step logical reasoning, fuzzy concepts, and uninterpretability, which limits their application in high-risk domains and leads to overconfidence.

Method used

It employs a multimodal input and feature extraction module, a multimodal fuzzy lexical generation module, a Transformer backbone module with a fuzzy inference layer, a parallel output head module, and a unified loss function and end-to-end training module to integrate differentiable fuzzy inference capabilities, thereby improving the rigor, interpretability, and reliability of the model's logical reasoning.

Benefits of technology

It enhances the model's ability to handle multi-step reasoning and fuzzy scenarios, provides interpretable decision paths and reliability assessments, and improves performance and security in complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121562784A_ABST
    Figure CN121562784A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model system based on fuzzy reasoning and a training method. The system comprises a multi-modal input and feature extraction module, a multi-modal fuzzy lexical element generation module, a Transform trunk module with a fuzzy reasoning layer, a parallel output header module and a unified loss function and end-to-end training module. The method comprises the following steps: extracting a multi-modal input feature and generating an initial vector; aligning and fusing the multi-modal features, and generating dynamic fuzzy lexical element embedding; performing context modeling and logical reasoning through a Transform trunk with a fuzzy reasoning layer; generating a task result and a reliability score by using a parallel output header; performing end-to-end training by adopting a unified loss function; according to the method, the robustness in a scene with inaccurate information or fuzzy information is improved; therefore, a clear and symbolized interpretation path conforming to human intuition is provided for a final decision, and the credibility of the model is greatly enhanced; and important safety guarantee is provided for high-risk application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence, and in particular relates to a large language model system and training method based on fuzzy reasoning. Background Technology

[0002] Large Language Models (LLMs) have achieved tremendous success in natural language generation and understanding. They can generate fluent, coherent text and achieve state-of-the-art performance on various NLP tasks. However, the success of these models is mainly based on learning statistical patterns from massive amounts of data, and their reasoning process is more like a "conjecture" based on data associations than a rigorous logical deduction. However, existing LLMs suffer from the following technical problems: Logical reasoning flaws: When dealing with tasks requiring multi-step, rigorous logical reasoning (such as mathematical word problems and logic puzzles), LLMs often commit factual errors or logical fallacies, leading to broken reasoning chains or incorrect conclusions; Inability to handle ambiguity: The real world is full of vague and uncertain concepts, such as descriptions like "elevated body temperature," "severe cough," and "mild difficulty breathing" in medical diagnosis. The meanings of these concepts are continuous and context-dependent, and standard LLMs based on discrete lexical units struggle to accurately quantify and process this kind of fuzzy information, limiting their application in high-risk fields such as healthcare and finance. Lack of interpretability: The decision-making process of an LLM is a "black box," and how billions or even trillions of parameters interact to produce the final output is invisible to the user. This lack of interpretability makes it difficult to trust its decisions and hinders effective debugging and optimization when the model malfunctions. Overconfidence: LLMs often assign extremely high confidence levels (probabilities) to their erroneous outputs, essentially "confidently making mistakes," making it difficult for users to judge the reliability of their outputs, potentially leading to serious consequences in critical decision-making scenarios.

[0003] Therefore, there is an urgent need for a new model architecture that can endogenously incorporate fuzzy reasoning capabilities into the LLM computational structure, achieve end-to-end optimization, and provide interpretable and reliable outputs. Summary of the Invention

[0004] Purpose of the invention: The purpose of this invention is to provide a large language model system and training method based on fuzzy reasoning; to deeply integrate differentiable fuzzy reasoning capabilities into the architecture of LLM, fundamentally improving the model's ability to handle uncertain information, the rigor of logical reasoning, the interpretability of the decision-making process, and the reliability of the output results.

[0005] The large language model system based on fuzzy inference described in this invention includes a multimodal input and feature extraction module, a multimodal fuzzy lexical generation module, a Transformer backbone module with a fuzzy inference layer, a parallel output head module, and a unified loss function and end-to-end training module. The multimodal input and feature extraction module is used to process text and image multimodal inputs, generate initial feature vectors, and complete cross-modal alignment; The multimodal fuzzy lexical generation module is used to fuse multimodal features and generate dynamic fuzzy lexical embeddings; The Transformer backbone module with fuzzy inference layer is used to embed fuzzy inference into the Transformer structure and achieve logic correction and context understanding through differentiable fuzzy operators. The parallel output head module includes a task output head and a reliability prediction head, used to generate task results and reliability scores; The unified loss function and end-to-end training module are used to optimize model parameters through the integrated loss function, thereby achieving collaborative learning of language modeling, fuzzy reasoning, and reliability assessment.

[0006] Furthermore, the generation of the initial feature vector to complete cross-modal alignment includes feature alignment enhancement and fuzzy word embedding generation; wherein feature alignment enhancement is optimized through the following loss function: The CLIP model is used to encode text and image inputs separately to obtain text feature vectors. ; and image feature vectors ,in The feature dimension is used to solve for the loss function: ; in, It is cosine similarity. It's a temperature parameter. This refers to the batch size; CLIP pre-trained models map semantically similar text and images to similar feature spaces. Based on CLIP features, a learnable alignment transformation matrix is ​​used. Enhance the alignment of text and image features: ; in, and These are the adjusted text and image feature vectors, respectively. By minimizing alignment loss Optimize.

[0007] Furthermore, the multimodal fuzzy lexical generation module fuses multimodal features and generates dynamic fuzzy lexical embeddings, which, combined with alignment enhancement features, introduce joint fuzzy lexical embeddings. Obtained using a weighted average method: ; in, It is an adjustable weight, As an input signal, it is transmitted to the fuzzy inference layer to calculate the membership distribution. .

[0008] Furthermore, the multimodal fuzzy lexical generation module fuses features through an adaptive weighting mechanism, and the weighting calculation formula is as follows: ; in, All are learnable parameters, σ is the sigmoid function, and is used to preserve... The weights sum to 1, and then normalization is performed.

[0009] Furthermore, the fuzzy inference layer is defined as follows: Calculate input using dynamically generated parameters. Membership degree: ; If the input is a multidimensional feature vector Then, the membership degree is calculated for each dimension separately, and the overall membership degree is obtained by aggregation: ; Design a Membership Parameter Generator Network (MPGN) to dynamically predict the parameters of a Gaussian function in real time. and MPGN accepts contextual features or fuzzy term embeddings. As input, and output parameters: ; in, These are the learnable parameters of the network. The MPGN output layers correspond to... and .

[0010] Furthermore, the fuzzy inference layer automatically generates a fuzzy rule set through ensemble reinforcement learning, and the reward function is defined as follows: ; in, It's the accuracy of the task. It is the complexity of the rule set. It's the weight.

[0011] Furthermore, the fuzzy inference layer reduces computational complexity through sparsity rule activation, and the activation strength calculation formula is as follows: For each rule Calculate its activation strength based on the membership degree of the input features: ; in, Input variables In fuzzy sets Membership degree; Introduce a learnable threshold, activating only those with strengths higher than [a certain threshold]. The rules are used in subsequent reasoning: ; threshold Optimization via gradient descent allows the initial value to be set to the median or mean of the activation strengths; to encourage sparsity, a sparsity loss term is introduced, penalizing the number of activation rules. ; in, Here are the regularization coefficients, and the total loss function is: ; in, It's a mission loss.

[0012] Furthermore, the reliability prediction head in the parallel output head module is optimized using the following calibration loss function: ; in, This is the weighting function, which is set as follows: ; in, It is a hyperparameter greater than 1, which increases the protection against incorrect predictions and results in high reliability. The severity of the punishment.

[0013] The large language model training method based on fuzzy reasoning described in this invention includes the following steps: (1) Extract multimodal input features and generate initial vectors; (2) Align and fuse multimodal features to generate dynamic fuzzy word embeddings; (3) Context modeling and logical reasoning are performed using the Transformer backbone with a fuzzy reasoning layer; (4) Use parallel output headers to generate task results and reliability scores; (5) End-to-end training is performed using a unified loss function, which is defined as follows: ; in, It is the standard cross-entropy loss. It is a hyperparameter for balancing the loss term.

[0014] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0015] (1) Enhanced reasoning ability: Through the embedded fuzzy reasoning layer, the model has gained the ability to perform logical operations, and can handle tasks that require multi-step reasoning and quantitative judgment more rigorously, surpassing powerful baseline models in complex reasoning tasks such as mathematics, medicine and finance.

[0016] (2) Excellent uncertainty handling: The dynamic fuzzy lexical embedding mechanism enables the model to flexibly quantify and understand fuzzy concepts in natural language, improving robustness in scenarios with imprecise or fuzzy information;

[0017] (3) High interpretability: When the model makes a decision, it can be traced back to the specific rules activated in the fuzzy inference layer and their strength, thus providing a clear, symbolic and intuitive explanation path for the final decision, which greatly enhances the credibility of the model.

[0018] (4) Reliability assessment of calibration: Through parallel reliability prediction head and dedicated calibration loss, the model can not only assess the reliability of its own output, but also its score is calibrated and highly correlated with the actual accuracy, providing important security for high-risk applications. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the overall architecture of the large language model system in an embodiment of the present invention; Figure 2 This is a reliability calibration diagram of the large language model output in an embodiment of the present invention. Detailed Implementation

[0020] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] like Figure 1 As shown, the large language model system based on fuzzy reasoning of this invention mainly includes the following stages of data processing:

[0022] 1. Introduction of the Large Language Model (F-LLM) framework

[0023] The architecture of LLM is designed to integrate fuzzy logic and large language models in the deepest, end-to-end way. Our design approach is to embed fuzzy reasoning into the Transformer's streaming computation without adding any external post-processing modules. The architecture includes five major steps of data processing and transformation: learning, quantization, transformation, reasoning, and querying.

[0024] 1.1 Multimodal Input and Feature Extraction

[0025] This is the source of the data flow. The starting point of the Large Language Model (F-LLM) is to integrate information from multiple modalities to obtain a more complete information representation, thereby characterizing those concepts in the real world that are not so clear. The content of this paper is decomposed into a series of lexical IDs by a standard lexical analyzer. These lexical IDs are then converted into initial text feature vectors through an embedding layer. The original image is fed into the network via a pre-trained ViT visual encoder, which directly generates image feature vectors aligned with the text semantic space. .

[0026] 1.2 Cross-modal alignment and multimodal fuzzy lexical generation

[0027] This stage integrates features from different modalities and transforms them into an intermediate form that enables fuzzy reasoning. Furthermore, this stage involves feature alignment, feature fusion, and dynamic fuzzy lexical generation.

[0028] 1.3 Transformer backbone with fuzzy inference layer

[0029] This part serves as a core computational unit for sequence modeling and context understanding. F-LLM is based on multi-layered Transformer Blocks, each with a self-attention mechanism and a feedforward network, while the fuzzy inference layer is integrated into the Transformer backbone structure. The internal operations of fuzzy inference include rule activation, fuzzy inference, and information aggregation. Information aggregation combines the result of fuzzy inference with the original hidden state to generate a logically corrected and enhanced new hidden state, which is then passed to the next Transformer layer. Finally, all fuzzy operators are differentiable, allowing gradients to flow through the inference implementation portion of the fuzzy inference layer, achieving end-to-end joint optimization between inference and the language model.

[0030] 1.4 Parallel Output Head

[0031] The final hidden state of the Transformer backbone of the fuzzy inference layer is fed into two parallel output heads: the task output head and the reliability prediction head. The task output head is used to complete the main task. The reliability prediction head, in the generation task, predicts the next word based on the probability distribution of each word in the output vocabulary.

[0032] 1.5 Unified Loss Function and End-to-End Training

[0033] The loss function used for model training is a unified loss function that combines these two types of losses.

[0034] Mission loss Standard cross-entropy loss is used to supervise the accuracy of the task output head prediction.

[0035] Reliability calibration loss The model is penalized by weighted Brier scores. When a model gives a high reliability score but makes a wrong prediction, it is penalized so that the model can have a better "self-awareness".

[0036] The total loss function is defined as follows: ,right During backpropagation, gradients flow throughout the network and back from the output head to the feature extractor, thereby simultaneously optimizing the parameters of the three parts: language modeling, fuzzy inference, and reliability evaluation, truly achieving mutual integration.

[0037] 2. Breakthrough in key technologies of the F-LLM framework

[0038] 2.1 Generation Method of Dynamic Embedding of Multimodal Fuzzy Lexical Units

[0039] Within the F-LLM framework, fusing multimodal data to obtain fuzzy lexical embeddings is an important approach to achieving the fusion of fuzzy reasoning and language modeling. The deep learning multimodal fuzzy lexical embedding generation framework proposed here can be divided into four steps: extracting feature vectors corresponding to each modality using a pre-trained model, learning fuzzy membership functions using a neural network and generating fuzzy sets, combining multimodal features and fuzzy sets, and generating unified fuzzy lexical embeddings.

[0040] 2.1.1 Data alignment issues

[0041] The semantics of text and images are not entirely consistent, resulting in noise in the fused features, which affects the quality of fuzzy lexical embeddings. This paper uses a high-capacity cross-modal pre-trained model to achieve feature alignment, fully leveraging its semantic alignment capabilities to extract shared representation spaces from text and images, and constructing fuzzy lexical embeddings based on these spaces.

[0042] Step 1: Cross-modal feature extraction

[0043] The CLIP model is used to encode text and image inputs separately to obtain text feature vectors. and image feature vectors ,in Given the feature dimension, the loss function is:

[0044] ;

[0045] in, It is cosine similarity. It's a temperature parameter. It refers to the batch size; CLIP pre-trained models map semantically similar texts and images to similar feature spaces.

[0046] Step 2: Feature Alignment Enhancement

[0047] Building upon CLIP features, a learnable alignment transformation matrix is ​​further employed. Enhance the alignment of text and image features:

[0048] ;

[0049] in, By minimizing alignment loss Optimize.

[0050] Step 3: Fuzzy Lexical Embedding Generation. Combine aligned features to generate joint fuzzy lexical embeddings. The weighted average yields the following:

[0051] ;

[0052] in, It is an adjustable weight. Next, we will... As an input signal, it is transmitted to the fuzzy inference layer to calculate the membership distribution. .

[0053] 2.1.2 Adaptive Weight Calculation Let the text features be Image features are Design a gating network to dynamically compute text.

[0054] Weights of image modalities:

[0055] ;

[0056] in, All are learnable parameters, σ is the sigmoid function, ensuring To ensure the weight sum is 1, the normalization process is as follows:

[0057] Feature fusion using adaptive weights:

[0058] ;

[0059] After fusion Input the fuzzy inference layer to generate the membership degree distribution. The parameters of the gating network can be obtained through parameter... Perform end-to-end optimization, while introducing regularization terms to prevent [further issues].

[0060] Prevent a certain modality from having too low a weight:

[0061] ;

[0062] in, It is a regularization coefficient that ensures the weights of the two modes are not excessively unbalanced. The adaptive weighting mechanism dynamically adjusts the importance of each mode to the fusion result.

[0063] 2.2 Fuzzy Inference Layer

[0064] Based on the F-LLM framework, research is conducted on three key research directions for fuzzy inference layers: deep learning-based membership function design, automatic generation of fuzzy rule sets through set reinforcement learning, and reducing computational complexity through sparsification rule activation.

[0065] Step 1: Choosing the form of the membership function

[0066] To ensure differentiability and computational efficiency, the Gaussian function is chosen as the basic form of the membership function because its gradient is simple to calculate and has good smoothness. The Gaussian membership function is defined as follows:

[0067] ;

[0068] in, It is the center point. It is the width parameter. These are the input feature values.

[0069] Step 2: Input Feature Extraction

[0070] In the F-LLM model, the input features can be a part of the fuzzy lexical embedding or other contextual features. The fused feature vector is obtained using the multimodal fuzzy lexical embedding generation method in the first part.

[0071] ;

[0072] in, and Alignment features of text and image modalities, It is adaptive weights. Each dimension or through the reduced features as Input the membership function.

[0073] Step 3: Dynamic Parameter Generation Network

[0074] Based on the aforementioned approach, a membership parameter generator (MPG) network was designed, which can be used to dynamically predict the parameters of Gaussian functions in real time. and MPGN accepts contextual features or fuzzy term embeddings. As input, and output parameters:

[0075] ;

[0076] in, These are the learnable parameters of the network. The MPGN output layers correspond to... and To ensure ,right The output applies soft positive constraints:

[0077] ;

[0078] Step 4: Membership Calculation

[0079] Calculate input using dynamically generated parameters. Membership degree:

[0080] ;

[0081] If the input is a multidimensional feature vector Then, the membership degree is calculated for each dimension separately, and the overall membership degree is obtained by aggregation:

[0082] ;

[0083] Step 5: Gradient Calculation and Optimization

[0084] Since both the Gaussian function and MPGN are differentiable, the gradient of the membership function can be calculated through backpropagation; assuming the loss function is... ,right and The gradient is:

[0085] ;

[0086] ;

[0087] These gradients are further passed to the parameters of MPGN via the chain rule. This enables end-to-end training.

[0088] Step 6: Task Adaptive Optimization

[0089] During training, the parameters of the membership function are optimized based on the task loss, and a regularization term is added to avoid... It becomes too small or too big.

[0090] ;

[0091] in, It is the regularization coefficient.

[0092] We use deep learning methods to design membership functions and use MPGN to dynamically generate the parameters of Gaussian membership functions, so that the membership functions can be adaptively adjusted according to context, task and other conditions, overcoming the limitations of the traditional manual design of membership functions.

[0093] 2.3 Ensemble Reinforcement Learning for Automatic Generation of Fuzzy Rule Sets

[0094] Modeling the rule generation problem as a reinforcement learning problem, where the state space... It consists of the current context features and the already generated rule set. Action space. The reward function is used to generate new rules. The rule set is determined based on task performance and its brevity. Here, action 'a' corresponds to a newly generated rule, and state 's' represents the currently generated rule set and context features. The state-action-state reward function can then be defined as:

[0095] ;

[0096] in, It's the accuracy of the task. It is the complexity of the rule set. It's the weight.

[0097] 2.4 Reducing computational complexity through sparsity rule activation

[0098] The computational complexity of fuzzy inference systems primarily stems from the rule activation and fuzzy aggregation steps, especially when the number of rules is large. Sparse rule activation significantly reduces computational overhead by decreasing the number of activated rules while maintaining inference performance. The detailed steps and mathematical derivations are as follows:

[0099] Step 1: Calculation of rule activation strength

[0100] For each rule Calculate its activation strength based on the membership degree of the input features:

[0101] ;

[0102] in, Input variables In fuzzy sets Membership degree.

[0103] Step 2: Sparsification of activation threshold

[0104] Introduce a learnable threshold Only activation intensity higher than The rules are used in subsequent reasoning:

[0105] ;

[0106] threshold This can be optimized using gradient descent, with the initial value set to the median or mean of the activation strength.

[0107] Step 3: Sparsification and Loss Regularization

[0108] To encourage sparsity, a sparsity loss term is introduced to penalize the number of activation rules:

[0109] ;

[0110] in, These are the regularization coefficients; the total loss function is:

[0111] ;

[0112] in, It is the task loss (such as classification or generation loss).

[0113] Robustness and reliability calibration: To examine its robustness, noise was first added to the FuzzyQuant-QA test set, changing "85" to "eighty-five" and "approximately 83". Test results showed that F-LLM performance decreased by 1.5% and GPT-3.5 performance decreased by 5.8%, as shown in Table 1. This indicates that the F-LLM dynamic membership function has strong adaptability to small changes in input. The expected calibration error value was then calculated, as shown in Table 2, and a reliability plot was drawn. Figure 2 As shown.

[0114] Table 1: Performance on major benchmark datasets (accuracy %)

[0115] Model GSM8K LogiQA MDB Financial PhraseBank Fuzzy Quant-QA LLaMA-2-13B 48.2 41.5 65.7 82.1 61.3 GPT-3.5-Turbo 71.8 55.3 74.2 88.5 89.6 LLaMA-2-13B + CoT 60.1 48.9 69.1 84.6 75.4 LLaMA-2-13B+ Self-Consistency 65.5 52.1 71.8 85.9 80.2 LLM +Post-hoc Fuzzy 66.2 53.0 72.5 86.3 82.5 LLM (Ours) 73.5 61.8 80.4 91.2 96.8

[0116] Table 2: Reliability Calibration Error (ECE)

[0117] Model MDB LogiQA LLaMA-2-13B + CoT 15.8 18.2 GPT-3.5-Turbo 11.3 14.5 F-LLM (Ours) 2.7 4.1

[0118] Ablation studies were conducted on the MDB dataset to verify the necessity of each component. The F-LLM results demonstrate the role of each module, with multimodal input being crucial to the results, as shown in Table 3, achieving state-of-the-art performance on tasks like MDB. The fuzzy inference layer significantly contributes to performance improvement and interpretability; removing the fuzzy inference layer results in a substantial performance decrease. Dynamic membership functions are more suitable for the current scenario than static membership functions, improving performance under these conditions. Reliability calibration has little impact on accuracy but a significant impact on ECE, improving the reliability of the model output.

[0119] Table 3: Results of F-LLM ablation studies (MDB, accuracy %)

[0120] Model variants accuracy ECE (%) F-LLM (Full) 80.4 2.7 w / o Multi-modal Input 73.1 (-7.3) 3.5 (+0.8) w / o Fuzzy Reasoning Layer 70.5 (-9.9) 9.8 (+7.1) w / Static Membership Function 76.8 (-3.6) 5.2 (+2.5) w / o Reliability Calibration 80.1 (-0.3) 12.4 (+9.7)

[0121] like Figure 2 As shown, this illustrates the relationship between the reliability (confidence level) of the model predictions and the actual accuracy.

[0122] Ideally, all points should fall on the diagonal. The expected calibration error of F-LLM is much lower than that of the baseline model, indicating that its reliability score has been well calibrated and has practical reference value.

[0123] In summary, this invention, through an innovative deep fusion architecture, successfully combines the reasoning ability and interpretability of fuzzy logic with the powerful language capabilities of large language models, while introducing a self-consistent reliability calibration mechanism, providing a new technical path for building more powerful, safer, and more trustworthy AI systems.

[0124] To address these issues, some research attempts to combine logic systems with LLM (Liquidity Management Model). Fuzzy logic, as a powerful mathematical tool for handling uncertainty and ambiguity, can simulate human thinking in situations with incomplete or ambiguous information. However, current research mostly employs a "shallow fusion" approach, such as using fuzzy logic as a post-processor or external controller for LLM outputs. In this approach, the LLM itself does not learn fuzzy reasoning capabilities, and the two cannot achieve deep collaborative learning and joint optimization, resulting in limited performance improvements.

Claims

1. A large language model system based on fuzzy reasoning, characterized in that, It includes a multimodal input and feature extraction module, a multimodal fuzzy lexical generation module, a Transformer backbone module with a fuzzy inference layer, a parallel output head module, and a unified loss function and end-to-end training module; The multimodal input and feature extraction module is used to process text and image multimodal inputs, generate initial feature vectors, and complete cross-modal alignment; The multimodal fuzzy lexical generation module is used to fuse multimodal features and generate dynamic fuzzy lexical embeddings; The Transformer backbone module with fuzzy inference layer is used to embed fuzzy inference into the Transformer structure and achieve logic correction and context understanding through differentiable fuzzy operators. The parallel output head module includes a task output head and a reliability prediction head, used to generate task results and reliability scores; The unified loss function and end-to-end training module are used to optimize model parameters through the integrated loss function, thereby achieving collaborative learning of language modeling, fuzzy reasoning, and reliability assessment.

2. The large language model system based on fuzzy reasoning according to claim 1, characterized in that, The generation of the initial feature vector and the completion of cross-modal alignment include feature alignment enhancement and fuzzy word embedding generation; wherein feature alignment enhancement is optimized through the following loss function: The CLIP model is used to encode text and image inputs separately to obtain text feature vectors. ; and image feature vectors ,in The feature dimension is used to solve for the loss function: ; in, It is cosine similarity. It's a temperature parameter. This refers to the batch size; CLIP pre-trained models map semantically similar text and images to similar feature spaces. Based on CLIP features, a learnable alignment transformation matrix is ​​used. Enhance the alignment of text and image features: ; in, and These are the adjusted text and image feature vectors, respectively. By minimizing alignment loss Optimize.

3. The large language model system based on fuzzy reasoning according to claim 1, characterized in that, The multimodal fuzzy lexical generation module fuses multimodal features and generates dynamic fuzzy lexical embeddings. Combined with alignment enhancement features, it introduces joint fuzzy lexical embeddings. Obtained using a weighted average method: ; in, It is an adjustable weight, As an input signal, it is transmitted to the fuzzy inference layer to calculate the membership distribution. .

4. The large language model system based on fuzzy reasoning according to claim 1, characterized in that, The multimodal fuzzy lexical generation module fuses features through an adaptive weighting mechanism, and the weighting calculation formula is as follows: ; in, All are learnable parameters, σ is the sigmoid function, and is used to preserve... The weights sum to 1, and then normalization is performed.

5. The large language model system based on fuzzy reasoning according to claim 1, characterized in that, The fuzzy inference layer is defined as follows: Calculate input using dynamically generated parameters. Membership degree: ; If the input is a multidimensional feature vector Then, the membership degree is calculated for each dimension separately, and the overall membership degree is obtained by aggregation: ; Design a Membership Parameter Generator Network (MPGN) to dynamically predict the parameters of a Gaussian function in real time. and MPGN accepts contextual features or fuzzy term embeddings. As input, and output parameters: ; in, These are the learnable parameters of the network. The MPGN output layers correspond to... and .

6. The large language model system based on fuzzy reasoning according to claim 1, characterized in that, The fuzzy inference layer automatically generates a fuzzy rule set through ensemble reinforcement learning, and the reward function is defined as follows: ; in, It's the accuracy of the task. It is the complexity of the rule set. It's the weight.

7. The large language model system based on fuzzy reasoning according to claim 1, characterized in that, The fuzzy inference layer reduces computational complexity through sparsity rule activation. The activation strength calculation formula is as follows: For each rule Calculate its activation strength based on the membership degree of the input features: ; in, Input variables In fuzzy sets Membership degree; Introduce a learnable threshold, activating only those with strengths higher than [a certain threshold]. The rules are used in subsequent reasoning: ; threshold Optimization via gradient descent allows the initial value to be set to the median or mean of the activation strengths; to encourage sparsity, a sparsity loss term is introduced, penalizing the number of activation rules. ; in, Here are the regularization coefficients, and the total loss function is: ; in, It's a mission loss.

8. The large language model system based on fuzzy reasoning according to claim 1, characterized in that, The reliability prediction head in the parallel output head module is optimized using the following calibration loss function: ; in, This is the weighting function, which is set as follows: ; in, It is a hyperparameter greater than 1, which increases the protection against incorrect predictions and results in high reliability. The severity of the punishment.

9. A method for training a large language model based on fuzzy reasoning, characterized in that, Includes the following steps: (1) Extract multimodal input features and generate initial vectors; (2) Align and fuse multimodal features to generate dynamic fuzzy word embeddings; (3) Context modeling and logical reasoning are performed using the Transformer backbone with a fuzzy reasoning layer; (4) Use parallel output headers to generate task results and reliability scores; (5) End-to-end training is performed using a unified loss function, which is defined as follows: ; in, It is the standard cross-entropy loss. It is a hyperparameter for balancing the loss term.