Asset evaluation method based on multi-modal large model
Through the asset valuation method of multimodal large models, the problems of unified embedding space and inter-modal distribution heterogeneity in multimodal data fusion in financial valuation scenarios are solved, deep representation and semantic alignment of images, text and structured data are achieved, and the synergy and reliability of the model are improved.
Patent Information
- Application Number
- CN202510793136.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing multimodal processing technologies lack explicit modeling of modal difference structures in financial assessment scenarios, resulting in confusion in the semantic hierarchy, failure to build a unified embedding space, incomparability between modalities, difficulty in structural alignment, fragmented training paths, inconsistent parameter optimization directions, and difficulty in coordinated convergence.
Through an asset evaluation method based on a multimodal large model, including the preprocessing of images, texts and structured data, high-dimensional feature vectors are extracted using image encoders, text encoders and structured data encoders, and mapped to the same common embedding space by constructing a mapper network. The optimal transfer method is used to measure the embedding distribution differences and minimize heterogeneity, and the training objectives are optimized by minimizing the loss function.
It achieves deep representation of various types of heterogeneous data, unified expression between modalities, and spatial comparability, improves the model's collaboration and deployment reliability, and solves the problems of insufficient modal feature expression, missing deep semantic information, unbalanced embedding, and unstable training.
Smart Images

Figure CN120671082A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence and financial technology, and specifically to an asset valuation method based on a multimodal large model. Background Art
[0002] In financial scenarios like asset management, corporate assessment, and risk monitoring, comprehensive judgments often require simultaneous reference to company image information, textual materials, and structured indicators. Different modalities offer complementary information and distinct perspectives, and the ability to rationally integrate and process them can significantly improve decision-making accuracy and efficiency.
[0003] Among existing multimodal processing technologies, some methods have achieved basic integration of image, text, and structured data. Most employ modality splicing, shared encoding, or simple weighted fusion strategies, achieving promising results in tasks such as classification and retrieval. These methods offer clear design principles and low engineering implementation costs, making them suitable for lightweight scenarios with similar modal feature distributions and high semantic consistency. They also offer advantages in structural simplicity and ease of deployment, enabling rapid deployment in some static evaluation systems.
[0004] However, the above-mentioned existing technologies have several key shortcomings when dealing with multimodal data in real financial assessment scenarios. First, they generally lack explicit modeling of modal difference structures, resulting in confusion in the semantic hierarchy. Second, they fail to build a unified embedding space, making different modalities incomparable and structural alignment difficult. Third, they lack a mechanism to resolve heterogeneity at the distribution level and rely solely on loss functions to compress distances, resulting in superficial and unstable effects. Finally, due to the fragmentation of training paths, each module often "does its own thing", with inconsistent parameter optimization directions, making it difficult for the overall model to converge collaboratively. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides an asset valuation method based on a multimodal large model, which solves the problems in the existing technology of the lack of a unified embedding space in the multimodal data fusion process, the difficulty in eliminating the heterogeneity of distribution between modalities, and the inability to coordinately optimize feature extraction and alignment.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: an asset evaluation method based on a multimodal large model, comprising the following steps: Preprocess the image data, text data, and structured data of the enterprise to be assessed to generate standardized multimodal input data; Based on the standardized multimodal input data, the image data, the text data, and the structured data are processed using an image encoder, a text encoder, and a structured data encoder, respectively, to extract a high-dimensional feature vector of each modality; By constructing a mapper network, the obtained high-dimensional feature vectors of each modality are mapped to the same common embedding space; The optimal transfer method is used to measure the embedding distribution difference between each modality and further remove the heterogeneity between modalities by minimizing the difference.
[0007] Preferably, the pretreatment includes: Normalize the image data and scale the pixel values of the image to a specified range; The text data is segmented and the word vectorization method is used to convert the text into a fixed-dimensional word vector representation; the structured data is normalized and the numerical data is mapped to a unified scale range.
[0008] Preferably, the multimodal input data includes: Image data, which represents the appearance of corporate assets, including product photos, factory facilities, and office environment images; Text data, which represents written information related to corporate assets, including corporate announcements, financial reports, patent documents, and industry analysis reports; Structured data represents the numerical information of corporate assets, including financial statements, balance sheets, and fixed asset lists.
[0009] Preferably, the high-dimensional feature vector includes: For image data, the high-dimensional feature vector is an image feature vector extracted by a convolutional neural network; For text data, the high-dimensional feature vector is a text representation extracted based on a pre-trained language model; For structured data, the high-dimensional feature vector is a numerical feature extracted by a multi-layer perceptron processing network.
[0010] Preferably, the mapper network comprises: The mapper network is trained by minimizing a loss function consisting of a weighted sum of alignment loss and valuation prediction error, as follows: in, is the modality alignment loss, is the asset valuation regression error, is the total loss function used to optimize the training objective of the entire mapper network, and λ is the loss weighting coefficient.
[0011] Preferably, the public embedding space includes: A unified high-dimensional vector space, where each modality can be mapped to this space and have similar semantic representations after being processed by the corresponding encoder; The common embedding space can align different modalities so that the feature vectors of images, text and structured data have consistent geometric structures and similar semantic expressions within the space; The structure of the common embedding space is dynamically adjusted through a learning process to adapt to the characteristics of different modalities and reduce the representation differences between the modalities.
[0012] Preferably, the optimal transmission method includes: The optimal transmission distance is used to measure the difference in embedding distribution between different modalities. The optimal transmission distance is defined as: where μ and ν are the embedding distributions of images, text, and structured data, respectively; d(x, y) is the distance metric in the embedding space; Π(μ, ν) is the total coupling that maximizes the alignment of different modal embeddings in the common space by minimizing this distance; γ∈Π(μ, ν) is the joint distribution that connects the distributions μ and ν, satisfying that the marginal distributions are μ and ν, respectively, and X×Y is the product space of the input space; is the p-order mean of the overall distance to ensure consistent units of the results; p is the power exponent parameter; inf is the minimum value, which means finding the minimum transmission cost among all possible couplings; dγ(x,y) represents the differential element integrated at (x,y) according to the joint distribution γ.
[0013] Preferably, the embedding distribution difference includes: The differences in the distribution generated by each modality in the embedding space are specifically manifested in the different geometric forms of the embedded features of images, text, and structured data in high-dimensional space; Image data usually presents a dense pixel structure, text data is represented by word embedding, and structured data is expressed by numerical features; This difference needs to be measured and minimized through an optimal transfer method so that features from different modalities can share similar semantic representations in a common embedding space.
[0014] Preferably, minimizing the difference comprises: By designing a joint optimization objective function, the alignment error between modalities and the asset valuation error are optimized; In the joint optimization objective, the modality alignment loss controls the distribution differences of different modality embedding vectors in the common embedding space through regularization methods.
[0015] Preferably, the heterogeneity between modalities includes: Different modal data are represented in different ways: image data is represented as a pixel matrix, text data is represented as a word vector, and structured data is represented as a numerical table data. Semantic differences between modalities and differences in representation between text and images at the semantic level lead to deviations in asset valuations due to the different expressive capabilities of different modalities.
[0016] The present invention provides an asset evaluation method based on a multimodal large model. It has the following beneficial effects: 1. This invention utilizes a multimodal encoder to extract high-dimensional semantic vectors from images, text, and structured data, achieving a more comprehensive, deep representation of heterogeneous data and semantic layer extraction. Compared to existing approaches that rely solely on shallow features or directly fuse them after preprocessing, this approach addresses the issues of insufficient expressiveness of modal features and the lack of deep semantic information.
[0017] 2. This invention utilizes a technical approach to construct modality-specific mappers and align them to a unified common embedding space, achieving unified representation and spatial comparability across modalities. Compared to existing multimodal fusion approaches that lack a unified representation framework and directly merge them, resulting in distributional dislocation, this approach addresses the issues of inability to directly compare modalities and structural distortion.
[0018] 3. This invention employs an optimal transmission method to measure and regularize the embedding distribution of each modality, achieving the goal of reducing modal heterogeneity at the distribution level. Compared to existing approaches that rely solely on simple alignment loss or manual feature normalization, this approach solves the challenges of unbalanced embeddings and large distribution offsets before modal fusion.
[0019] 4. This paper employs an end-to-end collaborative training mechanism for the embedder, mapper, and optimal transport modules, enabling unified optimization of feature learning and modality alignment. Compared to the existing approach of independent module training and disconnected parameter optimization, this approach addresses the issues of weak overall model coordination and unstable training, while also improving model reliability in actual deployment. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0022] Please see the attached Figure 1, an embodiment of the present invention provides an asset evaluation method based on a multimodal large model, comprising the following steps: S1, preprocessing image data, text data and structured data of the enterprise to be evaluated to generate standardized multimodal input data; Data preprocessing is a basic step in constructing standardized multimodal input data, which directly determines the effectiveness of the subsequent feature extraction module and the accuracy of the multimodal embedding alignment mechanism.
[0023] Typically, the raw data of the company being evaluated often suffers from inconsistent modalities, inconsistent scales, high noise levels, and unclear semantic distribution. Directly inputting this data into the model can lead to reduced feature extraction performance, impacting the accuracy of the final asset valuation. Therefore, in this invention, images, text, and structured data must first be standardized and preprocessed to ensure they meet the requirements for unified input format and provide high-quality input for the subsequent encoder network.
[0024] As an option, this preprocessing process can be performed immediately after data collection is completed, or it can be deployed in the enterprise data lake system in combination with the data storage process to facilitate long-term call and reuse.
[0025] For image data preprocessing, we first resample the original image using bilinear interpolation to a uniform image size of 224×224 pixels to meet the input requirements of mainstream convolutional neural network architectures. Next, we normalize the pixel values of the RGB channels of the image: Among them, I norm is the normalized image pixel value, I is the original image pixel value, μ is the channel mean vector of the training set image, and σ is the corresponding standard deviation vector. This normalization operation can effectively improve the convergence speed and stability of the feature extractor.
[0026] Specifically, in one possible implementation, after the image input is normalized, a ResNet-50 network can be used as the initial feature extractor. If industrial scenarios require this, a lightweight MobileNet architecture can be used instead to adapt to edge deployments.
[0027] In this embodiment, the preprocessing step for text data involves first performing basic word segmentation on the enterprise text information using a Chinese or English word segmentation tool. After word segmentation, the text sequence is mapped into a sequence of word vectors. The word embedding model used may include GloVe, Word2Vec, or more preferably, the BERT embedding layer based on the Transformer.
[0028] Then, the input tensor is generated by the following formula: T = Embed(W) + Pos(W); Where T represents the final input tensor, Embed(W) is the word embedding matrix, and Pos(W) is the corresponding positional encoding. The two are added together to form the Transformer input layer tensor. This approach helps capture semantic and sequential information in long texts.
[0029] Generally speaking, for typical text materials such as corporate announcements, financial summaries, and asset appraisal reports, the maximum input length can be set to 512 words, the truncation strategy is post-truncation, and the padding strategy is post-filling.
[0030] In this embodiment, the preprocessing of structured data includes but is not limited to numerical standardization, missing value filling, categorical variable encoding, etc.
[0031] Specifically, for corporate financial indicators (such as total assets, net fixed assets, net profit margin, etc.), the minimum-maximum normalization method is used to convert the original value x i Mapping to the interval 0,1: Among them, x i is the original value, is the normalized value of the index corresponding to the i-th sample, x min with x max are the minimum and maximum values of the indicator in the training set, respectively. This method is applicable to scenarios where most indicators have physical upper and lower limits or non-negative distributions.
[0032] Missing values in structured data can be imputed using mean values or, in more complex implementations, a KNN imputation strategy based on neighboring samples. Categorical fields (such as industry and asset classification level) can be encoded using one-hot encoding or embedded vector mapping (e.g., with a dimension of 16 or 32), depending on the downstream model capacity.
[0033] In some embodiments, the structured data is preprocessed and concatenated into a unified input feature vector, the dimension of which can change dynamically according to the number of selected fields, and the output dimension is usually fixed to 256 after processing by a multi-layer perceptron (MLP).
[0034] S2. Based on the standardized multimodal input data, the image data, text data, and structured data are processed using an image encoder, a text encoder, and a structured data encoder, respectively, to extract high-dimensional feature vectors of each modality; After completing the standardized preprocessing operation, in order to further achieve feature alignment and embedding fusion of cross-modal data, the standardized input data of different modalities needs to be input into a specially constructed encoder network for processing, so as to extract expressive high-dimensional feature vectors. This step is a direct follow-up to the aforementioned "standardized multimodal input data" and forms a bridge between the common embedding space construction link, assuming the core function of semantic representation of modal information. Therefore, each modality encoder needs to have a targeted structural design and parameter adjustment mechanism to adapt to the representation characteristics of images, text and structured data, and to ensure the stable operation of subsequent alignment and valuation modules.
[0035] In this embodiment, the image encoder is primarily used to extract features from the standardized image tensor. Typically, the image encoder employs a deep convolutional neural network (CNN), whose input is an RGB image tensor of size [3×224×224]. Alternatively, a ResNet-50 network structure can be used to extract high-level semantic features, with its output being a one-dimensional vector of dimension 2048.
[0036] In one possible implementation, the image encoder includes multiple residual blocks, each of which contains two 3×3 convolutional layers, a batch normalization layer (BatchNorm), and a ReLU activation function. Residual connections are used to enhance the network's gradient flow and prevent degradation in deep networks. The network structure can be formally described as follows: F img =φ CNN (I norm ); Among them, I norm is the normalized image pixel value, φ CNN Represents the convolutional neural network function mapping, F img is the extracted high-dimensional feature vector of the image.
[0037] Specifically, the training of the image encoder can be based on supervised learning or self-supervision. Auxiliary losses, such as TripletLoss or ContrastiveLoss, can be introduced during the training phase to enhance the discriminability of image features in the embedding space.
[0038] In this embodiment, the text encoder is used to process the word vector sequence obtained after word segmentation and embedding. The text data input format is [n×d t ], where n is the number of tokens, d t =512 is the word vector dimension. The text encoder is built based on the Transformer architecture and can use pre-trained models such as BERT or RoBERTa as the backbone network.
[0039] Generally speaking, the input of Transformer includes the sum of the embedding vector and the positional encoding. Its core structure is composed of alternating stacking of multi-head self-attention mechanism (Multi-headSelfAttention) and feed-forward network (Feed-forwardLayer).
[0040] In one possible implementation, the encoding process can be expressed as follows: F txt =φ BERT (T); Among them, φ BERT (T) represents the mapping function of the BERT model, F txt The output text high-dimensional semantic vector usually takes the CLS bit as the representation of the entire text.
[0041] Specifically, to improve embedding accuracy, the text encoder can be fine-tuned in the field of asset valuation. The training corpus includes annual reports, announcements, industry analysis, etc., to enhance the model's ability to understand professional terms and logical structures.
[0042] In this embodiment, the input of the structured data encoder is a s A numeric vector of d s Represents the number of normalized feature dimensions, such as 25 financial indicators. Generally, structured data encoders use a multi-layer perceptron (MLP) structure to extract deep numerical features through a combination of layer-by-layer linear transformations and nonlinear activation functions.
[0043] In one possible implementation, the structured encoder consists of three fully connected layers, each followed by a ReLU activation function, with a final output dimension of 256. The specific calculation process is as follows: F str =φ MLP (S norm )=W3(σ(W2(σ(W1S norm +b1))+b2))+b3; Among them, S norm is the normalized structured input vector; σ is the activation function; F str is the final structured data feature vector; φ MLP is a multi-layer perceptron network structure; W1 is the weight matrix of the first layer, used to map the input dimension to the hidden layer; W2 is the weight matrix of the second layer, used for further feature transformation; W3 is the weight matrix of the third layer, used to output the final vector; b1 is the bias vector of the first layer; b2 is the bias vector of the second layer; b3 is the bias vector of the third layer.
[0044] As an option, to prevent overfitting of numerical data, a Dropout mechanism can be introduced between encoder layers, and the dropout probability can be set between 0.3 and 0.5; at the same time, a batch normalization strategy is supported to stabilize the training process.
[0045] In some embodiments, to improve the consistency and comparability of the encoder output, each modal feature output vector may be subjected to L2 normalization: Among them, F m is the output feature of any modality, ||.||2 is the L2 norm, and the normalized feature It has unit length, which facilitates the alignment of subsequent similarity metrics with embeddings.
[0046] It should be emphasized that the parameters of each modal encoder can be jointly trained through the backpropagation algorithm. The training objectives include downstream asset valuation error and modal alignment loss. The details will be expanded in the subsequent module description.
[0047] S3, by building a mapper network, mapping the obtained high-dimensional feature vectors of each modality into the same common embedding space; After completing the feature extraction of each modality, the high-dimensional vectors of the obtained images, texts and structured data still have problems such as inconsistent semantic distribution and inconsistent geometric structure. It is often difficult to obtain ideal valuation results by directly performing modal fusion. Therefore, the present invention embeds the high-dimensional features of different modalities into the same common embedding space by constructing a unified mapper network to achieve semantic alignment and geometric consistency between modalities. This step serves as an intermediate bridge between feature extraction and modal fusion and optimal transmission processing. Its core goal is to standardize the representation semantics of the three types of modalities so that they can be conformally embedded in a shared space, thereby meeting the computational requirements of subsequent alignment and valuation.
[0048] In this embodiment, the mapper network consists of three parallel structural modules, corresponding to the embedding vectors of the image modality, text modality, and structured modality respectively. The input of each branch network is the high-dimensional feature representation of the modality, and the output is a unified feature vector mapped to the common space. In general, the dimension of the common embedding space is set to d c , which can be set to 256, 512, or 1024, depending on the representation capacity required by the downstream task.
[0049] In one possible implementation, the mapper network adopts a linear projection plus normalization structure, and the overall mapping relationship is as follows: Among them, F m is the output feature of any modality, W m is the weight matrix of the m-th mapper, bm is the bias vector, ψ m represents the mapper network function, Z m represents the public space embedding vector after mapping, and ||.||2 represents the L2 norm.
[0050] Specifically, the common embedding space has the ability to maintain semantic similarity and structural consistency. Its structure is dynamically adjusted during training through the alignment loss function, so that semantically similar data samples are closer in the space, while samples with large semantic differences are separated from each other.
[0051] As an option, the mapper network can further introduce nonlinear transformation modules, such as ReLU activation function and Dropout regularization mechanism, to improve nonlinear modeling capabilities and alleviate the risk of overfitting. Modeling can be performed in the following forms: Among them, W 1m is the first layer linear transformation parameter, W 2m is the second layer linear transformation parameter, F m is the output feature of any mode, Z m represents the public space embedding vector after mapping, ReLU is the activation function, represents the rectified linear unit, b 1m is the bias vector of the first layer, b 2m is the bias vector of the second layer.
[0052] In some embodiments, to ensure consistent projections of different modalities in the common space, a modality-sharing subspace constraint can be introduced. In this design, the image, text, and structured mapper networks share weights in some parameter layers, ensuring geometric compatibility among the three modalities in the common space.
[0053] In order to make the common embedding space learning process more stable, this embodiment uses a minimization loss function for training. The loss function includes the weighted sum of alignment loss and valuation prediction error, and the formula is: in, is the modality alignment loss, is the asset valuation regression error, is the total loss function used to optimize the training objective of the entire mapper network, and λ is the loss weighting coefficient.
[0054] In a specific design, the modality alignment loss can be modeled using contrastive learning, such as the InfoNCE loss: in, is the modality alignment loss; sin is the cosine similarity; τ is the temperature hyperparameter; used to adjust the degree of distribution smoothness; Embedding representation of images and texts for the same enterprise; log is the natural logarithm function; exp is the exponential function; used to construct a softmax distribution; is the text embedding of the jth sample; N is the total number of samples.
[0055] In some extended implementations, the public embedding space can be used as an explicit index space to provide fast retrieval capabilities for subsequent evaluation systems. For example, by constructing an inter-modal embedding distance matrix, it can be used to automatically identify anomalous corporate asset representations or to build a multimodal corporate similarity database.
[0056] In this embodiment, the common embedding space is not a static structure, and its geometric features will be continuously updated during the training iteration process. The parameters of the embedder network are iteratively optimized through optimization algorithms such as gradient descent.
[0057] Generally speaking, the embedding space is ultimately expressed as a continuous high-dimensional semantic space, in which enterprise samples of different modalities can present a clustering structure, and modal fusion or valuation prediction can be performed directly afterwards.
[0058] S4. Use the optimal transfer method to measure the embedding distribution difference between each modality and further remove the heterogeneity between modalities by minimizing the difference; After embedding multimodal features into a unified common space, although initial unification has been achieved at the geometric and semantic levels, significant differences in embedding distributions may still exist between different modalities. To further eliminate this heterogeneity and improve the consistency and robustness of multimodal information fusion, this paper introduces an optimal transmission method to measure the distance between the embedding distributions of each modality. Based on this minimization strategy, it forces the embedding distributions to converge towards consistency, thereby achieving consistency in the semantic distributions between modalities in the embedding space.
[0059] Structurally, this module follows the aforementioned mapper network output stage and is connected to the modal fusion and valuation regression stages. It is the alignment optimization core of the entire valuation system.
[0060] The optimal transmission distance is used to measure the difference in embedding distribution between different modalities. The optimal transmission distance is defined as: where μ and ν are the embedding distributions of images, text, and structured data, respectively; d(x, y) is the distance metric in the embedding space; Π(μ, ν) is the total coupling that maximizes the alignment of different modal embeddings in the common space by minimizing this distance; γ∈Π(μ, ν) is the joint distribution that connects the distributions μ and ν, satisfying that the marginal distributions are μ and ν, respectively, and X×Y is the product space of the input space; is the p-order mean of the overall distance to ensure consistent units of the results; p is the power exponent parameter; inf is the minimum value, which means finding the minimum transmission cost among all possible couplings; dγ(x,y) represents the differential element integrated at (x,y) according to the joint distribution γ.
[0061] Specifically, in the implementation of the present invention, in order to improve computational efficiency, the Sinkhorn distance is used as a differentiable alternative, and its formula is as follows: in, represents the optimal transmission loss between two distributions P and Q, ε is the regularization coefficient (temperature parameter), which controls the influence of the entropy regularization term, H(γ) is the entropy of γ, <γ,C> is the Frobenius inner product of γ and the cost matrix C, that is, the total transmission cost, Π(P,Q) is the set of all joint distributions such that the marginal distributions are P and Q, and γ is the joint distribution (coupling) matrix.
[0062] As an option, in this embodiment, optimal transmission constraints are applied to any two-to-one combinations of the three modalities. The final modal alignment loss can be defined as: in, is the modality alignment loss, which measures the degree of alignment of different modality embedding distributions in the common space, Z img is the set of embedding vectors of image modality, Z txt is the set of embedding vectors of text modality, Z str is the set of embedding vectors of structured data modality, three The terms represent the optimal transmission losses between the three pairs of modes: image-text, image structure, and text structure.
[0063] This loss function and the valuation loss together constitute the overall optimization goal during training, enabling the model to achieve consistent alignment of modal embeddings while maintaining valuation accuracy.
[0064] In one possible implementation, in order to improve the stability and regularization effect of training, a truncation or threshold processing mechanism can be introduced to the transmission matrix γ to retain only the most significant transmission path and reduce gradient noise interference.
[0065] In some embodiments, the embedding space can be further divided into multiple subspace regions, and the local optimal transmission distance is calculated for each region, followed by weighted aggregation. This mechanism is used to capture local structural heterogeneity and is applicable to local offset issues in enterprise asset valuation, such as scenarios where text descriptions and image content differ slightly but need to be jointly valued.
[0066] Specifically, to improve the discrimination ability after modal alignment, this embodiment also introduces an inter-modal discriminator to perform binary classification on embedding pairs that are from the same modality. The optimization goal is to minimize the OT alignment error while maximizing modal separability. This adversarial mechanism ensures that modal embeddings retain differentiated features while maintaining overall consistency, improving the clarity of the estimated discriminant boundary.
[0067] In terms of training strategy, the optimal transmission module and mapper network parameters are collaboratively optimized. Through an end-to-end backpropagation process, the modal embedding representation and its distribution alignment path are dynamically updated, ultimately achieving triple alignment at the semantic, structural, and distribution levels.
[0068] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. The asset evaluation method based on multimodal large model is characterized by: The following steps are involved: Preprocess the image data, text data, and structured data of the enterprise to be assessed to generate standardized multimodal input data; Based on the standardized multimodal input data, the image data, the text data, and the structured data are processed using an image encoder, a text encoder, and a structured data encoder, respectively, to extract a high-dimensional feature vector of each modality; By constructing a mapper network, the obtained high-dimensional feature vectors of each modality are mapped to the same common embedding space; The optimal transfer method is used to measure the embedding distribution difference between each modality and further remove the heterogeneity between modalities by minimizing the difference.
2. The asset evaluation method based on a multimodal large model according to claim 1 is characterized in that: The pretreatment includes: Normalize the image data and scale the pixel values of the image to a specified range; Perform word segmentation on text data and use word vectorization to convert the text into a fixed-dimensional word vector representation; Normalize structured data and map numerical data to a uniform scale range.
3. The asset evaluation method based on a multimodal large model according to claim 1 is characterized in that: The multimodal input data includes: Image data, which represents the appearance of corporate assets, including product photos, factory facilities, and office environment images; Text data, which represents written information related to corporate assets, including corporate announcements, financial reports, patent documents, and industry analysis reports; Structured data represents the numerical information of corporate assets, including financial statements, balance sheets, and fixed asset lists.
4. The asset evaluation method based on a multimodal large model according to claim 1 is characterized in that: The high-dimensional feature vector includes: For image data, the high-dimensional feature vector is an image feature vector extracted by a convolutional neural network; For text data, the high-dimensional feature vector is a text representation extracted based on a pre-trained language model; For structured data, the high-dimensional feature vector is a numerical feature extracted by a multi-layer perceptron processing network.
5. The asset evaluation method based on a multimodal large model according to claim 1 is characterized in that: The mapper network comprises: The mapper network is trained by minimizing a loss function consisting of a weighted sum of alignment loss and valuation prediction error, as follows: in, is the modality alignment loss, is the asset valuation regression error, is the total loss function used to optimize the training objective of the entire mapper network, and λ is the loss weighting coefficient.
6. The asset evaluation method based on a multimodal large model according to claim 1, characterized in that: The public embedding space includes: A unified high-dimensional vector space, where each modality can be mapped to this space and have similar semantic representations after being processed by the corresponding encoder; The common embedding space can align different modalities so that the feature vectors of images, text and structured data have consistent geometric structures and similar semantic expressions within the space; The structure of the common embedding space is dynamically adjusted through a learning process to adapt to the characteristics of different modalities and reduce the representation differences between the modalities.
7. The asset evaluation method based on a multimodal large model according to claim 1 is characterized in that: The optimal transmission method includes: The optimal transmission distance is used to measure the difference in embedding distribution between different modalities. The optimal transmission distance is defined as: where μ and ν are the embedding distributions of images, text, and structured data, respectively; d(x, y) is the distance metric in the embedding space; Π(μ, ν) is the total coupling that maximizes the alignment of different modal embeddings in the common space by minimizing this distance; γ∈Π(μ, ν) is the joint distribution that connects the distributions μ and ν, satisfying that the marginal distributions are μ and ν, respectively, and X×Y is the product space of the input space; is the p-order mean of the overall distance to ensure consistent units of the results; p is the power exponent parameter; inf is the minimum value, which means finding the minimum transmission cost among all possible couplings; dγ(x,y) represents the differential element integrated at (x,y) according to the joint distribution γ.
8. The asset evaluation method based on a multimodal large model according to claim 1 is characterized in that: The embedding distribution differences include: The differences in the distribution generated by each modality in the embedding space are specifically manifested in the different geometric forms of the embedded features of images, text, and structured data in high-dimensional space; Image data usually presents a dense pixel structure, text data is represented by word embedding, and structured data is expressed by numerical features; This difference needs to be measured and minimized through an optimal transfer method so that features from different modalities can share similar semantic representations in a common embedding space.
9. The asset evaluation method based on a multimodal large model according to claim 1, characterized in that: Minimizing the difference includes: By designing a joint optimization objective function, the alignment error between modalities and the asset valuation error are optimized; In the joint optimization objective, the modality alignment loss controls the distribution differences of different modality embedding vectors in the common embedding space through regularization methods.
10. The asset evaluation method based on a multimodal large model according to claim 1, characterized in that: The heterogeneity between modalities includes: Different modal data are represented in different ways: image data is represented as a pixel matrix, text data is represented as a word vector, and structured data is represented as a numerical table data. Semantic differences between modalities and differences in representation between text and images at the semantic level lead to deviations in asset valuations due to the different expressive capabilities of different modalities.
Citation Information
Cited By
Heterogeneous data fusion device and method for AI large model pre-training and medium
CN120974435A
Multi-modal feature alignment method and device for heterogeneous data and medium
CN121527790A
Multi-modal data-oriented credit large model joint modeling and collaborative reasoning method and device
CN121581236A
Multi-modal data-oriented credit large model joint modeling and collaborative reasoning method and device
CN121581236B