Cross-domain data asset intelligent pricing method and system based on multi-modal large model

By employing a multimodal large-scale model-based intelligent pricing method for cross-domain data assets, which combines automatic asset ontology identification, feature fusion, causal chain modeling, and collaborative game optimization, this approach addresses the shortcomings of existing technologies in multimodal modeling and poor cross-domain adaptability for data asset pricing. It achieves intelligent pricing that is highly accurate, interpretable, and risk-controllable.

CN120952834AInactive Publication Date: 2025-11-14HANGZHOU LEMANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511059856.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing data asset pricing methods lack the ability to model multimodal and complex data structures, making it difficult to fully reflect the true value of data assets. Furthermore, the models have insufficient generalization ability and poor domain adaptability, leading to pricing deviations when applied across domains. They also lack transparency and interpretability, making it difficult to support fair trading and risk control of high-value data assets.

Method used

A cross-domain data asset intelligent pricing method based on a multimodal large model is adopted. Through technologies such as automatic asset ontology identification, feature fusion, causal chain modeling, collaborative game optimization and domain invariant regularization, it realizes the fusion modeling and intelligent reasoning of multi-source information, thereby improving adaptability and generalization.

Benefits of technology

It achieves intelligent pricing of data assets with high precision, strong adaptability, interpretability and controllable risk, supports the circulation of data assets across domains and scenarios, improves the transparency and fairness of the pricing process, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952834A_ABST
    Figure CN120952834A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain data asset intelligent pricing method and system based on a multi-modal large model, and the method comprises the following steps: S1, obtaining multi-modal information, and generating an asset ontology vector and a fusion feature vector; s2, external feature information is collected and coded to generate external features used for pricing; s3, combining the fusion features and the external features to form pricing input features; s4, inputting an improved Visual BERT model, and generating pricing deep feature representation; s5, performing causal chain modeling and intervention reasoning, and outputting a causal chain path and an intervention result; s6, inputting a collaborative game and a constraint process, and outputting pricing characteristics; and S7, introducing a domain invariance regular term, and outputting a pricing result in combination with market feedback optimization parameters. The cross-domain data asset pricing method improves the intellectualization and adaptability of cross-domain data asset pricing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data asset management technology, and in particular to a cross-domain data asset intelligent pricing method and system based on a multimodal large model. Background Technology

[0002] Currently, with the continuous advancement of data elementization and data assetization, the importance of data assets in various industries such as finance, industry, healthcare, and transportation is becoming increasingly prominent. The trading, circulation, and valuation of data assets have become one of the core driving forces for the development of the digital economy. In existing technologies, mainstream data asset pricing methods are mostly based on single-modal data, mainly relying on structured data and expert rules, or depending on traditional statistical modeling and regression analysis. For example, some methods use asset attributes and transaction history as the main features, combined with linear regression, weighted scoring, etc., to perform static valuation of data assets. Other methods attempt to use expert experience or market average levels for correction to achieve simple data asset valuation. Although these methods have achieved certain results in scenarios with relatively complete structured data, they generally lack the ability to model multimodal and complex data structures, making it difficult to fully reflect the true value of data assets.

[0003] With the rapid development of artificial intelligence technology, multimodal learning models have been gradually introduced into the field of data processing. These models can simultaneously integrate different types of data information such as text, images, time series, and speech, improving the ability to express asset ontological characteristics and external influencing factors. In recent years, cross-modal large models such as vision-language joint pre-trained models (such as VisualBERT) have become a research hotspot in the field of intelligent valuation. These models have end-to-end multi-source information fusion capabilities and are expected to make up for the shortcomings of traditional data asset pricing methods in dealing with multidimensional and complex features. However, most existing multimodal large models only focus on the simple splicing and fusion of images and text, lacking a collaborative modeling mechanism for diverse and heterogeneous information such as asset ontological attributes, industry context, market dynamics, and expert knowledge. For external features such as market conditions, user behavior, and policy changes, traditional methods often process them in isolation, failing to achieve multi-level causal reasoning and collaborative optimization.

[0004] Furthermore, existing multimodal data asset pricing methods generally suffer from insufficient model generalization ability and poor domain adaptability at the algorithm level. Most models fail to fully incorporate domain-invariant regularization terms, making it difficult to cope with the distribution changes of data assets in different business scenarios or time periods. This leads to significant deviations in pricing results when applied across domains. The lack of causal chain reasoning and market game mechanisms also makes data asset valuation lack transparency and interpretability, making it difficult to support fair trading and risk control of high-value data assets. Existing solutions have significant defects in key aspects such as pricing feature fusion, causal relationship modeling, collaborative game optimization, and multi-domain consistency constraints, making it difficult to meet the actual needs of large-scale, cross-industry data asset circulation and highly reliable intelligent pricing.

[0005] Therefore, how to provide a method and system for intelligent pricing of cross-domain data assets based on a multimodal large model is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a cross-domain intelligent pricing method for data assets based on a multimodal large model. This invention comprehensively utilizes multimodal deep learning technologies such as artificial intelligence, natural language processing, computer vision, and temporal modeling. Through methods such as automatic asset ontology identification, feature fusion, causal chain modeling, collaborative game optimization, and domain invariant regularization, it realizes the fusion modeling and intelligent reasoning of multi-source information in the data asset pricing process, and has the advantages of strong adaptability, good generalization, and strong risk identification capability.

[0007] The cross-domain data asset intelligent pricing method based on a multimodal large model according to an embodiment of the present invention includes the following steps:

[0008] S1. Obtain multimodal information associated with data assets, use the ontology automatic recognition module to generate the original asset information set and asset ontology vector, and perform feature encoding and fusion to obtain the fused feature vector;

[0009] S2. Collect external feature information related to data assets, and use the feature processing module to encode and generate external features for pricing;

[0010] S3. Combine the fused feature vector with the external features used for pricing to form the pricing input features;

[0011] S4. Input the pricing input features into the improved VisualBERT model to generate a deep feature representation of pricing;

[0012] S5. Perform causal chain modeling on the deep feature representation of pricing to obtain the causal chain path, and perform intervention inference based on specified factors to generate intervention results;

[0013] S6. Input the deep feature representation of pricing, causal chain path and intervention result into the collaborative game mechanism, physical constraints and economic constraints process, perform feature configuration, consistency verification and risk detection, and output pricing features;

[0014] S7. Introduce domain-invariant regularization terms during model training, combine pricing features and market feedback to complete parameter optimization, and output data asset pricing results, causal chain paths and traceability information.

[0015] Optionally, the improved VisualBERT model:

[0016] Introduce asset ontology vectors as input content;

[0017] Add a fusion mechanism for text features, image features, table features, temporal features, and speech features;

[0018] An external feature information encoding and input mechanism is introduced to expand the representation of pricing input features;

[0019] Integrating causal chain modeling and factor intervention reasoning processes;

[0020] By employing a collaborative game mechanism, physical constraints, and economic constraints, we perform feature configuration, consistency verification, and risk detection on the deep characteristics of pricing, causal chain paths, and intervention results.

[0021] Introducing a domain-invariant regularization term during model training reduces the distribution difference between source domain features and target domain features, thereby improving the generalization ability of the pricing model.

[0022] Optionally, S1 specifically includes:

[0023] S11. Collect multimodal information associated with data assets, including text data, image data, tabular data, time-series data, and voice data;

[0024] S12. Denoise, normalize, and structure text data, image data, tabular data, time-series data, and voice data respectively to form a set of original asset information;

[0025] S13. Input the set of original asset information into the ontology automatic recognition module, extract the attribute, category, structure and hierarchical information of text data, image data, tabular data, time series data and voice data respectively, and encode each type of information to obtain the asset ontology vector.

[0026] S14. Input the asset ontology vector and the original asset information set into the feature encoding module to extract text features, image features, table features, time series features and speech features respectively. Perform feature fusion on the text features, image features, table features, time series features and speech features to obtain the fused feature vector.

[0027] Optionally, S2 specifically includes:

[0028] S21. Collect external characteristic information related to data assets, including historical transaction information, market information, expert pricing information, user behavior information, and policy context information;

[0029] S22. Perform preprocessing, normalization, and structuring operations on historical transaction information, market information, expert pricing information, user behavior information, and policy context information to form a set of external features;

[0030] S23. Input the set of external features into the feature processing module, encode the historical transaction information, market information, expert pricing information, user behavior information and policy context information respectively, and generate an external feature vector for pricing.

[0031] Optionally, S4 specifically includes:

[0032] S41. The pricing input features, including the fused feature vector and the external features used for pricing, are concatenated to form an input sequence, which is then input into the improved VisualBERT model.

[0033] S42. In the improved VisualBERT model, positional encoding is added to the input sequence, and all tokens are input into a multi-layer self-attention structure for joint feature modeling.

[0034] S43. In a multi-layer self-attention structure, a multi-head self-attention mechanism is used to interactively compute tokens with different modalities and external features;

[0035] S44. Pool and integrate all output tokens of the last layer of the self-attention structure to output the deep feature representation of pricing.

[0036] Optionally, S5 specifically includes:

[0037] S51. The deep feature representation of pricing is used as the node input to the causal chain modeling module. The structural learning algorithm is used to construct the dependency relationship between feature factors and pricing target variables, and generate a causal graph structure represented by a directed acyclic graph.

[0038] S52. Using the cause-effect graph structure, establish a set of structural equations to define the relationship between each characteristic factor and the pricing target variable, and clearly define the causal dependency path of each variable.

[0039] S53. Targeting the characteristic factor x that requires intervention k Based on Pearl's causal inference theory, intervention reasoning is performed using do calculus:

[0040]

[0041] in, Represents the characteristic factor x k Values Furthermore, given variable z, the probability distribution of output variable y, x -k Including intervention factor x k All other eigenfactors P(x) represents the pricing output probability when all factor values ​​are known. -k |z) represents the distribution of other characteristic factors given variable z;

[0042] S54. Output intervention inference values ​​based on the inference results to form a pricing causal path and intervention results based on causal chain modeling and factor intervention.

[0043] Optionally, S6 specifically includes:

[0044] S61. Input the deep feature representation of pricing, causal chain path, and intervention results into the collaborative game mechanism, and set a strategy space for each market participant, with each market participant having an objective function:

[0045] u i (x i ,x -i )=α i ·r i (x i ,x -i )-β i ·c i (x i ,x -i );

[0046] Among them, u i (x i ,x -i Let x represent the objective function of the i-th market participant. i Let x represent the strategy of the i-th market participant. -i Let r represent the strategies of all market participants except the i-th market participant. i (x i ,x -i Let ) represent the profit of the i-th market participant, and c i (x i ,x -i α represents the cost of the i-th market participant.i and β i These are the weight parameters;

[0047] S62. Perform joint optimization on the strategy variables of all market participants by solving the Nash equilibrium condition u. i To obtain the equilibrium strategy for each market participant, where, This represents the equilibrium strategy of the i-th market participant. This represents the equilibrium strategy of other market participants, u i Represent the objective function;

[0048] S63. The pricing features output by the collaborative game mechanism are configured with weights and input into the physical constraint function and economic constraint function for verification;

[0049] S64. Perform consistency verification and risk detection on the pricing features that have been tested through collaborative game mechanism, physical constraints and economic constraints, and output the pricing features.

[0050] Optionally, S7 specifically includes:

[0051] S71. During model training, obtain the pricing results and real pricing labels output by the model, compare the model output with the real labels, and calculate the main loss function according to the set loss metric method. The main loss function is used to measure the error between the model output and the actual pricing.

[0052] S72. Based on the characteristics of the source domain samples and the target domain samples, calculate the domain-invariant regularization loss. The domain-invariant regularization term adopts the maximum mean difference loss function:

[0053]

[0054] Among them, L MMD This represents the maximum mean difference loss, n. s Indicates the number of samples in the source domain. Let n represent the features of the i-th sample in the source domain. t Indicates the number of samples in the target domain. Let φ represent the feature of the j-th sample in the target domain, and let φ represent the feature mapping function.

[0055] S73. The main loss function and the domain-invariant regularization loss function are weighted and combined to form the total loss function. The model parameters are then optimized based on the total loss function, and the final output is the data asset pricing result, causal chain path and traceability information.

[0056] The cross-domain data asset intelligent pricing system based on a multimodal large model according to an embodiment of the present invention includes:

[0057] The multimodal data acquisition module is used to acquire text data, image data, tabular data, time-series data, voice data, and external feature information associated with data assets.

[0058] The automatic ontology recognition module is used to process multimodal information, extract asset attributes, categories, structures and hierarchical information, and generate a set of original asset information and an asset ontology vector.

[0059] The feature encoding and fusion module is used to extract text features, image features, table features, temporal features and speech features from the original asset information set and asset ontology vector, respectively, and perform feature fusion to generate a fused feature vector;

[0060] The feature processing module is used to preprocess, normalize, structure, and encode external feature information to generate external feature vectors for pricing.

[0061] The feature combination module is used to combine the fused feature vector with the external feature vector used for pricing to form the pricing input feature;

[0062] An improved VisualBERT modeling module is used to perform multimodal joint modeling and inference on pricing input features to generate deep feature representations of pricing.

[0063] The causal chain modeling and intervention reasoning module is used to build a causal graph structure and structural equation set based on the deep feature representation of pricing, infer the causal chain path of pricing, perform intervention reasoning on specified feature factors, and output the causal chain path and intervention results.

[0064] The collaborative game mechanism module is used to input the deep feature representation of pricing, causal chain path and intervention result into the collaborative game mechanism, set the strategy space and objective function of each market participant, and solve the pricing feature configuration weights through joint optimization and Nash equilibrium.

[0065] The physical constraint module and the economic constraint module are used to apply physical and economic constraints to the weights of pricing features and output pricing features that meet the constraints.

[0066] The consistency verification and risk detection module is used to perform consistency verification and risk detection on the weights of pricing features that meet the constraints, and output the pricing features.

[0067] The model training and domain invariance regularization module is used to optimize model parameters based on the main loss function and the maximum mean difference loss function during model training, so as to achieve consistency between the distribution of source domain features and target domain features, and output data asset pricing results, causal chain paths and traceability information.

[0068] The beneficial effects of this invention are:

[0069] This invention achieves unified fusion and deep modeling of various information such as text, images, tables, time series, and voice of data assets through a multimodal large model architecture. It effectively explores the synergistic value of asset ontology attributes and heterogeneous features such as external markets, policies, and expert knowledge. Compared with traditional methods, this invention not only has higher richness and completeness in feature expression, but also greatly improves the representation ability of intrinsic attributes and scene features of data assets through automatic ontology recognition and feature fusion mechanisms.

[0070] At the level of intelligent reasoning, this invention integrates causal chain modeling and intervention reasoning methods, which can automatically identify the causal relationships between various characteristic factors, support intervention analysis for key factors, significantly enhance the interpretability and transparency of the pricing process, and introduce a collaborative game mechanism, enabling the system to dynamically integrate the interests of different market participants and achieve multi-party optimal strategy linkage through Nash equilibrium, effectively improving the fairness and market adaptability of pricing decisions.

[0071] Furthermore, this invention employs domain-invariant regularization to constrain the model training process, achieving consistency in the feature distributions of the source and target domains. This effectively enhances the model's generalization ability across domains and scenarios. Combined with consistency verification and risk detection mechanisms, the system can detect and avoid abnormal pricing characteristics in advance, providing strong support for the highly reliable circulation and intelligent pricing of data assets. Overall, this invention significantly breaks through the dependence of existing pricing methods on single modalities, static rules, and weak generalization capabilities, and can meet the actual needs of the complex and ever-changing data asset market for high-precision, highly adaptable, interpretable, and risk-controllable intelligent pricing. Attached Figure Description

[0072] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0073] Figure 1 The flowchart shows the cross-domain data asset intelligent pricing method based on a multimodal large model proposed in this invention.

[0074] Figure 2 This is a schematic diagram of the improved VisualBERT model structure of the cross-domain data asset intelligent pricing method based on a multimodal large model proposed in this invention.

[0075] Figure 3 This is a flowchart illustrating the data asset pricing feature fusion and causal chain modeling process of the cross-domain data asset intelligent pricing method based on a multimodal large model proposed in this invention. Detailed Implementation

[0076] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0077] refer to Figure 1-3 A cross-domain data asset intelligent pricing method based on a multimodal large model includes the following steps:

[0078] S1. Obtain multimodal information associated with data assets, use the ontology automatic recognition module to generate the original asset information set and asset ontology vector, and perform feature encoding and fusion to obtain the fused feature vector;

[0079] S2. Collect external feature information related to data assets, and use the feature processing module to encode and generate external features for pricing;

[0080] S3. Combine the fused feature vector with the external features used for pricing to form the pricing input features;

[0081] S4. Input the pricing input features into the improved VisualBERT model to generate a deep feature representation of pricing;

[0082] S5. Perform causal chain modeling on the deep feature representation of pricing to obtain the causal chain path, and perform intervention inference based on specified factors to generate intervention results;

[0083] S6. Input the deep feature representation of pricing, causal chain path and intervention result into the collaborative game mechanism, physical constraints and economic constraints process, perform feature configuration, consistency verification and risk detection, and output pricing features;

[0084] S7. Introduce domain-invariant regularization terms during model training, combine pricing features and market feedback to complete parameter optimization, and output data asset pricing results, causal chain paths and traceability information.

[0085] This invention enables automatic pricing of multimodal data assets in cross-domain scenarios, greatly improving the intelligence and accuracy of data asset evaluation. By automatically collecting and processing multi-source information such as text, images, tables, time series, and voice, the system can fully restore the intrinsic attributes and multi-dimensional value characteristics of data assets, enabling the pricing model to have comprehensive perception, in-depth expression, and flexible adaptive capabilities, thereby promoting the efficient and reliable circulation of data assets across different industries and business scenarios.

[0086] In this embodiment, the improved VisualBERT model is:

[0087] Introduce asset ontology vectors as input content;

[0088] Add a fusion mechanism for text features, image features, table features, temporal features, and speech features;

[0089] An external feature information encoding and input mechanism is introduced to expand the representation of pricing input features;

[0090] Integrating causal chain modeling and factor intervention reasoning processes;

[0091] By employing a collaborative game mechanism, physical constraints, and economic constraints, we perform feature configuration, consistency verification, and risk detection on the deep characteristics of pricing, causal chain paths, and intervention results.

[0092] Introducing a domain-invariant regularization term during model training reduces the distribution difference between source domain features and target domain features, thereby improving the generalization ability of the pricing model.

[0093] This invention addresses the limitations of the traditional VisualBERT model by introducing asset ontology vectors as input and expanding its ability to input and express external features, further enabling joint encoding of complex data types. Through the integration of causal chain modeling, collaborative game theory mechanisms, and domain-invariant regularization terms, the model possesses stronger feature fusion, cross-domain generalization, and risk detection capabilities, effectively ensuring the scientific rigor, adaptability, and security of the data asset pricing process.

[0094] In this embodiment, S1 specifically includes:

[0095] S11. Collect multimodal information associated with data assets, including text data, image data, tabular data, time-series data, and voice data;

[0096] S12. Denoise, normalize, and structure text data, image data, tabular data, time-series data, and voice data respectively to form a set of original asset information;

[0097] S13. Input the original asset information set into the ontology automatic recognition module, extract attribute, category, structure, and hierarchical information from text data, image data, tabular data, time-series data, and voice data respectively, and encode each type of information to obtain the asset ontology vector:

[0098] V ont =[v text ,v img ,v tab ,v seq ,v aud ];

[0099] Among them, V ont Represents the asset ontology vector, v text Represents ontology information based on text data encoding, vimg Represents ontology information encoded from image data, v tab This represents ontology information based on tabular data encoding, v seq This represents ontology information based on time-series data encoding, v aud Represents ontology information encoded from speech data;

[0100] S14. Input the asset ontology vector and the original asset information set into the feature encoding module to extract text features, image features, table features, time series features and speech features respectively. Perform feature fusion on the text features, image features, table features, time series features and speech features to obtain the fused feature vector.

[0101] In the multimodal information acquisition and processing stage, this invention effectively enhances the ability to analyze and summarize data from different sources through automatic ontology recognition and multi-level feature fusion. The system not only achieves standardized processing of multimodal data but also automatically identifies data structure, hierarchy, and application attributes, outputting high-quality, expressive fused feature vectors. This lays a solid data foundation for subsequent intelligent modeling and pricing processes, improving the overall information richness and detail of the pricing system.

[0102] In this embodiment, S2 specifically includes:

[0103] S21. Collect external characteristic information related to data assets, including historical transaction information, market information, expert pricing information, user behavior information, and policy context information;

[0104] S22. Perform preprocessing, normalization, and structuring operations on historical transaction information, market information, expert pricing information, user behavior information, and policy context information to form a set of external features;

[0105] S23. Input the external feature set into the feature processing module, and encode historical transaction information, market information, expert pricing information, user behavior information, and policy context information respectively to generate an external feature vector for pricing:

[0106] F ext =[f his ,f mar ,f exp ,f usr ,f pol ];

[0107] Among them, F ext Let f represent the external feature vector used for pricing. his This indicates the characteristics based on the encoding of historical transaction information, f mar This indicates the characteristics of the market information encoding, f expThis indicates the characteristics based on expert pricing information encoding, f usr f represents the features encoded based on user behavior information. pol This represents the characteristics encoded based on policy context information.

[0108] This invention focuses on the efficient collection and structured processing of external feature information. It can automatically integrate multiple influencing factors such as market history, expert evaluation, user behavior, and policy changes, greatly enhancing the system's responsiveness to market dynamics, industry trends, and policy guidance. After refined coding and normalization, external features and ontological features are deeply integrated, providing rich, accurate, and timely decision-making basis for pricing models, making data asset valuation more closely aligned with market reality.

[0109] In this embodiment, S4 specifically includes:

[0110] S41. The pricing input features, including the fused feature vector and the external features used for pricing, are concatenated to form an input sequence, which is then input into the improved VisualBERT model.

[0111] S42. In the improved VisualBERT model, positional encoding is added to the input sequence, and all tokens are input into a multi-layer self-attention structure for joint feature modeling.

[0112] S43. In a multi-layer self-attention structure, a multi-head self-attention mechanism is used to interactively compute tokens with different modalities and external features:

[0113]

[0114] Where Attention(Q,K,V) represents the multi-head self-attention output, Q is the query matrix, K is the key matrix, V is the value matrix, and d k The dimension of the key vector;

[0115] S44. Pool and integrate all output tokens of the last layer of the self-attention structure to output the deep feature representation of pricing.

[0116] In the modeling stage of pricing input features, this invention effectively breaks down the semantic barriers between internal attributes and external influencing factors by concatenating and jointly expressing multimodal features and external features. Utilizing an improved VisualBERT structure, a multi-layer self-attention mechanism, and a feature pooling strategy, the model can uncover complex interactions and potential value contributions between high-dimensional features, ensuring that the final pricing features comprehensively reflect the ontology and environmental value of data assets.

[0117] In this embodiment, S5 specifically includes:

[0118] S51. The deep feature representation of pricing is used as the node input to the causal chain modeling module. The structural learning algorithm is used to construct the dependency relationship between feature factors and pricing target variables, and generate a causal graph structure represented by a directed acyclic graph.

[0119] S52. Using the cause-effect graph structure, establish a set of structural equations to define the relationship between each characteristic factor and the pricing target variable, and clearly define the causal dependency path of each variable.

[0120] S53. Targeting the characteristic factor x that requires intervention k Based on Pearl's causal inference theory, intervention reasoning is performed using do calculus:

[0121]

[0122] in, Represents the characteristic factor x k Values Furthermore, given variable z, the probability distribution of output variable y, x -k Including intervention factor x k All other eigenfactors P(x) represents the pricing output probability when all factor values ​​are known. -k |z) represents the distribution of other characteristic factors given variable z. This formula comes from Pearl's causal inference theory.

[0123] S54. Output intervention inference values ​​based on the inference results to form a pricing causal path and intervention results based on causal chain modeling and factor intervention.

[0124] This invention employs causal chain modeling and intervention inference methods, achieving for the first time the mining of causal relationships and causal path analysis of characteristic factors in data asset pricing. The system can automatically establish causal graph structures and structural equation sets, and combined with Pearl causal inference theory, conduct intervention simulations for key variables to reveal the core driving factors affecting pricing, thereby improving the transparency, traceability, and scientific explanatory power of the data asset pricing process.

[0125] In this embodiment, S6 specifically includes:

[0126] S61. Input the deep feature representation of pricing, causal chain path, and intervention results into the collaborative game mechanism, and set a strategy space for each market participant, with each market participant having an objective function:

[0127] u i (x i ,x -i )=α i ·r i (x i ,x -i )-βi ·c i (x i ,x -i );

[0128] Among them, u i (x i ,x -i Let x represent the objective function of the i-th market participant. i Let x represent the strategy of the i-th market participant. -i Let r represent the strategies of all market participants except the i-th market participant. i (x i ,x -i Let ) represent the profit of the i-th market participant, and c i (x i ,x -i α represents the cost of the i-th market participant. i and β i These are the weight parameters;

[0129] S62. Perform joint optimization on the strategy variables of all market participants by solving the Nash equilibrium condition u. i To obtain the equilibrium strategy for each market participant, where, This represents the equilibrium strategy of the i-th market participant. This represents the equilibrium strategy of other market participants, u i Let represent the objective function. This condition is the classic mathematical definition of Nash equilibrium, ensuring that each market participant, with the strategies of other participants fixed, can obtain an objective function value no lower than that of any other feasible strategy by adopting the equilibrium strategy.

[0130] S63. The pricing features output by the collaborative game mechanism are configured with weights and input into the physical constraint function and economic constraint function for verification. The physical constraint function and economic constraint function are as follows:

[0131] f phy (F deep )≤λ phy f eco (F deep )≤λ eco ;

[0132] Among them, f phy Let λ represent the physical constraint function. phy f represents the physical constraint threshold. eco Let λ represent the economic constraint function. eco F represents the economic constraint threshold. deep This represents the deep features of pricing;

[0133] S64. Perform consistency verification and risk detection on the pricing features that have been tested through collaborative game mechanism, physical constraints and economic constraints, and output the pricing features.

[0134] In the collaborative optimization and constraint verification stages, this invention incorporates the interests of all market participants into the model's optimization objective through a collaborative game mechanism, and utilizes Nash equilibrium to solve for pricing strategies, achieving a balance of interests among all parties. By combining physical and economic constraints, the system can proactively avoid risks and irrational allocations in data asset pricing, further enhancing the security, stability, and fairness of data asset circulation, and promoting the healthy and orderly development of the data market.

[0135] In this embodiment, S7 specifically includes:

[0136] S71. During model training, obtain the pricing results and real pricing labels output by the model, compare the model output with the real labels, and calculate the main loss function according to the set loss metric method. The main loss function is used to measure the error between the model output and the actual pricing.

[0137] S72. Based on the characteristics of the source domain samples and the target domain samples, calculate the domain-invariant regularization loss. The domain-invariant regularization term adopts the maximum mean difference loss function:

[0138]

[0139] Among them, L MMD This represents the maximum mean difference loss, n. s Indicates the number of samples in the source domain. Let n represent the features of the i-th sample in the source domain. t Indicates the number of samples in the target domain. Let φ represent the feature of the j-th sample in the target domain, and let φ represent the feature mapping function.

[0140] S73. The main loss function and the domain-invariant regularization loss function are weighted and combined to form the total loss function. The model parameters are then optimized based on the total loss function, and the final output is the data asset pricing result, causal chain path and traceability information.

[0141] This invention incorporates domain-invariant regularization terms during model training and optimization, effectively addressing model performance fluctuations caused by differences in data distribution across domains. By optimizing feature distribution consistency through the maximum mean difference loss function and dynamically adjusting model parameters in conjunction with market feedback, it significantly improves the generalization ability of the pricing model in different domains and scenarios, providing strong support for the cross-industry and cross-regional circulation of data assets.

[0142] A cross-domain data asset intelligent pricing system based on a multimodal large model includes:

[0143] The multimodal data acquisition module is used to acquire text data, image data, tabular data, time-series data, voice data, and external feature information associated with data assets.

[0144] The automatic ontology recognition module is used to process multimodal information, extract asset attributes, categories, structures and hierarchical information, and generate a set of original asset information and an asset ontology vector.

[0145] The feature encoding and fusion module is used to extract text features, image features, table features, temporal features and speech features from the original asset information set and asset ontology vector, respectively, and perform feature fusion to generate a fused feature vector;

[0146] The feature processing module is used to preprocess, normalize, structure, and encode external feature information to generate external feature vectors for pricing.

[0147] The feature combination module is used to combine the fused feature vector with the external feature vector used for pricing to form the pricing input feature;

[0148] An improved VisualBERT modeling module is used to perform multimodal joint modeling and inference on pricing input features to generate deep feature representations of pricing.

[0149] The causal chain modeling and intervention reasoning module is used to build a causal graph structure and structural equation set based on the deep feature representation of pricing, infer the causal chain path of pricing, perform intervention reasoning on specified feature factors, and output the causal chain path and intervention results.

[0150] The collaborative game mechanism module is used to input the deep feature representation of pricing, causal chain path and intervention result into the collaborative game mechanism, set the strategy space and objective function of each market participant, and solve the pricing feature configuration weights through joint optimization and Nash equilibrium.

[0151] The physical constraint module and the economic constraint module are used to apply physical and economic constraints to the weights of pricing features and output pricing features that meet the constraints.

[0152] The consistency verification and risk detection module is used to perform consistency verification and risk detection on the weights of pricing features that meet the constraints, and output the pricing features.

[0153] The model training and domain invariance regularization module is used to optimize model parameters based on the main loss function and the maximum mean difference loss function during model training, so as to achieve consistency between the distribution of source domain features and target domain features, and output data asset pricing results, causal chain paths and traceability information.

[0154] This invention systematically integrates functional modules such as multimodal data acquisition, ontology recognition, feature fusion, intelligent modeling, causal reasoning, collaborative game theory, and constraint optimization to form a closed-loop intelligent pricing system. The overall solution considers data asset representation, information fusion, scientific pricing, risk control, and cross-domain adaptability, comprehensively enhancing the market value discovery capability of data assets and providing a solid technical foundation for large-scale data asset circulation and highly reliable intelligent pricing in the context of the digital economy.

[0155] Example 1:

[0156] To verify the feasibility of this invention in practice, it was applied to a healthcare platform deeply involved in the data element market, focusing on the asset pricing issue of medical imaging data. The platform operates on a multimodal data-driven model, dealing with raw data supplies from hospitals, imaging centers, and third-party data service providers, as well as responding to diverse data needs from pharmaceutical companies and research institutions. In practice, significant disputes often arise regarding the pricing of different data assets due to differences in data sources, annotation methods, and collection standards. For example, even with lung CT data, factors such as whether there is expert annotation, data resolution, patient follow-up information, and policy guidance can significantly impact the actual price.

[0157] Before adopting this invention, platforms often relied on manual annotation dimensions, data count, and expert experience for simple weighting. In many cases, due to the failure to take into account historical market transaction data and external policy fluctuations, price settings were prone to deviating from the true market level. The operations team often had to repeatedly coordinate with hospitals and purchasers, and interpreting the "reasonable price" of the dataset was time-consuming and laborious. Some datasets with high circulation rates, due to the lack of sufficient external features and scenario information, showed significant price jumps when circulating across regions, affecting the enthusiasm of partners to continue participating.

[0158] After its launch, the platform first performs unified cleaning of previously collected medical images, structured medical records, monitoring sequences, and audio recordings. After standardization processing, the multimodal data is automatically labeled and hierarchically categorized, and features closely related to asset pricing, such as text, images, and time series, are extracted based on the application scenario. The platform further retrieves market transaction records, regional supply and demand information, authoritative expert assessments, and user activity data from recent years, inputting this information along with the multimodal features into the intelligent pricing system. The model automatically calculates the fused features and identifies the variables most significantly affecting prices through causal reasoning.

[0159] For example, for a batch of lung nodule image data with high-quality annotations, the platform system identified three dimensions that have the greatest impact on the final transaction price: "annotation completeness," "expert rating labels," and "market activity." The pricing team only needs to set intervention scenarios in the system to simulate the impact of these factors on the price. For example, increasing annotation completeness by 5% will increase the average transaction price by 12%, while reducing expert labels will increase price volatility, and the system will automatically issue a risk warning. The platform transforms the actual demands of data providers, service providers, and purchasers into objective functions. The intelligent system balances the interests of all parties through game theory optimization and automatic application of market equilibrium constraints. The pricing results are transparent and traceable. Ultimately, the pricing history, adjustment basis, and risk warnings for each asset can be traced. Table 1 shows the comparison results of some indicators of medical imaging data asset transactions on the platform in the recent period:

[0160] Table 1 Comparison of Key Indicators for Medical Imaging Data Asset Transactions

[0161]

[0162]

[0163] The results show that smart pricing significantly improved the platform's average transaction price and market circulation efficiency, and greatly increased user acceptance of pricing ranges and processes. Actual feedback indicates that the new smart pricing system not only reduced human intervention and subjective disputes, but also gave hospitals, service providers, and purchasers greater confidence in their collaborations, resulting in a significant improvement in the overall activity and transaction transparency of the data asset market.

[0164] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A cross-domain data asset intelligent pricing method based on a multimodal large model, characterized in that, Includes the following steps: S1. Obtain multimodal information associated with data assets, use the ontology automatic recognition module to generate the original asset information set and asset ontology vector, and perform feature encoding and fusion to obtain the fused feature vector; S2. Collect external feature information related to data assets, and use the feature processing module to encode and generate external features for pricing; S3. Combine the fused feature vector with the external features used for pricing to form the pricing input features; S4. Input the pricing input features into the improved VisualBERT model to generate a deep feature representation of pricing; S5. Perform causal chain modeling on the deep feature representation of pricing to obtain the causal chain path, and perform intervention inference based on specified factors to generate intervention results; S6. Input the deep feature representation of pricing, causal chain path and intervention results into the collaborative game mechanism, physical constraints and economic constraints process, perform feature configuration, consistency verification and risk detection, and output pricing features; S7. Introduce domain-invariant regularization terms during model training, combine pricing features and market feedback to complete parameter optimization, and output data asset pricing results, causal chain paths and traceability information.

2. The intelligent pricing method for cross-domain data assets based on a multimodal large model according to claim 1, characterized in that, The improved VisualBERT model: Introduce asset ontology vectors as input content; Add a fusion mechanism for text features, image features, table features, temporal features, and speech features; An external feature information encoding and input mechanism is introduced to expand the representation of pricing input features; Integrating causal chain modeling and factor intervention reasoning processes; By employing a collaborative game mechanism, physical constraints, and economic constraints, we perform feature configuration, consistency verification, and risk detection on the deep characteristics of pricing, causal chain paths, and intervention results. Introducing a domain-invariant regularization term during model training reduces the distribution difference between source domain features and target domain features, thereby improving the generalization ability of the pricing model.

3. The intelligent pricing method for cross-domain data assets based on a multimodal large model according to claim 1, characterized in that, S1 specifically includes: S11. Collect multimodal information associated with data assets, including text data, image data, tabular data, time-series data, and voice data; S12. Denoise, normalize, and structure text data, image data, tabular data, time-series data, and voice data respectively to form a set of original asset information; S13. Input the set of original asset information into the ontology automatic recognition module, extract the attribute, category, structure and hierarchical information of text data, image data, tabular data, time series data and voice data respectively, and encode each type of information to obtain the asset ontology vector. S14. Input the asset ontology vector and the original asset information set into the feature encoding module to extract text features, image features, table features, time series features and speech features respectively. Perform feature fusion on the text features, image features, table features, time series features and speech features to obtain the fused feature vector.

4. The intelligent pricing method for cross-domain data assets based on a multimodal large model according to claim 1, characterized in that, S2 specifically includes: S21. Collect external characteristic information related to data assets, including historical transaction information, market information, expert pricing information, user behavior information, and policy context information; S22. Perform preprocessing, normalization, and structuring operations on historical transaction information, market information, expert pricing information, user behavior information, and policy context information to form a set of external features; S23. Input the set of external features into the feature processing module, encode the historical transaction information, market information, expert pricing information, user behavior information and policy context information respectively, and generate an external feature vector for pricing.

5. The intelligent pricing method for cross-domain data assets based on a multimodal large model according to claim 1, characterized in that, S4 specifically includes: S41. The pricing input features, including the fused feature vector and the external features used for pricing, are concatenated to form an input sequence, which is then input into the improved VisualBERT model. S42. In the improved VisualBERT model, positional encoding is added to the input sequence, and all tokens are input into a multi-layer self-attention structure for joint feature modeling. S43. In a multi-layer self-attention structure, a multi-head self-attention mechanism is used to interactively compute tokens with different modalities and external features; S44. Pool and integrate all output tokens of the last layer of the self-attention structure to output the deep feature representation of pricing.

6. The intelligent pricing method for cross-domain data assets based on a multimodal large model according to claim 1, characterized in that, S5 specifically includes: S51. The deep feature representation of pricing is used as the node input to the causal chain modeling module. The structural learning algorithm is used to construct the dependency relationship between feature factors and pricing target variables, and generate a causal graph structure represented by a directed acyclic graph. S52. Using the cause-effect graph structure, establish a set of structural equations to define the relationship between each characteristic factor and the pricing target variable, and clearly define the causal dependency path of each variable. S53. Targeting the characteristic factor x that requires intervention k Based on Pearl's causal inference theory, intervention reasoning is performed using do calculus: in, Represents the characteristic factor x k Values Furthermore, given variable z, the probability distribution of output variable y, x -k Including intervention factor x k All other eigenfactors P(x) represents the pricing output probability when all factor values ​​are known. -k |z) represents the distribution of other characteristic factors given variable z; S54. Output intervention inference values ​​based on the inference results to form a pricing causal path and intervention results based on causal chain modeling and factor intervention.

7. The intelligent pricing method for cross-domain data assets based on a multimodal large model according to claim 1, characterized in that, S6 specifically includes: S61. Input the deep feature representation of pricing, causal chain path, and intervention results into the collaborative game mechanism, and set a strategy space for each market participant, with each market participant having an objective function: u i (x i ,x -i )=α i ·r i (x i ,x -i )-β i ·c i (x i ,x -i ); Among them, u i (x i ,x -i Let x represent the objective function of the i-th market participant. i Let x represent the strategy of the i-th market participant. -i Let r represent the strategies of all market participants except the i-th market participant. i (x i ,x -i Let ) represent the profit of the i-th market participant, and c i (x i ,x -i α represents the cost of the i-th market participant. i and β i These are the weight parameters; S62. Perform joint optimization on the strategy variables of all market participants by solving the Nash equilibrium condition u. i To obtain the equilibrium strategy for each market participant, where, This represents the equilibrium strategy of the i-th market participant. This represents the equilibrium strategy of other market participants, u i Denotes the objective function; S63. The pricing features output by the collaborative game mechanism are configured with weights and input into the physical constraint function and economic constraint function for verification; S64. Perform consistency verification and risk detection on the pricing features that have been tested through collaborative game mechanism, physical constraints and economic constraints, and output the pricing features.

8. The intelligent pricing method for cross-domain data assets based on a multimodal large model according to claim 1, characterized in that, Specifically, S7 includes: S71. During model training, obtain the pricing results and real pricing labels output by the model, compare the model output with the real labels, and calculate the main loss function according to the set loss metric method. The main loss function is used to measure the error between the model output and the actual pricing. S72. Based on the characteristics of the source domain samples and the target domain samples, calculate the domain-invariant regularization loss. The domain-invariant regularization term adopts the maximum mean difference loss function: Among them, L MMD This represents the maximum mean difference loss, n. s Indicates the number of samples in the source domain. Let n represent the features of the i-th sample in the source domain. t Indicates the number of samples in the target domain. Let φ represent the feature of the j-th sample in the target domain, and let φ represent the feature mapping function. S73. The main loss function and the domain-invariant regularization loss function are weighted and combined to form the total loss function. The model parameters are then optimized based on the total loss function, and the final output is the data asset pricing result, causal chain path and traceability information.

9. A cross-domain data asset intelligent pricing system based on a multimodal large model, comprising executing the cross-domain data asset intelligent pricing method based on a multimodal large model as described in any one of claims 1 to 8, characterized in that, include: The multimodal data acquisition module is used to acquire text data, image data, tabular data, time-series data, voice data, and external feature information associated with data assets. The automatic ontology recognition module is used to process multimodal information, extract asset attributes, categories, structures and hierarchical information, and generate a set of original asset information and an asset ontology vector. The feature encoding and fusion module is used to extract text features, image features, table features, temporal features and speech features from the original asset information set and asset ontology vector, respectively, and perform feature fusion to generate a fused feature vector; The feature processing module is used to preprocess, normalize, structure, and encode external feature information to generate external feature vectors for pricing. The feature combination module is used to combine the fused feature vector with the external feature vector used for pricing to form the pricing input feature; An improved VisualBERT modeling module is used to perform multimodal joint modeling and inference on pricing input features to generate deep feature representations of pricing. The causal chain modeling and intervention reasoning module is used to build a causal graph structure and structural equation set based on the deep feature representation of pricing, infer the causal chain path of pricing, perform intervention reasoning on specified feature factors, and output the causal chain path and intervention results. The collaborative game mechanism module is used to input the deep feature representation of pricing, causal chain path and intervention result into the collaborative game mechanism, set the strategy space and objective function of each market participant, and solve the pricing feature configuration weights through joint optimization and Nash equilibrium. The physical constraint module and the economic constraint module are used to apply physical and economic constraints to the weights of pricing features and output pricing features that meet the constraints. The consistency verification and risk detection module is used to perform consistency verification and risk detection on the weights of pricing features that meet the constraints, and output the pricing features. The model training and domain invariance regularization module is used to optimize model parameters based on the main loss function and the maximum mean difference loss function during model training, so as to achieve consistency between the distribution of source domain features and target domain features, and output data asset pricing results, causal chain paths and traceability information.