Multi-modal large language model data management system applied to medical field
By introducing a multimodal large language model data management system into the medical data management system, and using blockchain technology and contract model, the problems of medical data sharing incentives and information asymmetry are solved, and efficient and secure sharing and utilization of medical data are achieved.
Patent Information
- Application Number
- CN202411999906.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
When applying generative artificial intelligence and retrieval enhancement generation technology to medical data management in the medical Internet of Things, the prior art faces the problems of data sharing incentives, information asymmetry, and low data quality and sharing efficiency.
A multimodal large language model data management system is proposed, including a data transmission module, a data sharing contract customization module, an optimal contract generation module, a multimodal large language model pre-training module, a hybrid multimodal search enhancement generation module and a multimodal large language model inference optimization module. The system uses blockchain technology and contract model to realize the incentive mechanism for data sharing, and generates the optimal contract strategy through deep reinforcement learning to improve data quality and sharing efficiency.
It realizes efficient sharing and utilization of medical data, solves the problem of information asymmetry, improves data quality and sharing efficiency, and ensures the security and efficiency of data processing and sharing processes.
Smart Images

Figure CN119943242A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing and Internet of Things technology, and in particular to a multimodal large language model data management method for the medical field. Background Art
[0002] Deep reinforcement learning technology is a powerful fusion technology that integrates the excellent feature extraction ability of deep learning and the excellent decision-making ability of reinforcement learning, so that it can directly give the optimal decision output based on the input multi-dimensional data, forming an end-to-end decision control system. Deep reinforcement learning based on the generative diffusion model is a cutting-edge technology fusion. The generative diffusion model generates new data samples by gradually adding noise and reverse denoising. It has powerful generation capabilities and modeling capabilities for complex distributions. Deep reinforcement learning based on the generative diffusion model has great application potential in many fields. For example, in urban data management, it can help formulate more flexible and adaptable comprehensive governance plans [CN118780445A]; in robot simulation, it can highly restore the operating state of complex robots, use the learned optimal strategy to make decisions, and provide effective reference for practical applications [CN202311688339.0]. However, although deep reinforcement learning based on the generative diffusion model has achieved remarkable results in many application scenarios, its application in the field of medical data management, especially in solving the optimal contract of the contract theory model for data sharing pricing based on the generative diffusion model, still needs further exploration and innovation to meet the special needs and high standards of medical data management.
[0003] Contract theory, as an important branch of economics, is mainly used to study how to design optimal contracts to calibrate incentives between parties with asymmetric information. This theory has been widely used in many fields, such as wireless communications and artificial intelligence. In the context of data sharing, information asymmetry is prevalent because data holders usually have more data information than data users. Contract theory plays a key role in this context. It can effectively incentivize data sharing by ensuring that both parties can benefit from the exchange. For example, in healthcare applications, WYBLim et al. published a two-stage incentive mechanism that fully considers the user's willingness to participate and meets inter-period incentive compatibility. This dynamic contract design not only meets the essential constraints, but also obtains higher profits than the uniform pricing scheme [IEEE Internet of Things Journal, vol. 8, no. 23, pp. 16 853–16 862, 2020]. In the context of mobile AI-generated content networks with drones, J. Wen et al. published a contract theory model based on Age of Information (AoI) to incentivize fresh data contribution between drones [IEEE, 2023, pp. 1–6]. In addition, J. Kang et al. proposed an effective incentive mechanism that combines reputation and contract theory to encourage high-reputation mobile devices with high-quality data to participate in model learning in federated learning scenarios [IEEE Internet of Things Journal, vol. 6, no. 6, pp. 10 700–10 714, 2019]. Although contract theory has been fully applied in existing technologies, the application of diffusion-based contract theory in incentivizing data sharing remains unexplored. This means that there is still room for development and innovation in solving the incentive problem in data sharing and further improving the efficiency and effectiveness of data sharing.
[0004] Retrieval Augmented Generation (RAG) is an emerging technology with outstanding performance advantages. It can integrate additional information sources such as external knowledge bases and use relevant retrieval data in context to enhance user prompts, thereby significantly improving the accuracy and reliability of the output of the Large Language Model (LLM). As an innovative means, RAG enables LLM to obtain the latest information without retraining and produce reliable output through retrieval-based generation. In related research, some authors introduced RAG and proved that it can effectively improve the accuracy and relevance of generated text by incorporating retrieved documents into the generation process [Advances in Neural Information Processing Systems, vol. 33, pp. 9459–9474, 2020.]. At the same time, some authors proposed a hybrid RAG method that combines sentence windows and parent-child methods, and experimentally verified that this method is superior to existing RAG methods [2024 10th International Conference on Web Research (ICWR). IEEE, 2024, pp. 22–26.]. At the same time, RAG has gradually shown good development prospects in the field of healthcare applications. For example, in the interpretation of medical knowledge, it can optimize the interpretation of clinical guidelines for liver disease with the help of external medical knowledge [NPJ Digital Medicine, vol.7, no.1, p.102, 2024.]. In addition, some authors retrieve similar image-text pairs based on image-text comparison similarity, and use the retrieval attention module to fuse the representation of images and questions with the retrieved images and texts, showing effectiveness in simple biomedical visual question answering [Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp.547–556.]. However, although RAG has achieved certain results in many fields, there is still room for further optimization and innovation in the face of specific complex application scenarios, such as medical data management involving different task scenarios.
[0005] Healthcare systems have developed rapidly with the help of cutting-edge technologies such as cloud computing, the Internet of Things, and Artificial Intelligent (AI). The Internet of Medical Things (IoMT) has become a key interconnected architecture, responsible for collecting and transmitting important medical data, and has led to the generation, storage, and analysis of a large amount of healthcare big data, including omics data, clinical records, and electronic health records. Generative Artificial Intelligent (GAI), as a new technology of artificial intelligence, has been widely used in the field of the Internet of Things. Large Language Models (LLM) can achieve general language generation and natural language processing tasks, and Multi-Modal Large Language Models (MLLM) can help patients better understand their health status. Retrieval Augmented Generation (RAG) technology improves the reliability and accuracy of the GAI model by retrieving facts from an external knowledge base, retrieves relevant data based on the similarity of the alignment vector of the query, enhances user prompts, and makes the responses generated by MLLM more accurate and more contextual. The integration of MLLM and RAG has been applied in many fields. In order to make MLLM perform better in different medical tasks, it is necessary to improve data quality and modify internal modules. Summary of the invention
[0006] The purpose of the present invention is to overcome the many challenges faced by the prior art when applying generative artificial intelligence, retrieval enhanced generation and other technologies to medical data management in the medical Internet of Things. Based on the above purpose, the present invention first provides a multimodal large language model data management system in the medical field, and the system includes the following modules:
[0007] (1) a data transmission module, which includes a data transmission module for connecting the main chain as a data service provider, the relay chain as a blockchain management platform, and multiple subchains as data providers, wherein the main chain sends a data collection task to the relay chain, and the subchain receives the task and uploads local multimodal data according to the contract between it and the main chain. After verification, the relay chain transmits the data to the main chain for training. After the training is completed, the main chain transmits the training data to the relay chain, which distributes it to the subchain, and rewards the subchain according to the contract to complete the transaction;
[0008] (2) a data sharing contract customization module, which is used to generate the contract described in the data transmission module. The contract is a contract type improved by the data service provider, taking into account the non-negative utility and optimal utility of the data provider and the maximum expected utility of the data service provider;
[0009] (3) an optimal contract generation module, which is composed of a contract generation network, a contract quality network, a target contract generation network, and a target contract quality network. The target contract generation network is trained by a deep reinforcement learning method based on GDM to generate a contract strategy based on the current environmental state information, and the contract quality network generates the optimal contract strategy after evaluation;
[0010] (4) a multimodal large language model pre-training module, which is installed on the main chain and uses a supervised learning method to train the multimodal data provided by the data provider;
[0011] (5) A hybrid multimodal retrieval enhancement generation module, which is installed on the data provider and is used to store the multimodal large language model data from the main chain training locally. When the data provider uploads the original modal data for task query, the module converts the original modal data into a task query vector and calculates the similarity score between the task query vector and the multimodal large language model database vector by the hybrid multimodal retrieval enhancement generation method, retrieves the vector that best matches the task query and prioritizes it, and optimizes and synthesizes coherent prompts through prompt word engineering, and then inputs the improved prompts into the multimodal large language model training optimization module;
[0012] (6) A multimodal large language model inference optimization module. After receiving an improvement suggestion from a data provider to combine the original multimodal task query with the retrieved multimodal data, the multimodal large language model inference optimization module connects each modal input to its respective pre-trained encoder model and uses an adapter module to unify all processed embedding data. Finally, the pre-trained weights are used to generate corresponding content according to the input and different tasks.
[0013] In an optional technical solution, in the data transmission module, when data transmission requests are made to the main chain, relay chain and sub-chain, the relay chain is provided with a cross-chain mechanism verification procedure. Data transmission can only be carried out after verification is passed. The interaction requests between the data provider and the data service provider are all recorded in the blockchain.
[0014] In an optional technical solution, in the data sharing contract customization module, the utility of the data provider refers to:
[0015]
[0016] Among them, U H (R k ,f k ), represents the utility function of the data provider, k refers to the type of contract k, R k To obtain the corresponding reward, δ k is the data transmission cost, f k The frequency of data update;
[0017] The maximum expected utility of the data service provider is:
[0018]
[0019]
[0020] Among them, the constraints that must be satisfied by the utility are that the utility of the data provider is non-negative, that is, for any contract of type k, it must be satisfied At the same time, the data provider of type k needs to select the corresponding type of contract (f k ,R k ), rather than other contracts The utility of the corresponding type of contract is also greater than or equal to the utility of other contracts, that is, In optimizing the objective function In the s (f,R) is the maximum utility function of the data service provider, Q k is the probability that the data provider chooses a contract of type k, with the constraint that the sum of these probabilities is equal to 1, where β>0, indicating that the same k The associated unit profit, S k is the evaluation function of the MLLM output performance based on the information age,
[0021]
[0022] Where α is the overall zero-shot accuracy of the MLLM. is a medical data quality metric based on AoI, defined as:
[0023]
[0024] in Indicates the maximum allowed value of AoI, represents the average AoI, which is calculated as:
[0025]
[0026] θ mIt represents the length of a single time slot in each data update cycle, and t is the length of a single process time slot.
[0027] In a more preferred technical solution, the optimal contract generation steps of the optimal contract generation module are as follows:
[0028] S401. Initialize playback buffer Contract Generation Network∈ w , Contract Quality Network Target contract generation network ∈′ ω′ , Target Contract Quality Network
[0029] S402. For all loops E max In each training cycle e, a random process is initialized To facilitate the exploration of contract design;
[0030] S403. There are Z in a cycle e max For a single step z, first observe the current state space s z ;
[0031] S404. Set it to Gaussian noise and pass De-noising to generate contract design
[0032] S405. Execution contract design And observe the reward r z ;
[0033] S406. Storage records To playback buffer
[0034] S407. From playback buffer Randomly select a small batch of N records from
[0035] S408. Update contract quality network and the contract generation network ∈ w , and finally update the target contract generation network ∈′ ω′ and Target Contract Quality Network
[0036] ω′←ηω+(1-η)ω′
[0037]
[0038] S409. End the loop and obtain the best contract generation network ∈ w ;
[0039] S410. After the model is trained, the data service provider will generate the optimal contract strategy a based on the state space containing the information of each data provider. 0 .
[0040] After the multimodal large language model training module completes the training, it accesses the MLLM application programming interface through the main chain and the relay chain. The data provider chooses the weight or application programming interface to use the MLLM service according to the capabilities of its own computing equipment.
[0041] In an optional technical solution, the hybrid multimodal retrieval enhancement generation module works as follows:
[0042] S601. Each data provider uses a corresponding embedding model to convert different modal data into vectors specific to each modality according to different data types, and uses a structured query language tool to store these vectors in a local knowledge base;
[0043] S602. When receiving a task query, the hybrid multimodal RAG system converts the query into a vector using the same embedding model as step S601, then calculates the similarity metric function between the task query vector and the vectors in the knowledge base, retrieves the topK vectors with the similarity metric function scores and prioritizes them;
[0044] S603. Based on the similarity metric function scores of the vectors in each knowledge base, the proportion of each similarity metric function is introduced as a ranking consideration to optimize the ranking of step S602;
[0045] S604. After completing the retrieval process, the hybrid multimodal RAG system uses prompt word engineering to optimize and synthesize a coherent prompt that combines the original multimodal task query with the retrieved multimodal data, and then uses this improved prompt as the input of the MLLM;
[0046] S605. After receiving multimodal input, MLLM connects each modality input to its respective pre-trained encoder model and uses an adapter module to unify all processed embeddings. Finally, MLLM uses pre-trained weights to generate corresponding content based on the input and different tasks.
[0047] More preferably, the optimization of step S603 is to further filter the results using multimodal information similarity:
[0048]
[0049] where f i(·) represents the similarity measurement function between the task query and the source data in the database, which can be freely determined according to the requirements of the specific task. x1 and x2 are the unimodal data and source data corresponding to the task query in the database, respectively. The weight factor w is used. i Characterize each similarity metric function f i (·) when the results are re-ranked and filtered by MIS, and then the context in the cue word is expanded using the retrieved refinement information.
[0050] Secondly, the present invention provides a method for managing medical data using the above-mentioned multimodal large language model data management system, wherein the data service provider is a service center that provides medical data services, and the data provider is a hospital that has data service needs, and the method comprises the following steps:
[0051] S801. Contract generation step: generating an optimal contract type between the service center and the hospital that takes into account the non-negative utility and optimal utility of the hospital and the maximum expected utility of the service center;
[0052] S802. Data upload step: The hospital uploads the medical data of the original modality data to the service center according to the contract type, and obtains corresponding rewards according to the contract type;
[0053] S803. Multimodal large language model pre-training: The service center uses a supervised learning method to train the original modality data uploaded by the hospital, and distributes the training data to the hospital's multimodal vector database;
[0054] S804. Hybrid multimodal retrieval enhancement generation: The hospital performs task queries on its own multimodal vector database based on specific tasks, optimizes the task query results using the hybrid multimodal RAG module and prompt word engineering, and inputs the multimodal large language model;
[0055] S805. Multimodal Large Language Model Optimization Training: MLLM connects each modality input to its respective pre-trained encoder model and uses an adapter module to unify all processed embedding data. Finally, MLLM uses pre-trained weights to generate corresponding content based on the input and different tasks.
[0056] Third, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for managing medical data.
[0057] Finally, the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for managing medical data when executing the program.
[0058] The excellent technical effects of the present invention are embodied in:
[0059] 1. In the prior art, the training of medical multimodal large language models requires hospitals in various places to share a large amount of real-time high-quality medical data, but there is information asymmetry between data holders and service providers, which affects the quality and sharing efficiency of data. The present invention uses a contract model to implement a more efficient data sharing incentive mechanism between the medical health center S and local hospitals H, encouraging local hospitals H to share high-quality and fresh data to solve the problem of information asymmetry. In addition, the information age index is used to evaluate the quality information of both parties to the contract to ensure that both parties to the contract can obtain high-quality rewards or service resources.
[0060] 2. The medical data sharing environment is dynamically changing, and existing technologies may show limitations in dealing with such changes. The present invention designs a flexible contract strategy based on the reinforcement learning algorithm of GDM, which is mainly composed of a contract generation network and a contract quality network. After the generative model has undergone a certain amount of training, the target contract generation network will generate a contract strategy based on the current environmental state information, and the contract quality network will generate the best contract strategy after evaluation, enabling it to adapt to the ever-changing data sharing environment. This adaptability ensures efficient data sharing and utilization in various complex situations, and improves the stability and reliability of the entire system.
[0061] 3. When the medical multimodal large language model (MLLM) processes different types of medical tasks, the model may lead to inaccurate inferences due to data set bias. The present invention greatly improves the diagnostic performance of MLLM for different medical tasks by introducing a hybrid multimodal RAG system and optimizing multimodal input, so that it can retrieve data that meets specific indicators according to the task type and data modality type and apply them to different downstream medical tasks, thereby providing more accurate medical diagnosis and personalized services. At the same time, cross-chain technology and a hybrid multimodal RAG framework are used to achieve safe and efficient medical data management and application. The system uses cross-chain technology to provide security for each data transmission and weight distribution request between local hospitals H and medical and health centers S, jointly ensuring the security and efficiency of the data processing and sharing process, making medical multimodal data tamper-proof and traceable. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 . A schematic diagram of the structure of the medical data management system of the present invention;
[0063] Figure 2 .Schematic diagram of DRL algorithm based on GDM;
[0064] Figure 3 .Schematic diagram of the hybrid multimodal RAG module structure;
[0065] Figure 4 .Performance comparison between different solutions in data sharing;
[0066] Figure 5 .Performance comparison between GDM and DRL-PPO in optimal contract design;
[0067] Figure 6 .Optimal contracts designed by GDM and DRL-PPO;
[0068] Figure 7 .Performance comparison of the management method of the present invention under different medical data cases;
[0069] Figure 8 . Output result diagram of the specific working example of hybrid multimodal RAG. DETAILED DESCRIPTION
[0070] The present invention will be further described below in conjunction with specific embodiments, and the advantages and features of the present invention will become clearer as the description proceeds. However, these embodiments are merely exemplary and do not constitute any limitation on the scope of protection defined by the claims of the present invention.
[0071] Abbreviations:
[0072] AoI: Age of Information;
[0073] API: Programming Interface;
[0074] DRL: Deep reinforcement learning method;
[0075] GDM: Generative Diffusion Model;
[0076] IoMT: Internet of Medical Things;
[0077] LLAVA-Med: Large Language and Visual Assistant for Biomedicine;
[0078] MIS: multimodal information similarity;
[0079] MLLM: Multimodal Large Language Model;
[0080] RAG: Retrieval-augmented generation;
[0081] Example 1. Construction of a multimodal large language model data management system
[0082] 1. System Description
[0083] 1) System initialization: All modal medical data of local hospitals H are stored in a specific database, and local hospitals H and medical and health centers S have partial data information, forming an information environment with asymmetric information. At the same time, medical and health center S builds a blockchain platform based on the information for safe sharing of medical data.
[0084] 2) In the blockchain data sharing platform, the medical and health center S is the service provider, and the local hospitals H are the data providers. Local hospitals H provide medical data to the medical and health center S, which trains the MLLM based on the medical data and returns monetary rewards and weight files or application interfaces to local hospitals H. In the blockchain platform, each hospital H has an account and corresponding assets (e.g., tokens). The medical and health center S and each hospital H send and receive requests through the blockchain platform and use the cross-chain mechanism to verify the requests. All kinds of interaction requests between the medical and health center S and each hospital H are recorded in the blockchain. In order to prevent malicious attackers from destroying the data sharing system, all participants in the data sharing system must obtain personal identity authentication from an authoritative trust agency before entering the system. Only participants who have passed the authentication are allowed to enter the data sharing platform. Verifying the identity of the participants is conducive to further improving the security of data sharing.
[0085] 3) A hybrid multimodal RAG system is installed in each hospital H in each region. The internal data and files of the system are stored locally in each hospital H. While protecting the privacy of sensitive data, it can also help improve the performance of MLLM services.
[0086] 4) Because the present invention is aimed at the medical data sharing scenario of IoMT, and the blockchain platform is used as an auxiliary tool to ensure the security of data transmission, the data sharing platform is carried out in the virtual space.
[0087] 2. System Overview
[0088] (1) Data transmission module
[0089] like Figure 1As shown, in the health center, a robust aggregated MLLM is developed by training a large amount of high-quality multimodal medical data. During data collection, the medical health center S will collect medical care data from hospitals H in different regions, but due to privacy, patient willingness and reward structure, the hospital may not upload all the data, and most hospitals do not have the powerful computing power required to train MLLM, which requires alternative methods. Therefore, in the present invention, cross-chain technology is used to allow hospital H to safely upload medical data and trade with the medical health center S. A robust medical health center S uses the main chain to manage medical care data collection and model updates, and multiple sub-chains handle specific tasks of hospitals H in different regions. The sub-chain uses IoMT to collect real-time patient data and manage related tasks. The specific operation is that the main chain sends a data collection task to the relay chain. After the sub-chain receives the task, hospital H uploads local multimodal medical data according to the contract. The relay chain verifies that the cross-chain request is successful and returns a confirmation to allow upload. After the data is transmitted to the main chain, the medical health center S trains the MLLM. After completion, the medical health center S accesses the MLLM application programming interface (Application Programming Interface, API) through the relay chain and gives monetary compensation according to the data contribution.
[0090] (2) Medical data sharing contract technology module
[0091] In order to develop a robust aggregated MLLM in the medical and health center, a large amount of high-quality multimodal medical data needs to be trained. After the medical and health center S issues the task through the blockchain management platform, considering the information asymmetry between local hospitals H and medical and health center S, hospital H will choose to upload its own multimodal medical data according to the contract provided by medical and health center S. The implementation steps are as follows:
[0092] 1) Define the utility function U of each hospital H H : Hospitals H in various places select one of the k contracts provided by the medical and health center S and receive the corresponding reward R k , considering the data transmission cost δ k With data update frequency f k , the utility function of hospitals in various regions is U H :
[0093]
[0094] 2) Define the MLLM satisfaction function S in the medical and health center S k : In order to evaluate the impact of high-quality real-time data on MLLM performance, the present invention uses an evaluation function based on information age to assess MLLM output performance, which is defined as:
[0095]
[0096] Where α is the overall zero-shot accuracy of the MLLM. is a medical data quality metric based on AoI, defined as:
[0097]
[0098] in Indicates the maximum allowed value of AoI, represents the average AoI, which is calculated as:
[0099]
[0100] Here θ m It represents the length of a single time slot in each data update cycle, and t is the length of a single process time slot.
[0101] 3) Define the utility function U of the medical and health center S s : Due to information asymmetry, the medical and health center S only knows the total number and type distribution of hospitals H in various places, but does not have detailed information about the type of each hospital H.
[0102]
[0103] Where Q k is the probability that hospitals H choose type k contracts, with the constraint that the sum of these probabilities is equal to 1. β>0 means that k Associated unit profit.
[0104] 4) Customized contract constraints: In order to prevent local hospitals H from providing low-quality data in pursuit of higher returns, a robust method is needed to maintain the quality of MLLM services. In this scenario, the medical and health center S takes the lead in designing a set of contract terms and provides them to K hospitals (healthcare data holders). Each hospital H selects the most appropriate contract terms according to its type and uses Indicates that, where f k represents the update frequency of type k healthcare data holder, R k is the reward given to healthcare data holders of type k as an incentive for their contribution. To ensure that each healthcare data holder chooses the most favorable contract item for his type, the designed contract must comply with incentive compatibility (IC) and individual rationality (IR) constraints.
[0105] The IR constraint is the contractual item for the k-type healthcare data holder to guarantee non-negative utility:
[0106]
[0107] The IC constraint is that a healthcare data holder of type k will choose a contract item tailored to its type (f k ,R k ), rather than any other contractual item Right now:
[0108]
[0109] 5) Determine the optimization goal: The present invention hopes to ensure that hospitals H in various places have reasonable utility while maximizing the expected utility of medical and health centers S, that is:
[0110]
[0111] (3) Optimal Contract Generation Module
[0112] Once the contract model is defined, mathematical methods can be used to solve the optimization goal. However, traditional mathematical solutions may not be able to effectively adapt to the dynamic changes of the environment during data sharing. This paper proposes a deep reinforcement learning method (DRL) based on GDM to determine the optimal contract. Compared with the DRL algorithm that directly optimizes model parameters, GDM can enhance contract design through an iterative process of denoising the initial distribution. The diffusion model network maps the environmental state to the contract design, thereby forming a contract design policy expressed in parameters. The specific steps are as follows:
[0113] 1) Initialize the playback buffer Contract Generation Network∈ w , Contract Quality Network Target contract generation network ∈′ ω′ , Target Contract Quality Network
[0114] 2) For all cycles E max In each training cycle e, a random process is initialized To facilitate the exploration of contract design.
[0115] 3) There are Z in a cycle e max For a single step z, first observe the current state space s z . Where s z Contains the environment status at the current step z, such as: M is the total number of data providers, K is the total number of data provider types, is the maximum value of the average AOI, and They represent the contract probability set and contract class set selected by the data provider respectively, and are randomly generated according to the number of noise addition steps T.
[0116] 4) Set it to Gaussian noise and pass De-noising to generate contract design
[0117] 5) Execution contract design And observe the reward r z .
[0118] 6) Storage records To playback buffer
[0119] 7) From the playback buffer Randomly select a small batch of N records from
[0120] 8) Update the contract quality network and the contract generation network ∈ w , and finally update the target contract generation network ∈′ ω′ and Target Contract Quality Network Where η represents the soft update coefficient.
[0121] ω′←ηω+(1-η)ω′
[0122]
[0123] 9) End the cycle and get the best contract generation network ∈ w , the above process is as follows Figure 2 shown.
[0124] 10) After the model is trained, the medical and health center S will i The state space of information is used to generate the optimal contract strategy a 0 .
[0125] (4) Multimodal large language model pre-training module
[0126] The multimodal large language model training module is mounted on the main chain and uses a supervised learning method to train the multimodal data provided by the data provider. During the training process, a segmented training method is used to train the multimodal large language model. First, freeze the language model part, and use medical image data to train the corresponding visual model encoder; secondly, freeze the visual model encoder part, and use medical text information to further incrementally train the pre-trained language model; finally, fine-tune the adapter module uniformly according to the multimodal information such as text-image equivalence, so that the adapter module can process data information of different modalities at the same time. Through supervised learning and segmented training methods, the model can learn the association and mapping relationship between data of different modalities, thereby improving the ability to integrate and understand multimodal information.
[0127] (5) Hybrid multimodal RAG module
[0128] After the medical and health center S uses supervised learning methods to train a multimodal large language model (MLLM), local hospitals H can select weights or application interfaces (APIs) based on their own computing equipment capabilities to use MLLM to serve different medical tasks. The MLLM trained by the medical and health center S has generalized medical capabilities, but it may perform poorly in specific medical tasks due to data set bias. In order to further enhance the output performance of MLLM in different medical tasks, local hospitals H can build a multimodal vector database based on their own data. And use the hybrid multimodal RAG module and prompt word engineering to improve the performance of MLLM in different medical tasks.
[0129] Figure 3 The working principle diagram of the hybrid multi-modal RAG module is given, and the specific operations are as follows:
[0130] 1) Hospitals H in different regions use corresponding embedding models to convert different modality data into vectors specific to each modality according to different data types, and store these vectors in the local knowledge base using structured query language tools.
[0131] 2) When the MLLM receives a task query, the hybrid multimodal RAG system converts the query into a vector using the same embedding model as in step 1). The system then calculates the similarity score between the task query vector and the vectors in the knowledge base, retrieves the topK vectors that best match the task query, and prioritizes them.
[0132] 3) To avoid information overload when all multimodal data are directly input into the MLLM, the hybrid multimodal RAG system further filters the results by applying the multimodal information similarity (MIS) indicator, which is as follows:
[0133]
[0134] where f i (·) represents the similarity measurement function between the task query and the source data in the database, which can be freely determined according to the requirements of the specific task. x1 and x2 are the unimodal data and source data corresponding to the task query in the database, respectively, and the weight factor w is used. i Characterize each similarity metric function f i (·) When the results are re-ranked and filtered by MIS, the context in the prompt is then expanded using the retrieved optimized healthcare information.
[0135] (6) After completing the retrieval process, the hybrid multimodal RAG system adopts cue word engineering to optimize and synthesize a coherent cue that combines the original multimodal task query with the retrieved multimodal medical data, and then uses this improved cue as the input of the MLLM.
[0136] (7) Multimodal large language model reasoning optimization module
[0137] After receiving multimodal input, MLLM connects each modality input to its respective pre-trained encoder model and uses an adapter module to unify all processed embedding data. Finally, MLLM uses pre-trained weights to generate corresponding content based on the input and different tasks.
[0138] Example 2. Comparison of the performance of the contract solution proposed by the present invention, the greedy solution and the random solution
[0139] The contract scheme of the present invention is implemented based on the diffusion model, such as Figure 2 As shown, specifically, the system first initializes the playback buffer, the value network (i.e., the contract quality network), the policy network (i.e., the contract generation network), the target value network (i.e., the target contract quality network), and the target policy network (i.e., the target contract quality network). In a step of a training cycle, the diffusion model-based network is trained according to the environment s z As input, and the contract Set to Gaussian noise, and use the MLP (multi-layer perceptron) in the policy network to denoise and generate contract design After repeating the steps, the environment state (such as: M is the total number of data providers, K is the total number of data provider types, is the maximum value of the average AOI, and The contract probability set and contract class set are selected by the data provider respectively, and are randomly generated according to the number of noise steps T), and the contract, reward and other information are placed in the playback buffer. Then, two gradient-based optimizers (strategy optimizer and value function optimizer) are used to soft-update the policy network and value network according to S408. The value network can verify the action effect of the policy network, and the use of the dual value network can reduce the overestimation of the reward value and improve the learning stability. After the cycle is completed, the target policy network and the target value network can be obtained. Finally, the data service provider will generate the optimal contract strategy a based on the state space containing the information of each data provider. 0 .
[0140] The performance optimization scheme based on contract theory proposed in the present invention is applicable to the scenario including multiple different types of data providers and one data service provider (in this implementation example, we simulate 10 data providers, two different types of data providers). In this scheme, by adopting the contract theory method, using two learning rates of 1×10 -6 The optimizers (strategy optimizer and value function optimizer) train the contract generation network (i.e., strategy network) and the contract quality network (i.e., value network) respectively. The network parameters are updated using the S408 soft update method, and the soft update coefficient is set to 5×10 -3 The data provider’s utility is considered as a reward to motivate them to provide high-quality real-time data.
[0141] In the contract theory scheme of the present invention, the available information of the data service provider is divided into two situations: one is that the data service provider has complete information, and the other is that the data service provider has incomplete information. In the case of information asymmetry, some data providers may obtain more rewards by providing low-quality or delayed data. To prevent this from happening, the present invention designs compatibility constraints (IC constraints) to limit the submission of data that does not meet the requirements. In the case where the data service provider has complete information, since the data information can be perceived in real time, the data service provider does not need to apply IC constraints to achieve higher overall utility.
[0142] In order to verify the performance advantages of the present invention, the present invention compares the scheme based on contract theory with the random selection scheme and the greedy selection scheme. In the random scheme, the data provider randomly selects contracts, resulting in large fluctuations in contract quality and data utility; in the greedy scheme, the data provider selects the contract type with the highest utility. Although the short-term benefits are maximized, the utility in some stages is even lower than that of the random selection scheme because the contract type does not match the actual data situation. In the experiment, a random number range of [-2,2] is used to simulate the utility of the data provider under the random selection scheme; the maximum value of 10 is selected in the greedy algorithm; and the contract theory scheme simulates different types of contract benefits based on the arithmetic progression of [0,4] to more realistically reflect the relationship between contract benefits and utility benefits. The random scheme and the greedy scheme are not trained, and the above rules are used by default in each round of iteration. Figure 4 The performance utility comparison between various schemes is shown in the iterative process of continuous data collection.
[0143] like Figure 4 As shown in the figure, the scheme for dealing with information asymmetry based on contract theory proposed in the present invention always outperforms the random selection and greedy selection schemes in terms of performance. For the contract generation network and the contract quality network, the same learning rate of 1×10 -6 , soft update coefficient 5×10 -3, the number of denoising steps of the diffusion model is 5, and the maximum capacity of the playback buffer is 10 6 , when the training data batch is 512, the utility of the contract theory mechanism (blue) in the present invention is significantly higher than that of the information asymmetry model in the complete information scenario. In particular, in the scenario where the medical and health center S (data service provider) cannot fully understand the types of hospitals H (data providers) in various places, it can still significantly improve the utility.
[0144] Although in theory, a complete information scenario can enable the data service provider to provide the best contract terms by knowing the exact type of each data provider, this is difficult to achieve in real scenarios. Even in a complete information environment, a rational data provider (such as hospital H) may manipulate rewards through misleading data, resulting in a reduction in the subjectivity of the data service provider's utility. Therefore, the contract theory model proposed in the present invention shows higher reliability and practicality in dealing with information asymmetry, and can maximize utility in real scenarios.
[0145] Example 3. Performance of GDM and Proximal Policy Optimization in Optimal Contract Design (DRL-PPO)
[0146] In Example 3, in order to compare the superiority of deep reinforcement learning based on the diffusion model, the present invention compares its performance with the current mainstream deep reinforcement learning method proximal policy optimization (DRL-PPO). The implementation of the scheme based on the diffusion model is the same as that of Example 2, while DRL-PPO uses MLP neural network as decision and value judgment compared to GDM. The environmental parameters and working parameters of the two schemes are consistent, such as: the policy network and the value network use the same learning rate 1×10 -6 , the maximum capacity of the playback buffer is 10 6 , the training data batch is 512, and the number of training cycles is 200. Both train their own policy networks and value networks, using the utility value of the data provider as a reward to encourage it to provide more high-quality real-time data.
[0147] The performance of GDM is compared with that of proximal policy optimization in optimal contract design (DRL-PPO). Figure 5 As shown in the figure, both models are able to continuously obtain rewards in a complex and changing environment until convergence. It is worth noting that under the same parameter settings, the final test return of GDM is significantly higher than that of DRL-PPO, allowing the medical and health center S (data service provider) to always ensure greater utility. This is due to the fine-grained policy adjustment in the diffusion process, which effectively reduces the impact of randomness and noise. Specifically, when solving the optimal goal of the data service provider, the N IR constraints are first simplified according to the recursive relationship of the IC constraints, and at the same time, according to the variable R of the data provider, the N IR constraints are simplified according to the recursive relationship of the IC constraints. k , δk , f k The IC constraints are reduced by the monotonicity and the progressive relationship between upward and downward incentive compatible constraints. The simplified IR constraints and IC constraints are iterated to obtain the optimal reward of the data provider and bring it into the objective function. Finally, the optimal solution is solved through the diffusion denoising process. In addition, the flexibility and robustness of the contract design strategy are enhanced through diffusion exploration to prevent it from falling into suboptimal solutions. Therefore, this superior performance proves the ability of GDM to capture complex patterns and connections between environmental observations, and it can effectively reduce the complexity of the relationship between local hospitals H (data providers) and medical and health centers S. And in Figure 6 The optimal contracts designed by GDM and DRL-PPO are compared in . In this scheme, we consider 10 data providers M, divided into two types K, and set the parameters as M = 10 and K = 2. For the two types of data providers θ1 and θ2, their data provision frequency values are randomly selected between the intervals [1,6] and [13,18] respectively. In addition, the maximum timeliness tolerance The data are randomly generated in the range of [30,60]. To evaluate the utility of the data service provider, the parameters α, β, and t are set to 39.9, 10, and 2, respectively, while Q1 and Q2 are randomly generated according to the Dirichlet distribution. Considering the state of the environment, the GDM-based model, enhanced by exploration during the denoising process, produces an optimal contract design that provides a utility value of 280.85 for the medical health center S, which is higher than the 233.2 achieved by DRL-PPO. This advantage comes from the ability of GDM to generate near-optimal contracts. In addition, as the types of hospitals H increase in various locations, the rewards they receive also increase. However, DRL-PPO shows consistent variations for the types, indicating a tendency towards local optimal solutions. At the same time, the numerical analysis highlights the practical feasibility and superior performance of the proposed GDM-based scheme.
[0148] Scenario description
[0149] 1. Medical data providers (hospitals, clinics, etc.): They have access to a large amount of medical data of patients, but each holder has differences in data quality and frequency.
[0150] 2. Medical data service providers (data analysis companies, AI companies, etc.): Use the data provided by the data holder for modeling and analysis to provide better medical AI services.
[0151] Contract design goals
[0152] 1. Incentives for data providers: Encourage data holders to provide high-quality, real-time updated data.
[0153] 2. Constraints on data service providers: Obtain data that meets the needs at a reasonable cost and maximize service utility.
[0154] Parameter settings
[0155] 1) Number of data providers: M = 10;
[0156] 2) Data provider type: K = 2;
[0157] 3) The data provide frequency values: θ1 and θ2, randomly drawn from the intervals [1,6] and [13,18] respectively;
[0158] 4) Maximum timeliness tolerance: Randomly generated in the range [30,60];
[0159] 5) Multimodal large language model reasoning accuracy: α = 39.9;
[0160] 6) Unit profit related to type: β = 10;
[0161] 7) Data update provides interval slot: t=2.
[0162] Optimal contract generation steps
[0163] 1. According to the utility setting of the present invention, the utility functions of the data provider and the data service provider are set.
[0164] 2. Handling information asymmetry: Since the data service provider cannot directly observe the real data information of the data provider, it is necessary to design IC constraints and IR constraints through contracts to induce the data provider to voluntarily choose high-quality and high-frequency data provision methods while obtaining reasonable benefits.
[0165] 3. Solve the target equation under the premise of satisfying the constraints: Specifically, when solving the optimal goal of the data service provider, first simplify the N IR constraints according to the recursive relationship of the IC constraints, and then simplify the N IR constraints according to the variable R of the data provider. k , δ k , f k The monotonicity of the algorithm and the progressive relationship between upward and downward incentive compatibility constraints are used to reduce IC constraints. The simplified IR constraints and IC constraints are iterated to obtain the optimal reward for the data provider and bring it into the objective function. Finally, the optimal solution is solved through the diffusion denoising process.
[0166] 4. Optimize contract parameters: The data service provider balances the utility between the data provider and the data service provider by adjusting the network parameters of the diffusion model to meet the IR and IC constraints and maximize its own benefits.
[0167] The effect of the contract
[0168] 1. Data holders can choose the optimal data quality and frequency while satisfying their own interests.
[0169] 2. Data service providers can obtain high-quality medical data at a reasonable cost, thereby improving service effectiveness and customer satisfaction.
[0170] 3. This optimal contract design based on contract theory can achieve mutual benefit and win-win results for both parties in an environment of information asymmetry and promote efficient data exchange and cooperation.
[0171] Example 4. LLAVA-Med based on hybrid multimodal RAG and other
[0172] The experiment of the hybrid multimodal RAG system uses LLAVA-Med (for biomedical large-scale language and visual assistants), GPT4-o as the base multimodal large model, and uses indicators such as Response artificial intelligence (RAI) and semantic similarity (SS) as objective criteria, and combines multiple indicators into a unified score - the relative LLM score. The formula is as follows:
[0173] ζ=λ·RAI+ν·SS,
[0174] Among them, the parameters λ and ν are both 0.5.
[0175] like Figure 7 As shown, we compare the performance of the proposed framework under different medical data cases. Our results show that the hybrid multimodal RAG enables LLAVA-Med to consistently score above 0.9, especially in the X-ray cases of users 1 and 2 with known causes, maintaining high-quality answers and stability. In contrast, the other MLLMs show a decrease in output quality due to the interference of disease factors. In the scenarios where users 3 and 4 are normal and have no specific causes, the MLLMs are able to obtain high scores and give reasonable judgments. However, in the case of user 5, who is normal but has an X-ray that can be easily misjudged by doctors, the other MLLMs show a higher misjudgment rate. In contrast, the hybrid multimodal RAG continues to produce high-quality output by matching similar disease conditions, Figure 8A specific practical example is shown. Experimental results show that the data retrieved by Hybrid Multimodal RAG provides valuable information for answering questions. Meanwhile, Table 1 summarizes all the scores, indicating that Hybrid Multimodal RAG helps LLAVA-Med maintain a consistent high score (LLAVA-Med-Hybrid RAG), demonstrating its strong performance in different scenarios. This indicates that Hybrid Multimodal RAG effectively considers the quality of retrieved information by leveraging the features of multimodal data such as images and text. The retrieved relevant healthcare data can help MLLM through contextual relationships, enabling MLLM to provide reliable and robust outputs due to its strong contextual learning ability.
[0176] Table 1: Relative LLM scores of different methods
[0177]
[0178]
Claims
1. A multimodal large language model data management system, characterized in that: The system includes the following modules: (1) a data transmission module, wherein the data transmission module includes a data transmission module for connecting the main chain as a data service provider, the relay chain as a blockchain management platform, and multiple subchains as data providers, wherein the main chain sends a data collection task to the relay chain, and the subchain receives the task and uploads local multimodal data according to the contract between it and the main chain. After verification, the relay chain transmits the multimodal data to the main chain for training. After the training is completed, the main chain transmits the training data to the relay chain, which distributes it to the subchain, and rewards the subchain according to the contract to complete the transaction; (2) a data sharing contract customization module, which is used to generate the contract described in the data transmission module. The contract is a contract that takes into account the non-negative utility and optimal utility of the data provider and the maximum expected utility of the data service provider among the contract types provided by the data service provider; (3) an optimal contract generation module, which is composed of a contract generation network, a contract quality network, a target contract generation network, and a target contract quality network. The target contract generation network is trained by a deep reinforcement learning method based on GDM. The target contract generation network generates a contract strategy based on the current environmental state information, and the contract quality network generates the optimal contract strategy after evaluation. (4) a multimodal large language model pre-training module, which is installed on the main chain and uses a supervised learning method to train the multimodal data provided by the data provider; (5) A hybrid multimodal retrieval enhancement generation module, which is installed on the data provider and is used to store the multimodal large language model data from the main chain training locally. When the data provider uploads the original multimodal data for task query, the module converts the original multimodal data into a task query vector, and calculates the similarity score between the task query vector and the multimodal large language model database vector by the hybrid multimodal retrieval enhancement generation method, retrieves the vector that best matches the task query and prioritizes it, and optimizes and synthesizes coherent prompts through prompt word engineering, and then inputs the improved prompts into the multimodal large language model reasoning optimization module; (6) A multimodal large language model inference optimization module. After receiving an improvement suggestion from a data provider to combine the original multimodal task query with the retrieved multimodal data, the multimodal large language model inference optimization module connects each modal input to its respective pre-trained encoder model and uses an adapter module to unify all processed embedding data. Finally, the pre-trained weights are used to generate corresponding content according to the input and different tasks.
2. The multimodal large language model data management system according to claim 1, characterized in that: In the data transmission module, when data transmission requests are made to the main chain, relay chain and sub-chain, the relay chain is equipped with a cross-chain mechanism verification procedure. Data transmission can only be carried out after verification. The interaction requests between the data provider and the data service provider are all recorded in the blockchain.
3. The multimodal large language model data management system according to claim 1, characterized in that: In the data sharing contract customization module, the utility of the data provider refers to: Among them, U H (R k ,f k ), represents the utility function of the data provider, k refers to the type of contract k, R k To obtain the corresponding reward, δ k is the data transmission cost, f k The frequency of data update; The maximum expected utility of the data service provider is: Among them, the constraints that must be satisfied by the utility are that the utility of the data provider is non-negative, that is, for any contract of type k, it must be satisfied At the same time, the data provider of type k needs to select the corresponding type of contract (f k ,R k ), rather than other contracts The utility of the corresponding type of contract is also greater than or equal to the utility of other contracts, that is, In optimizing the objective function In the s (f,R) is the maximum utility function of the data service provider, Q k is the probability that the data provider chooses a contract of type k, with the constraint that the sum of these probabilities is equal to 1, where β>0, indicating that the same k The associated unit profit, S k is the evaluation function of the MLLM output performance based on the age of information AoI, where α is the overall zero-shot accuracy of the MLLM, where is a data quality metric based on AoI, defined as: in Indicates the maximum allowed value of AoI, represents the average AoI, which is calculated as: θ m It represents the length of a single time slot in each data update cycle, and t is the length of a single process time slot.
4. The multimodal large language model data management system according to claim 3, characterized in that: The optimal contract generation steps in the optimal contract generation module are as follows: S401. Initialize playback buffer Contract Generation Network∈ w , Contract Quality Network Target contract generation network ∈′ ω′ , Target Contract Quality Network S402. For all loops E max In each training cycle e, a random process is initialized To facilitate the exploration of contract design; S403. There are Z in a cycle e max For a single step z, first observe the current state space s z ; S404. Set it to Gaussian noise and pass De-noising to generate contract design S405. Execution contract design And observe the reward r z ; S406. Storage records To playback buffer D; S407. From playback buffer Randomly select a small batch of N records from S408. Update contract quality network and the contract generation network ∈ w , and finally update the target contract generation network ∈′ ω′ and Target Contract Quality Network ω′←ηω+(1-η)ω′ S409. End the loop and obtain the best contract generation network ∈ w ; S410. After the model is trained, the data service provider will generate the optimal contract strategy a based on the state space containing the information of each data provider. 0 .
5. The multimodal large language model data management system according to claim 1, characterized in that: After the multimodal large language model training module completes the training, it accesses the MLLM application programming interface through the main chain and the relay chain. The data provider chooses the weight or application programming interface to use the MLLM service according to the capabilities of its own computing equipment.
6. The multimodal large language model data management system according to claim 1, characterized in that: The working steps of the hybrid multimodal retrieval enhancement generation module are as follows: S601. Each data provider uses a corresponding embedding model to convert different modal data into vectors specific to each modality according to different data types, and uses a structured query language tool to store these vectors in a local knowledge base; S602. When receiving a task query, the hybrid multimodal RAG system converts the query into a vector using the same embedding model as step S601, then calculates the similarity metric function between the task query vector and the vectors in the knowledge base, retrieves the topK vectors with the similarity metric function scores and prioritizes them; S603. Based on the similarity metric function scores of the vectors in each knowledge base, the proportion of each similarity metric function is introduced as a ranking consideration to optimize the ranking of step S602; S604. After completing the retrieval process, the hybrid multimodal RAG system uses prompt word engineering to optimize and synthesize a coherent prompt that combines the original multimodal task query with the retrieved multimodal data, and then uses this improved prompt as the input of the MLLM; S605. After receiving multimodal input, MLLM connects each modality input to its respective pre-trained encoder model and uses an adapter module to unify all processed embeddings. Finally, MLLM uses pre-trained weights to generate corresponding content based on the input and different tasks.
7. The multimodal large language model data management system according to claim 6, characterized in that: The optimization of step S603 is to further filter the results using multimodal information similarity: where f i (·) represents the similarity measurement function between the task query and the source data in the database, which can be freely determined according to the requirements of the specific task. x1 and x2 are the unimodal data and source data corresponding to the task query in the database, respectively. The weight factor w is used. i Characterize each similarity metric function f i (·) when the results are re-ranked and filtered by MIS, and then the context in the cue word is expanded using the retrieved refinement information.
8. A method for managing medical data using the multimodal large language model data management system described in any one of claims 1 to 7, characterized in that: The data service provider is a service center that provides medical data services, and the data provider is a hospital that has data service needs. The method includes the following steps: S801. Contract generation step: using a diffusion model-based reinforcement learning method between the service center and the hospital to generate an optimal contract type that takes into account the non-negative utility and optimal utility of the hospital and the maximum expected utility of the service center; S802. Data upload step: The hospital uploads multimodal data to the service center according to the contract type and obtains corresponding rewards according to the contract type; S803. Multimodal large language model pre-training step: The service center uses a supervised learning method to train the multimodal data uploaded by the hospital, and distributes the pre-trained weight file of the multimodal large language model to the hospital data service platform; S804. Hybrid multimodal retrieval enhancement generation step: The hospital performs task queries on the multimodal vector database built by itself based on specific tasks, optimizes the task query results using the hybrid multimodal RAG module and prompt word engineering, and inputs the multimodal large language model; S805. Multimodal Large Language Model Optimization Inference Optimization Steps: MLLM connects each modal input to its respective pre-trained encoder model and uses an adapter module to unify all processed embedding data. Finally, MLLM uses pre-trained weights to generate corresponding content based on the input and different tasks.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for managing medical data as claimed in claim 8 is implemented.
10. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method for managing medical data according to claim 8 when executing the program.
Citation Information
Patent Citations
Offline reinforcement learning method based on diffusion model
CN117669689A
Smart contract security analysis method and system based on multi-modal technology
CN116958767A
Medical knowledge relation extraction method and system based on large language model fine tuning and retrieval enhancement generation
CN118569263A
Multi-modal data fusion control method and device, equipment and medium
CN118734250A
Cited By
Multi-source heterogeneous data multi-mode mixed retrieval method and system based on large model reasoning
CN120509496A