A method and device for constructing an identification set of multimodal digital content on a chain
Through multimodal models and smart contract monitoring technology, a dynamic multimodal digital content identification set is constructed, which solves the problem of multimodal content monitoring and classification in the blockchain environment and realizes efficient understanding and classification of various content forms on the blockchain.
Patent Information
- Application Number
- CN202410527639.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-29
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-04-29
AI Technical Summary
Existing technologies are unable to effectively handle the dynamic changes and diversity of multimodal digital content in blockchain environments, resulting in difficulties in content monitoring and classification.
Using multimodal models and smart contract event monitoring technology, a dynamic multimodal digital content identification set is constructed through cross-modal conversion, text feature extraction, dimensionality reduction processing and cluster analysis, and KeyBERT and TF-IDF are used to extract keywords for topic classification.
It achieves a comprehensive understanding and accurate classification of multimodal content, and is able to monitor and respond to dynamic changes in digital content on the blockchain in real time, adapting to technological progress and content evolution.
Smart Images

Figure CN118427348B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computers, artificial intelligence, and blockchain technology, and in particular to a method, device, storage medium, and electronic device for constructing an identification set of multimodal digital content on a chain. Background Art
[0002] In the current digital age, emerging blockchain technologies such as Matic, Ethereum, and Optimism have revolutionized decentralized asset trading and data storage. Leveraging a globally distributed network of nodes, these technologies decentralize data storage, bringing unprecedented freedom and security to digital content. However, this decentralized structure also poses significant challenges for content regulation. Of particular concern is the dynamic evolution of digital content on blockchains, including media files. Initially compliant content can become inappropriate or inappropriate over time, potentially triggering legal and ethical issues and even social unrest. Current content monitoring technologies, which are mostly based on static sets of content identifiers, struggle to adapt to the highly virtual, diverse, and dynamic nature of content in blockchain environments.
[0003] Despite the numerous advantages offered by blockchain technology, existing content monitoring methods struggle to cope with the dynamic and diverse content found in blockchain environments. A key issue is that these methods rely on static sets of content identifiers and lack adaptability to the rapid evolution of emerging digital content. Furthermore, existing techniques are unable to effectively handle multimodal content, such as images, video, and audio, limiting their application for comprehensive content understanding and classification.
[0004] Current technology does not fully support multimodal content. At the same time, most existing technologies are general frameworks and do not achieve good results in processing new types of on-chain digital content. Summary of the Invention
[0005] In response to the current problem of missing content identification on the new digital chain, the present invention proposes a targeted multimodal digital content identification set discovery and construction technology.
[0006] Specifically, if Figure 3 As shown, the present invention proposes a method for constructing an identification set of multimodal digital content on a chain, which includes:
[0007] Smart contract event monitoring step 1: Use the blockchain interaction library to monitor metadata update events in the target smart contract. When the metadata of the target smart contract is updated, perform the cross-modal conversion step;
[0008] In the cross-modal conversion step 2, a multimodal model is used to analyze the updated metadata and convert it into a text description.
[0009] Text feature extraction step 3, extracting semantic information of the text description and the metadata to obtain its high-dimensional semantic features;
[0010] Dimensionality reduction processing step 4, converting the high-dimensional semantic features into a low-dimensional space to obtain low-dimensional semantic features;
[0011] Cluster analysis step 5: clustering the semantic information of the text description and the metadata according to the low-dimensional semantic features to obtain multiple clusters, each cluster corresponding to a new topic;
[0012] In the topic extraction and classification step 6, keywords are extracted from each cluster as the topic content of the new topic, and each new topic and its corresponding topic content are used as the digital content identification set construction result of the target smart contract.
[0013] The method for constructing an identification set of on-chain multimodal digital content, wherein the multimodal model is the closed-source model GPT-4, Gemini and the open-source QwenVL, LLaVA.
[0014] The method for constructing an identification set of on-chain multimodal digital content, wherein the metadata includes images, and / or text, and / or videos.
[0015] In the method for constructing an identification set of multimodal digital content on a chain, the subject extraction and classification step includes:
[0016] Based on the TF-IDF score, multiple candidate keywords for each new topic are extracted. The features of the candidate keywords are fine-tuned using KeyBERT to obtain candidate features for each candidate keyword. Representative documents are extracted for each new topic, and the representative documents are embedded and averaged to create the topic features of the new topic. Multiple candidate keywords with the highest similarity between the candidate features and the topic features are selected as the topic content of the new topic.
[0017] The present invention also proposes a device for constructing an identification set of multimodal digital content on a chain, which includes:
[0018] The smart contract event monitoring module uses the blockchain interaction library to monitor metadata update events in the target smart contract. When the metadata of the target smart contract is updated, the cross-modal conversion module is executed;
[0019] The cross-modal conversion module uses a multimodal model to analyze the updated metadata and convert it into a text description;
[0020] A text feature extraction module extracts the semantic information of the text description and the metadata to obtain its high-dimensional semantic features;
[0021] The dimensionality reduction processing module converts the high-dimensional semantic features into a low-dimensional space to obtain low-dimensional semantic features;
[0022] A cluster analysis module clusters the semantic information of the text description and the metadata according to the low-dimensional semantic features to obtain multiple clusters, each cluster corresponding to a new topic;
[0023] The topic extraction and classification module extracts keywords from each cluster as the subject content of the new topic, and uses each new topic and its corresponding subject content as the digital content identification set construction result of the target smart contract.
[0024] The device for constructing an identification set of multimodal digital content on the chain, wherein the multimodal model is the closed-source model GPT-4, Gemini and the open-source QwenVL, LLaVA.
[0025] The device for constructing an identification set of multimodal digital content on the chain, wherein the metadata includes images, and / or text, and / or videos.
[0026] The apparatus for constructing an identification set of multimodal digital content on the chain, wherein the subject extraction and classification module includes:
[0027] Multiple candidate keywords for each new topic are extracted based on the TF-IDF score. The features of the candidate keywords are fine-tuned using KeyBERT to obtain candidate features for each candidate keyword. Representative documents are extracted for each new topic, and the representative documents are embedded and averaged to create the topic features of the new topic. The candidate keyword with the highest similarity between the candidate feature and the topic feature is selected as the topic content of the new topic.
[0028] The present invention also proposes an electronic device, which includes the above-mentioned device for constructing an identification set of multimodal digital content on a chain.
[0029] The electronic device is connected to an information display device, which is used to display the construction results of the digital content identification set using display parameters and attributes set by the user or through an artificial intelligence model.
[0030] The present invention also proposes a storage medium for storing a computer program for executing the method for constructing an identification set of multimodal digital content on the chain.
[0031] From the above scheme, it can be seen that the advantages of the present invention are:
[0032] 1. Multimodal content understanding
[0033] This invention uses advanced multimodal models to understand and analyze multiple types of digital content, including text, images, video, and audio. This multimodal approach enables more comprehensive content analysis, capturing key information across diverse content formats and enabling more accurate topic construction and content categorization.
[0034] 2. Dynamic Evolution and Contract Monitoring
[0035] Smart contract event monitoring technology enables this patent to monitor and respond to dynamic changes in NFT and other digital asset metadata on the blockchain in real time, thereby achieving efficient and accurate management of decentralized digital content.
[0036] 3. Adaptive topic construction and classification
[0037] The patented technical architecture is highly modular and flexible, allowing each step to be optimized or replaced based on technological developments and specific needs. This not only ensures the continued effectiveness of the method, but also enables it to adapt to future technological advancements and the continuous evolution of new on-chain content.
[0038] 4. Disassembly and technological updateability
[0039] Combining multimodal understanding and topic analysis, this technology can adaptively discover and construct emerging content themes and meticulously categorize these themes. This adaptive capability is achieved by continuously learning from the dynamically evolving blockchain content, ensuring the timeliness and effectiveness of the technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 An overall flow chart for executing the present invention;
[0041] Figure 2 is an execution flow chart of a specific embodiment of the present invention;
[0042] Figure 3 Flow chart of the method of the present invention;
[0043] Figure 4 This is a module diagram of the device of the present invention;
[0044] Figure 5 This is a schematic structural diagram of a first electronic device of the present invention;
[0045] Figure 6 This is a schematic diagram of the application environment structure of the first electronic device of the present invention;
[0046] Figure 7 This is a schematic structural diagram of a second electronic device according to the present invention.
[0047] Reference numerals:
[0048] 1-Smart contract event monitoring steps;
[0049] 2- Cross-modal conversion step;
[0050] 3-Text feature extraction step;
[0051] 4-Dimensionality reduction processing step;
[0052] 5-Cluster analysis steps;
[0053] 6-Topic extraction and classification step;
[0054] 100-Smart contract event monitoring module;
[0055] 200-cross-modal conversion module;
[0056] 300-text feature extraction module;
[0057] 400- Dimensionality reduction processing module;
[0058] 500-cluster analysis module;
[0059] 600-Topic extraction and classification module;
[0060] A-First electronic device;
[0061] B-chain multimodal digital content identification set construction device;
[0062] C-data acquisition equipment;
[0063] D-information display device;
[0064] 1000- second electronic device;
[0065] Ⅰ-computing unit;
[0066] II-ROM;
[0067] III-RAM;
[0068] IV-bus;
[0069] V-interface;
[0070] VI-input unit;
[0071] VII-output unit;
[0072] VIII-Storage medium;
[0073] IX-Communication unit. DETAILED DESCRIPTION
[0074] It should be noted that the processor described in the present invention is the control center of an electronic device and can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).
[0075] Optionally, the processor can perform various functions of the electronic device by running or executing a software program stored in the memory, and calling data stored in the memory.
[0076] In a specific implementation, as an embodiment, the processor may include one or more CPUs. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). Electronic devices may include: servers, desktop computers, laptops, smartphones, tablet computers, embedded computers, etc., wherein the embedded computers include vehicles and robots, etc.
[0077] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.
[0078] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto, and the actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or a combination of certain components, or a different arrangement of components.
[0079] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0080] It should also be understood that the term "and / or" in this document simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " in this document generally indicates an "or" relationship between the related objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.
[0081] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0082] It should also be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0083] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.
[0084] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0085] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0086] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0087] Based on in-depth research into decentralized digital content and its monitoring needs, the inventors recognized that existing technologies failed to adequately address the multimodality and dynamic nature of content in blockchain environments. This insight led them to explore a novel research path: developing a technology for constructing intelligent identifier sets for digital content that can understand and adapt to content evolution. By combining the latest advances in artificial intelligence technologies such as deep learning, natural language processing, and computer vision, this invention aims to create an advanced multimodal content analysis and identifier construction technology capable of monitoring changes in digital content. This technology is capable of processing and understanding not only text but also various forms of content, including images, video, and audio, enabling efficient discovery and construction of identifier sets for the rich and diverse content in blockchain environments.
[0088] Based on this, the present invention proposes a detachable technical architecture of a multimodal large model-topic model. Each module in the architecture can be completed using the most advanced corresponding technology currently available. The present invention names this technology MultiModalKeyTopic Construction (multimodal content key topic construction). First, a multimodal model is used to describe the new on-chain digital image content in detail, and other modalities, such as image, video, and audio to text conversion, are completed. Then, the framework proposed by the present invention is used for topic analysis to discover and construct a set of digital content identifiers, and classify them. The entire process then uses smart contract event monitoring technology to monitor changes in digital content metadata. Whenever digital content is dynamically updated, incremental training models, new data analysis, and new identifier discovery and construction are carried out in a timely manner.
[0089] To illustrate the above-mentioned features and effects of the present invention more clearly and easily, the following embodiments are specifically described below with reference to the accompanying drawings. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are for illustrative purposes only. The scope of protection of the present invention is not limited to the disclosed embodiments; the present invention is defined by the appended claims.
[0090] As the present invention Figure 1 and 2 As shown, the implementation process of this technical solution can be divided into the following main steps:
[0091] Smart Contract Event Monitoring: Leveraging blockchain interaction libraries like Web3.js, we monitor metadata updates in real time within target smart contracts. Once we detect that an NFT (Non-Fungible Token) has been updated, as indicated by the smart contract's token ID, we retrieve the updated metadata. If the digital content of the metadata has changed, the system automatically triggers subsequent processing.
[0092] Cross-modal conversion: Apply a multimodal model to analyze updated non-text digital content (such as images, videos, and audio) and convert it into detailed text descriptions, paving the way for subsequent text analysis. This multimodal model can be a closed-source model such as GPT-4 or Gemini and / or an open-source model such as QwenVL or LLaVA.
[0093] Text embedding extraction: For all text content, whether original text or text obtained through cross-modal conversion, sentence-level embedding extraction is performed using a semantic extraction model such as bge-large-zh-v1.5 to capture rich semantic information.
[0094] Dimensionality reduction: Using dimensionality reduction techniques such as Incremental PCA, high-dimensional text embeddings are effectively converted into low-dimensional space representations, which not only reduces the computational burden but also retains key semantic features, facilitating subsequent steps.
[0095] Cluster analysis: Apply clustering algorithms such as MiniBatchKMeans to effectively cluster text based on embedding after dimensionality reduction. Each cluster represents a potential new topic. This step lays the foundation for topic extraction and classification.
[0096] Topic extraction and classification: Using techniques such as TF-IDF, keywords are extracted from each cluster to clearly define each newly discovered topic. These keywords constitute the identification set of new on-chain digital content and are the core of the discovery results of this technical solution. Specifically:
[0097] This paper uses KeyBERT+TF-IDF to extract keywords. First, the top N keywords for each topic are extracted based on the TF-IDF score. Then, KeyBERT is used to fine-tune the keyword representation and embed these keywords using the same embedding model used to embed documents. Next, representative documents are extracted for each topic, embedded, and averaged to create the topic embedding (features) for that topic. Finally, the similarity between the embedding of the candidate keywords and the topic embedding is compared to select the final keyword.
[0098] Incremental training and topic updates: Combined with smart contract monitoring, this technical solution implements incremental training of topic models. It can update the model based on newly added data, discover and build new topics, and thus flexibly respond to the dynamic evolution of digital content.
[0099] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in conjunction with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.
[0100] like Figure 4 As shown, the present invention also proposes an on-chain multimodal digital content identification set construction device B, which includes:
[0101] The smart contract event monitoring module 100 uses the blockchain interaction library to monitor metadata update events in the target smart contract. When the metadata of the target smart contract is updated, the cross-modal conversion module is executed.
[0102] The cross-modal conversion module 200 uses a multimodal model to analyze the updated metadata and convert it into a text description;
[0103] A text feature extraction module 300 extracts semantic information from the text description and the metadata to obtain high-dimensional semantic features;
[0104] Dimensionality reduction processing module 400 converts the high-dimensional semantic features into a low-dimensional space to obtain low-dimensional semantic features;
[0105] A cluster analysis module 500 clusters the semantic information of the text description and the metadata according to the low-dimensional semantic features to obtain a plurality of clusters, each cluster corresponding to a new topic;
[0106] The topic extraction and classification module 600 extracts keywords from each cluster as the topic content of the new topic, and uses each new topic and its corresponding topic content as the digital content identification set construction result of the target smart contract.
[0107] The device for constructing an identification set of multimodal digital content on the chain, wherein the multimodal model is the closed-source model GPT-4, Gemini and the open-source QwenVL, LLaVA.
[0108] The device for constructing an identification set of multimodal digital content on the chain, wherein the metadata includes images, and / or text, and / or videos.
[0109] The apparatus for constructing an identification set of multimodal digital content on the chain, wherein the subject extraction and classification module includes:
[0110] Multiple candidate keywords for each new topic are extracted based on the TF-IDF score. The features of the candidate keywords are fine-tuned using KeyBERT to obtain candidate features for each candidate keyword. Representative documents are extracted for each new topic, and the representative documents are embedded and averaged to create the topic features of the new topic. The candidate keyword with the highest similarity between the candidate feature and the topic feature is selected as the topic content of the new topic.
[0111] like Figure 5 As shown, the present invention further proposes a first electronic device A in another embodiment, which includes the identification set construction device B of the on-chain multimodal digital content.
[0112] like Figure 6 As shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to collect and monitor whether the metadata in the target smart contract is updated, and the information display device D is used to display the digital content identification set construction results obtained by the analysis of the present invention.
[0113] The information display device D can organize and process the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be manually preset, for example, the data output by the first electronic device A is visually displayed, which can be based on the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll, etc. The user is presented with the key information specified by the user, such as new topics, new topic content keywords, etc. The user can understand this information more promptly without having to visit the secondary page or scroll the page, saving the user's operation. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the user's key information based on the user's previous usage habits, such as viewing time, number of clicks, number of edits, etc., and then automatically present the user with rich and necessary key information.
[0114] The present invention also proposes a storage medium VIII in another embodiment for storing a computer program for executing the identification set construction method of the multimodal digital content on the chain. It should be understood that the storage medium in the embodiment of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM).
[0115] Figure 7 A schematic block diagram of a second electronic device 1000 that can be used to implement an embodiment of the present invention is shown. The second electronic device 1000 electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein. The second electronic device 1000 may be the same as or different from the first electronic device A.
[0116] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory II (ROM) or a computer program loaded from a storage medium VIII into a random access memory (RAM) III. Various programs and data required for the operation of the device 1000 can also be stored in the RAM III. The computing unit I, ROM II, and RAM III are connected to each other via a bus IV. An input / output (I / O) interface V is also connected to the bus IV.
[0117] Multiple components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard and mouse; an output unit VII, such as various types of displays and speakers; a storage medium VIII, such as a magnetic disk and optical disk; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0118] Computing unit I can be various general and / or special processing components with processing and computing capabilities. Some examples of computing unit I include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. Computing unit I performs the various methods and processes described above, such as method steps 1-6. For example, in some embodiments, the method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via ROM II and / or communication unit IX. When the computer program is loaded into RAM III and executed by computing unit I, one or more steps of the method described above can be performed. Alternatively, in other embodiments, computing unit I can be configured to execute the method in any other appropriate manner (e.g., by means of firmware).
[0119] In summary, the present invention is based on the pre-processing of multimodal large models and can be applied to multimodal new on-chain digital content; based on smart contract event monitoring technology, it can monitor changes in digital content metadata; based on the detachability of modules, it can follow new technologies in real time; based on TF-IDF (based on word frequency-inverse document frequency) technology, it can automatically construct and update new on-chain digital content identification sets.
[0120] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.
Claims
1. A method for constructing an identification set of multimodal digital content on a chain, characterized in that: include: The smart contract event monitoring step utilizes the blockchain interaction library to monitor metadata update events in the target smart contract. When the metadata of the target smart contract is updated, the cross-modal conversion step is executed; the metadata can be images, text, and / or videos. In the cross-modal conversion step, a multimodal model is used to analyze the updated metadata and convert it into a text description; A text feature extraction step extracts the semantic information of the text description and the metadata to obtain its high-dimensional semantic features; Dimensionality reduction processing step, converting the high-dimensional semantic features into a low-dimensional space to obtain low-dimensional semantic features; A cluster analysis step of clustering the semantic information of the text description and the metadata according to the low-dimensional semantic features to obtain multiple clusters, each cluster corresponding to a new topic; The topic extraction and classification step extracts keywords from each cluster as the subject content of the new topic, and uses each new topic and its corresponding subject content as the result of constructing the digital content identification set of the target smart contract; The topic extraction and classification steps include: Based on the TF-IDF score, multiple candidate keywords for each new topic are extracted. The features of the candidate keywords are fine-tuned using KeyBERT to obtain candidate features for each candidate keyword. Representative documents are extracted for each new topic, and the representative documents are embedded and averaged to create the topic features of the new topic. Multiple candidate keywords with the highest similarity between the candidate features and the topic features are selected as the topic content of the new topic.
2. The method for constructing an identification set of multimodal digital content on a chain according to claim 1, characterized in that: The multimodal models include closed-source models GPT-4, Gemini and open-source QwenVL, LLaVA.
3. A device for constructing an identification set of multimodal digital content on a chain, characterized in that: include: The smart contract event monitoring module uses the blockchain interaction library to monitor metadata update events in the target smart contract. When the metadata of the target smart contract is updated, the cross-modal conversion module is executed; the metadata can be images, text, and / or videos. The cross-modal conversion module uses a multimodal model to analyze the updated metadata and convert it into a text description; A text feature extraction module extracts the semantic information of the text description and the metadata to obtain its high-dimensional semantic features; The dimensionality reduction processing module converts the high-dimensional semantic features into a low-dimensional space to obtain low-dimensional semantic features; A cluster analysis module clusters the semantic information of the text description and the metadata according to the low-dimensional semantic features to obtain multiple clusters, each cluster corresponding to a new topic; The topic extraction and classification module extracts keywords from each cluster as the subject content of the new topic, and uses each new topic and its corresponding subject content as the result of constructing the digital content identification set of the target smart contract; The topic extraction classification module includes: Based on the TF-IDF score, multiple candidate keywords for each new topic are extracted. The features of the candidate keywords are fine-tuned using KeyBERT to obtain candidate features for each candidate keyword. Representative documents are extracted for each new topic, and the representative documents are embedded and averaged to create the topic features of the new topic. Multiple candidate keywords with the highest similarity between the candidate features and the topic features are selected as the topic content of the new topic.
4. The apparatus for constructing an identifier set for multimodal digital content on a chain as claimed in claim 3, wherein: The multimodal models include closed-source models GPT-4, Gemini and open-source QwenVL, LLaVA.
5. An electronic device, characterized in that: Includes the identification set construction device for on-chain multimodal digital content as described in claim 3 or 4.
6. The electronic device according to claim 5, wherein: The electronic device is connected to an information display device, which is used to display the construction result of the digital content identification set using display parameters and attributes set by a user or through an artificial intelligence model.
7. A storage medium for storing a computer program for executing the method for constructing an identification set of on-chain multimodal digital content as described in claim 1 or 2.
Citation Information
Patent Citations
End-to-end news program structuring method and structuring framework system thereof
CN110012349A
Cross-chain resource mapping and management method and system
CN116361292A