An art digital interactive and auction notarization system based on knowledge distillation
Patent Information
- Application Number
- CN202610870804.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-11
AI Technical Summary
当前,依托人工智能、三维建模及网络平台开展艺术品线上导览、虚拟拍卖与权属验证的技术应用日益广泛,但现有技术方案在实际落地过程中,仍存在诸多亟待解决的技术瓶颈,制约了艺术品数字化交互与拍卖存证的规模化、专业化应用
1.本发明通过知识蒸馏构建轻量化专家模型,将大规模预训练模型的核心知识迁移至小参数网络,有效降低多模态数字人交互所需的算力开销与硬件配置要求,显著减少推理时延,实现低算力终端设备下的实时响应,不仅提高了系统在移动端、边缘端、普通服务器等多样化环境中的部署适配性与运行稳定性,也降低了部署成本,提升了整体运行效率与规模化应用能力。
Smart Images

Figure CN122736723A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of blockchain and e-commerce technology, specifically to a digital interactive and auction evidence storage system for artworks based on knowledge distillation. Background Technology
[0002] With the rapid development of digital technology and the internet economy, digital display, online interaction, and auction transactions of artworks have become important directions for the digital transformation of the cultural industry. Currently, the application of technologies such as artificial intelligence, 3D modeling, and online platforms for online guided tours, virtual auctions, and ownership verification of artworks is becoming increasingly widespread. However, existing technical solutions still face many technical bottlenecks that urgently need to be addressed in practical implementation, which restricts the large-scale and professional application of digital interaction and auction evidence preservation of artworks.
[0003] First, existing highly realistic digital human systems, when combined with large language models, generally suffer from high computational consumption and significant interaction latency. To achieve high-definition rendering, real-time speech synthesis, and motion-driven operation of digital humans, traditional solutions rely on large-scale general-purpose pre-trained models and high-end server clusters. Terminal devices struggle to handle complex computational tasks, resulting in lag and audio-visual asynchrony on the user end. This fails to meet the real-time interaction needs of lightweight scenarios such as mobile devices, limiting the widespread application of digital humans in online art services.
[0004] Secondly, current general-purpose AI models used in the art field lack support from specialized domain expertise. These general-purpose models only possess basic natural language understanding capabilities and have not undergone specialized training in areas such as art authentication, dating, stylistic analysis, and valuation, making it difficult to provide professional answers that meet industry standards. Furthermore, existing models primarily output standardized text responses, failing to replicate the explanation styles and intonations of industry experts. This results in a mechanical and rigid digital human interaction that fails to meet users' demands for a professional and immersive art tour experience.
[0005] Furthermore, the online art market suffers from a lack of trust mechanisms and low automation. Existing online auction platforms largely rely on human customer service and offline expert verification to complete the transaction process. The processes of guiding bidding, verifying authenticity, and confirming the sale involve a high degree of human intervention, resulting in cumbersome and inefficient procedures. Simultaneously, information regarding art ownership, transfer records, and authentication data is mostly stored on centralized servers, posing risks of data tampering and loss. The lack of real-time, transparent, and immutable evidence preservation and traceability mechanisms makes it difficult to guarantee the authenticity and security of online art transactions, hindering the standardized development of the online art auction market. Therefore, a solution is proposed. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a knowledge distillation-based digital interactive and auction evidence preservation system for artworks, which solves the problems mentioned in the background section.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a digital interactive and automated auction evidence storage system for artworks driven by knowledge distillation and multimodal AI, comprising: The underlying data source is used to provide data on the history, authentication standards, and transaction records of artworks; The cognitive computing layer is a lightweight art expert model generated based on knowledge distillation. It connects to the underlying data source and adopts a teacher-student model architecture to transfer the art expertise and expert expression style parameters of the general large model to the lightweight model. The core logic layer, namely the semantic-emotion-action mapping scheduling engine, interacts bidirectionally with the cognitive computing layer to extract the emotional weight of the output text, match the body movement parameters in the preset 3D skill library, and generate multimodal collaborative instructions. The output layer is the driving engine for the 3D simulation digital avatar. It receives instructions from the core logic layer and drives the digital avatar to perform voice narration and body movements, rendering the digital twin of the artwork. The external interface layer includes a blockchain notarization interface and a real-time bidding data interface, which connects to the blockchain platform and the auction bidding data stream to realize the on-chain notarization of transaction data and the synchronization of bidding status. The client-side display layer is used to present the interactive screen of the digital avatar, the digital twin of the artwork, and the blockchain-based evidence information. The system achieves a closed loop of automated guided tours, auction bidding, and ownership verification of artworks through the linkage of the cognitive computing layer, the core logic layer, and the presentation output layer, combined with data interaction from the external interface layer.
[0008] Preferably, a general large-scale pre-trained model is obtained as the teacher model; a feature dataset of the vertical domain of artworks is constructed, including text knowledge, expert explanation audio and body movement data; knowledge distillation training is performed to extract the logical weights, professional knowledge and expert expression style parameters of the teacher model; a lightweight student model with fewer parameters than the teacher model is generated; the expert action features are mapped to a 3D skill index table; and the lightweight student model is deployed to the cloud or edge for real-time inference.
[0009] Preferably, the semantic-emotion-action mapping scheduling engine is configured to: receive the output text of a lightweight art expert model, parse the semantics and quantify the emotion weight; preset a 3D skill library including gazing, appreciation gestures, gavel-dropping, and lecturing postures; establish a mapping relationship between emotion weights, semantic tags and body movement parameters; retrieve matching 3D skills and generate action-driven instructions and speech synthesis parameters.
[0010] Preferably, the automated auction linkage logic is implemented by the core logic layer in conjunction with the real-time bidding data interface, including: The dynamic strategy generation module captures data on bidding frequency, bid amount, and bidding duration, and triggers corresponding skills according to preset thresholds. The state closed-loop execution module monitors the transaction signal and drives the digital substitute to perform the hammer-dropping action and generate a transaction data packet.
[0011] Preferably, the blockchain evidence storage interface of the external interface layer realizes a three-in-one traceability system of chain, person, and thing, including: Retrieve and render the artwork's on-chain hash value, ownership certificate, and circulation record; upload the digital avatar's recommendation, explanation, and auction data to the blockchain for evidence storage; After a transaction is completed, a transaction data packet is sent to the blockchain platform, generating a proof receipt which is then sent back to the client.
[0012] Preferably, the knowledge distillation loss function of the lightweight art expert model consists of three weighted components: KL divergence loss between the teacher model and the student model outputs; Real label supervised loss for art vertical domain datasets; mean squared error loss for expert expression style parameters; By jointly optimizing the loss function, a lightweight model that maintains professional accuracy can be obtained.
[0013] Preferably, the 3D skill library includes a basic action library and a scene-customized action library: A basic motion library, including nodding, shaking the head, pointing gestures, and smiling. A scene-customizable action library, including actions for auction hammer falling, bidding guidance, and authenticity verification; Motion parameters can be adjusted according to the digital avatar's appearance and interactive scenarios.
[0014] Preferably, the data interaction sequence in the automated auction scenario includes: the user initiates a bid signal to the automated auction scheduling system; the scheduling system identifies the bid status and triggers the bid confirmation skill; when the bidding is suspended, the knowledge base is retrieved to execute the value analysis skill; after the transaction signal is triggered, the hammer is dropped and a transaction data packet is generated; the transaction data packet is sent to the blockchain platform to complete the on-chain process; the blockchain returns a proof receipt and renders the proof information on the client.
[0015] Preferably, the virtual exhibition hall in the client display layer supports multi-terminal adaptation: mobile terminal, adopts lightweight rendering mode to display digital twin interaction, artwork information and evidence summary; web terminal, supports high-definition 3D rendering to display interactive area, digital twin detail area, evidence detail area and bidding list area; the virtual exhibition hall is equipped with interactive entry point to support user questions, bidding and information query.
[0016] Preferably, the application methods of the system include: digital art tour: user inquiries, lightweight model generates explanatory text, scheduling engine drives digital stand-ins to perform corresponding actions and display blockchain-stored information; automated auction: digital stand-ins provide opening commentary, respond to bids and adjust their speech and actions, hammer down after a sale and complete data uploading to the blockchain; ownership traceability: user inputs a number or scans a code, the system retrieves blockchain records, and the digital stand-ins provide explanations and visualizations.
[0017] This invention provides a digital interactive and auction evidence preservation system for artworks based on knowledge distillation. It has the following beneficial effects: 1. This invention constructs a lightweight expert model through knowledge distillation, transferring the core knowledge of a large-scale pre-trained model to a small-parameter network. This effectively reduces the computing power and hardware configuration requirements for multimodal digital human interaction, significantly reduces inference latency, and enables real-time response on low-computing-power terminal devices. It not only improves the system's deployment adaptability and operational stability in diverse environments such as mobile terminals, edge terminals, and ordinary servers, but also reduces deployment costs and enhances overall operational efficiency and scalable application capabilities.
[0018] 2. This invention deeply integrates professional knowledge of the vertical field of art and expert expression style parameters into the model training process. Relying on the semantic-emotion-action mapping scheduling mechanism, the digital human can stably output professional explanation content that meets industry standards and synchronously match the corresponding body movement output. This avoids the cumbersome process of relying on manual scripts and manual action editing in traditional solutions, enhances the professionalism, consistency and immersion of digital interaction of art, and improves user experience and service quality.
[0019] 3. This invention deeply integrates digital human interaction behavior, auction process data, and blockchain evidence storage mechanisms to achieve real-time on-chain storage, tamper-proof evidence storage, and full-chain traceability of transaction data, bidding records, interaction content, and ownership information. This effectively reduces manual verification costs, minimizes information tampering and trust risks, and ensures the authenticity, security, transparency, and compliance of the entire online art transaction process, providing reliable technical support for the digital art transaction ecosystem. Attached Figure Description
[0020] Figure 1 This is a diagram illustrating the overall architecture of the present invention; Figure 2 This is a flowchart of the intelligent system data processing of the present invention. Detailed Implementation
[0021] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Example: See appendix Figure 1 - Appendix Figure 2 This invention provides a knowledge distillation-based digital interactive and auction evidence storage system for artworks. The digital interactive and automated auction evidence storage system for artworks based on knowledge distillation and multimodal AI driven by this invention adopts a layered architecture design. It forms a complete technical link from data support, intelligent reasoning, logic scheduling, multimodal output, external connection to terminal display. It can be adapted to cloud deployment, edge deployment or cloud-edge collaborative deployment mode to meet the computing power requirements and response time requirements under different application scenarios.
[0023] During system operation, the underlying data source serves as the data foundation of the entire system, continuously integrating and maintaining multi-dimensional, structured art-related data. This includes not only basic art archive information but also authentication standard texts verified by industry experts, high-definition image comparison data, material composition test reports, historical transaction price curves, auction records, ownership change certificates, and judicial filing materials. All types of data are standardized and stored in a unified data format to ensure data integrity, consistency, and accessibility, providing reliable data support for upper-level intelligent model training and real-time inference.
[0024] The lightweight art expert model constructed by the cognitive computing layer relies on knowledge distillation technology to achieve model compression and knowledge transfer. First, a large-scale pre-trained language model with strong natural language understanding and general knowledge generation capabilities is selected as the teacher model. At the same time, a large-scale and comprehensive feature dataset for the vertical field of art is constructed. This dataset not only contains massive amounts of text knowledge, but also includes professional audio samples of explanations from senior industry appraisal experts, art historians, and senior auctioneers, as well as body motion capture data, intonation feature parameters, and expression style labels. Then, it is jointly trained through the knowledge distillation algorithm. During the training process, KL divergence loss, label supervision loss, and style parameter loss are weighted and fused. While preserving the semantic understanding ability and professional knowledge system of the teacher model, the number of model parameters and computational complexity are significantly reduced. The final lightweight model can run stably on ordinary servers, edge computing devices, and even mobile hardware environments, achieving millisecond-level response speed, while the output content has professional accuracy and expert-level expression style.
[0025] The semantic-emotion-action mapping scheduling engine implemented in the core logic layer serves as a key middleware connecting intelligent reasoning and multimodal output. It receives text content output by a lightweight art expert model in real time, performs deep semantic analysis using natural language processing technology, identifies the emotional tendencies, semantic intentions, and information weights contained in the text, and quantifies and generates corresponding emotional weight values accordingly. The system pre-builds a rich and clearly categorized 3D skill library, which stores a large number of motion capture and optimized body movement parameters, covering basic actions required for daily interactions as well as customized actions required for specific scenarios such as auctions, appraisals, and tracing. The scheduling engine automatically matches the parsed emotional weights and semantic tags to the corresponding body movement parameters through preset mapping rules, generating collaborative instructions that include action sequences, action amplitudes, action durations, and speech synthesis parameters, ensuring the coherence of multimodal output.
[0026] The 3D simulation digital avatar driving engine in the output layer receives collaborative instructions from the core logic layer and drives the digital avatar model to perform corresponding body movements and voice broadcasts in real time. At the same time, it loads and renders a high-precision digital twin of the artwork. The digital avatar model can be customized according to the actual application scenario, including appearance, clothing style, body shape, etc. During the execution of body movements, the details of the movements can be adjusted in real time according to the instruction parameters. The voice broadcast is generated using high-naturalness speech synthesis technology and output synchronously with the body movements. The digital twin is built based on 3D scanning and modeling technology, accurately restoring the appearance, texture details, color characteristics and size proportions of the artwork, providing users with a highly realistic visual experience.
[0027] The external interface layer integrates a blockchain evidence storage interface and a real-time bidding data interface, which are responsible for the data interaction between the system and external platforms. The blockchain evidence storage interface can connect to mainstream consortium blockchain underlying platforms, supporting operations such as data hash calculation, block on-chaining, certificate generation, information retrieval and verification, ensuring that all key data is tamper-proof and traceable. The real-time bidding data interface enables data interoperability with the online auction business platform, acquiring and parsing user bidding information, bidding timestamps, bidding frequency, price change range and other data in real time, providing real-time data input for the automated scheduling of the auction process.
[0028] The virtual exhibition hall interface presented by the client-side display layer supports multi-terminal adaptation and access for mobile devices and web browsers. The mobile terminal adopts lightweight rendering technology to prioritize smooth interaction, with a simple and intuitive interface that highlights the interactive screen of the digital twin, the core information summary of the artwork, and the brief identification of the blockchain evidence. The web terminal supports high-definition 3D rendering and multi-window parallel display, which can simultaneously present the interactive area of the digital twin, the detailed display area of the digital twin of the artwork, the display area of the blockchain evidence details, and the real-time bidding list area. The virtual exhibition hall interface has convenient interactive entry points, allowing users to initiate inquiries, participate in bidding, or query traceability information through various methods such as text input, voice input, and touch screen operation. The system responds quickly, and the interactive experience is smooth and natural.
[0029] In practical applications, when users seek digital guided tours of artworks, they submit queries through the virtual exhibition hall interface regarding the authenticity, creation date, artistic style, craftsmanship, and historical value of the artwork. Upon receiving the request, the underlying data source quickly retrieves relevant archives and authentication data for the corresponding artwork. The lightweight artwork expert model in the cognitive computing layer performs semantic understanding and knowledge reasoning on the query request, generating professional, accurate, and expert-style explanatory text. The scheduling engine in the core logic layer performs semantic and sentiment analysis on the explanatory text, matches corresponding body movement parameters, and generates collaborative instructions. The presentation output layer drives the digital avatar to synchronously execute corresponding body movements and voice explanations, while simultaneously rendering and displaying the digital twin of the artwork. The client interface displays the artwork's corresponding blockchain hash value, ownership certificate number, and key evidence information in real time, completing the entire interactive response process.
[0030] When the system runs the automated virtual auction process, after the auction officially starts, the digital avatar completes the opening introduction and artwork promotion according to the preset process. During the auction, the real-time bidding data interface continuously receives user bidding data and transmits it to the core logic layer. The scheduling engine analyzes data such as bidding frequency, bidding interval, and price increase in real time, dynamically identifying the bidding status. When the bidding is intense and active, the digital avatar's bidding guidance actions and incentive messages are automatically matched and triggered. When the bidding stalls, the knowledge base content is automatically retrieved to generate artwork value analysis text and trigger corresponding explanation actions to guide the bidding to continue. When the bidding ends, the digital avatar immediately executes the standard hammer-down action to confirm the auction result. Subsequently, the system automatically organizes and generates a transaction data package containing artwork information, bidding records, transaction price, transaction timestamp, etc. The data is stored on the blockchain through the blockchain notarization interface. The blockchain platform generates a unique notarization receipt and feeds it back to the system. The client interface updates and displays the transaction result information and blockchain notarization certificate in real time, realizing the fully automated closed-loop operation of the auction process.
[0031] When users need to verify the ownership of an artwork, they can initiate a traceability query request by entering the artwork's unique number or scanning the digital twin's exclusive QR code. After receiving the request, the underlying data source retrieves the artwork's full-cycle historical data from creation, authentication, exhibition, transaction to ownership change. The lightweight artwork expert model in the cognitive computing layer generates a clear and detailed traceability explanation text. The scheduling engine in the core logic layer matches and displays the parameters of the digital twin's body movements, driving the digital twin to complete the traceability explanation and action display. At the same time, the client interface visually presents the artwork's blockchain records, ownership transfer trajectory, authentication reports, and transaction vouchers, achieving transparent display and reliable verification of the artwork's full-chain traceability information.
[0032] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A digital interactive and automated auction evidence storage system for artworks based on knowledge distillation and multimodal AI-driven methods, characterized in that, include: The underlying data source is used to provide data on the history, authentication standards, and transaction records of artworks; The cognitive computing layer is a lightweight art expert model generated based on knowledge distillation. It connects to the underlying data source and adopts a teacher-student model architecture to transfer the art expertise and expert expression style parameters of the general large model to the lightweight model. The core logic layer, namely the semantic-emotion-action mapping scheduling engine, interacts bidirectionally with the cognitive computing layer to extract the emotional weight of the output text, match the body movement parameters in the preset 3D skill library, and generate multimodal collaborative instructions. The output layer is the driving engine for the 3D simulation digital avatar. It receives instructions from the core logic layer and drives the digital avatar to perform voice narration and body movements, rendering the digital twin of the artwork. The external interface layer includes a blockchain notarization interface and a real-time bidding data interface, which connects to the blockchain platform and the auction bidding data stream to realize the on-chain notarization of transaction data and the synchronization of bidding status. The client-side display layer is used to present the interactive screen of the digital avatar, the digital twin of the artwork, and the blockchain-based evidence information. The system achieves a closed loop of automated guided tours, auction bidding, and ownership verification of artworks through the linkage of the cognitive computing layer, the core logic layer, and the presentation output layer, combined with data interaction from the external interface layer.
2. The system according to claim 1, characterized in that, The knowledge distillation process of the cognitive computing layer includes: acquiring a general large-scale pre-trained model as a teacher model; constructing a vertical domain feature dataset for artworks, including textual knowledge, expert audio explanations, and body movement data; performing knowledge distillation training to extract the logical weights, professional knowledge, and expert expression style parameters of the teacher model; generating a lightweight student model with fewer parameters than the teacher model; mapping expert action features to a 3D skill index table; and deploying the lightweight student model to the cloud or edge for real-time inference.
3. The system according to claim 1, characterized in that, The semantic-emotion-action mapping scheduling engine is configured to: receive the text output by the lightweight art expert model, parse the semantics and quantify the emotion weight; preset a 3D skill library including gazing, appreciation gestures, gavel-dropping, and lecturing postures; establish a mapping relationship between emotion weights, semantic tags and body movement parameters; retrieve matching 3D skills and generate action-driven instructions and speech synthesis parameters.
4. The system according to claim 1, characterized in that, The automated auction linkage logic is implemented by the core logic layer in conjunction with the real-time bidding data interface, including: The dynamic strategy generation module captures data on bidding frequency, bid amount, and bidding duration, and triggers corresponding skills according to preset thresholds. The state closed-loop execution module monitors the transaction signal and drives the digital substitute to perform the hammer-dropping action and generate a transaction data packet.
5. The system according to claim 1, characterized in that, The blockchain evidence storage interface of the external interface layer enables three-in-one traceability of the chain, people, and things, including: Retrieve and render the artwork's on-chain hash value, ownership certificate, and circulation record; upload the digital avatar's recommendation, explanation, and auction data to the blockchain for evidence storage; After a transaction is completed, a transaction data packet is sent to the blockchain platform, generating a proof receipt which is then sent back to the client.
6. The system according to claim 2, characterized in that, The knowledge distillation loss function of the lightweight art expert model consists of three weighted parts: KL divergence loss between the teacher model and the student model outputs; Real label supervised loss for art vertical domain datasets; mean squared error loss for expert expression style parameters; By jointly optimizing the loss function, a lightweight model that maintains professional accuracy can be obtained.
7. The system according to claim 3, characterized in that, The 3D skill library includes a basic action library and a scene-customized action library: A basic motion library, including nodding, shaking the head, pointing gestures, and smiling. A scene-customizable action library, including actions for auction hammer falling, bidding guidance, and authenticity verification; Motion parameters can be adjusted according to the digital avatar's appearance and interactive scenarios.
8. The system according to claim 4, characterized in that, The data interaction sequence in the automated auction scenario includes: the user initiates a bid signal to the automated auction scheduling system; the scheduling system identifies the bid status and triggers the bid confirmation skill; when the bidding is paused, the knowledge base is retrieved to execute the value analysis skill; after the transaction signal is triggered, the hammer is dropped and a transaction data packet is generated; the transaction data packet is sent to the blockchain platform to complete the on-chain process; the blockchain returns a proof receipt and renders the proof information on the client.
9. The system according to claim 1, characterized in that, The virtual exhibition hall in the client display layer supports multi-terminal adaptation: mobile terminal, adopts lightweight rendering mode to display digital twin interaction, artwork information and evidence summary; web terminal, supports high-definition 3D rendering, displays interactive area, digital twin detail area, evidence detail area and bidding list area; the virtual exhibition hall has an interactive entry point to support user questions, bidding and information query.
10. The system according to any one of claims 1 to 9, characterized in that, The application methods of the system include: digital art tours: user inquiries, lightweight models generate explanatory text, scheduling engines drive digital avatars to perform corresponding actions and display blockchain-stored information; automated auctions: digital avatars provide opening commentary, respond to bids and adjust their language and actions, and the hammer falls after a sale and the data is uploaded to the blockchain; ownership traceability: users input a number or scan a code, the system retrieves blockchain records, and the digital avatar provides explanations and visualizations.