Multi-modal design work intelligent classification and retrieval system based on deep learning
By integrating image, text, audio and other data through multimodal deep learning technology, a cross-modal correlation analysis system is constructed, which solves the feature extraction and data security issues of the design work classification and retrieval system, and realizes efficient and accurate multimodal design resource management and personalized recommendations.
Patent Information
- Application Number
- CN202510949522.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing design work classification and retrieval systems have limitations in feature extraction and intelligent analysis, making it difficult to fully characterize the complexity of design works. Data security and collaboration issues also restrict the efficient use of design resources, making it impossible to effectively integrate multimodal data for accurate retrieval and management.
It uses multi-scale convolutional neural networks and Transformer architecture to fuse multimodal data such as images, text, and audio, constructs a multimodal knowledge graph for cross-modal association analysis, combines incremental learning and reinforcement learning for classification and recommendation, integrates blockchain technology to ensure data security, and supports cross-regional collaboration and copyright traceability.
It achieves accurate extraction and deep understanding of the multimodal features of design works, improves the accuracy and efficiency of classification and retrieval, supports personalized recommendations and stable system operation, and ensures data security and collaboration efficiency.
Smart Images

Figure CN120654072A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal deep learning retrieval, and in particular to a multimodal design work intelligent classification and retrieval system based on deep learning. Background Art
[0002] With the booming development of the digital design industry, the number of design works has exploded, covering multiple fields such as graphic design, industrial design, and UI design. Designers, companies, and research institutions face the problem of low efficiency when searching for target works within massive design resources. Traditional keyword-based search methods rely on manual annotation, which is not only time-consuming and labor-intensive, but also difficult to accurately capture the deep semantics and stylistic characteristics of design works. For example, when searching for "furniture design with oriental aesthetics," relying solely on text keywords often misses works that are not annotated with relevant terms but actually meet the requirements, resulting in a deviation between the search results and user needs.
[0003] Existing design work classification and retrieval systems face numerous technical bottlenecks. Regarding feature extraction, single-modal analysis methods cannot fully characterize the complexity of design works. For example, it is difficult to understand the design concept conveyed by design description text through image recognition alone, while simple text analysis cannot capture the visual features of the image. This modality-separated processing method leads to information loss and reduces the accuracy of classification and retrieval. At the level of intelligent analysis, traditional machine learning algorithms lack the ability to dynamically learn design trends and user preferences, making it difficult to adapt to the rapidly changing needs of the design field and insufficiently identifying emerging design styles and niche design categories.
[0004] Data security and collaboration issues also hinder the efficient use of design resources. Design works often contain a company's core creativity and trade secrets, and there is a risk of data leakage during the sharing process across teams and institutions. At the same time, the design resource formats and standards are not unified across different platforms, making data interoperability difficult and forming information islands. In addition, existing systems mostly focus on processing static design works, and have weak analysis capabilities for dynamic multimodal data such as design process videos and designer explanation audio, and are unable to fully tap the potential value of design resources. Therefore, there is an urgent need to develop a system that can integrate multimodal data, realize intelligent classification and accurate retrieval, in order to improve the efficiency of design resource management. Summary of the Invention
[0005] The present invention proposes a multimodal design work intelligent classification and retrieval system based on deep learning to solve the problems mentioned in the above-mentioned prior art.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A multimodal design work intelligent classification and retrieval system based on deep learning, including the following modules:
[0008] Multimodal feature extraction module: This module uses a multi-scale convolutional neural network, integrating a visual Transformer and a residual network, to extract image texture, color, and geometric features. It combines BERT with a multi-head attention mechanism to perform text sentiment analysis and keyword extraction. It also uses Mel-frequency cepstral coefficients and spectral centroids to extract innovative audio features for emotion recognition.
[0009] Cross-modal association analysis module: Builds a multimodal knowledge graph, uses graph neural networks to model the semantic associations between images, text, and audio, develops a dynamic attention fusion mechanism, and uses innovative formulas to: Adaptively adjust the weight distribution of different modal features, W t,i,a Represent the feature weights of text t, image i, and audio a respectively; m is the modal variable, S t,i,a is the semantic association score between the corresponding modality and the target task, λ t,i,a is the adjustment coefficient, and at the same time innovates the spatiotemporal attention mechanism to capture the correlation between temporal information and spatial features in design works;
[0010] Multi-dimensional classification engine: Design a hierarchical cascade classifier architecture, using lightweight MobileNetV3 for coarse classification at the bottom layer and integrating Deep Forest and XGBoost algorithms for fine-grained classification at the top layer; develop an incremental learning framework and continuously optimize the classification model through a federated learning mechanism;
[0011] Semantic Enhanced Retrieval Module: Builds a design intent understanding model, combines historical user interaction data with the current search context, and performs fuzzy matching and intent correction of query semantics. It also innovates multimodal hybrid retrieval strategies, supporting cross-modal retrieval methods such as text retrieval from images and image retrieval from text.
[0012] Intelligent recommendation system: Develop a recommendation engine based on reinforcement learning, using the Actor-Critic architecture to dynamically adjust the recommendation strategy. Combined with a design trend prediction model, it enables personalized design recommendations. A design work value assessment module uses an attention mechanism to quantify the contribution of each feature to design quality and generate an interpretable rating report.
[0013] System operation and maintenance monitoring module: Build a digital twin model of device performance to monitor server CPU utilization, memory usage, and network throughput in real time; develop an abnormal behavior detection algorithm and achieve early warning of system failures by combining isolation forests with autoencoders.
[0014] Furthermore, the multimodal feature extraction module also includes: developing a multispectral image analysis unit, integrating RGB, near-infrared, and ultraviolet spectral imaging technologies to extract the material characteristics of the design works; innovating a gesture action recognition module, using a skeleton point detection algorithm and a spatiotemporal graph convolutional network to analyze the semantics of gestures in the design creation process.
[0015] Furthermore, the cross-modal association analysis module also includes: constructing a design style transfer network to achieve automatic conversion of different design styles through a generative adversarial network; developing a multimodal knowledge distillation mechanism to encode expert knowledge into a graph structure representation to achieve cross-domain design knowledge transfer.
[0016] Furthermore, the multi-dimensional classification engine also includes: designing a hierarchical attention classifier, realizing hierarchical representation of design elements through capsule networks, and supporting step-by-step classification from macro style to micro details; innovating a small sample learning module, using prototype networks combined with data enhancement technology to classify rare design categories.
[0017] Furthermore, the semantic enhancement retrieval module further includes: introducing the formula P(I|U,F) refers to the probability of inferring design intent I when the user search behavior U and the design work feature F are known. P(U|I,F) is the conditional probability of U appearing given I and F. P(I,F) is the joint probability of I and F. i traverses all design intents, and the denominator serves as a normalizer. This formula analyzes causal relationships, assists the design intent inference engine in making recommendations, and cooperates with the knowledge question-answering system to serve users.
[0018] Furthermore, the intelligent recommendation system also includes: a design trend prediction unit, which builds a time series prediction model to predict design trends 6 months in advance; a user feedback optimization mechanism, which uses a preference learning algorithm to update user portraits in real time.
[0019] Furthermore, the system operation and maintenance monitoring module also includes: developing an adaptive load balancing algorithm to dynamically adjust server cluster resource allocation based on reinforcement learning; innovating system security protection mechanisms to implement user behavior audits through federated zero-knowledge proofs and verify system compliance without leaking privacy.
[0020] Furthermore, it also includes: a design collaboration workshop module that integrates real-time multi-person annotation tools and version control mechanisms to support collaborative creation by cross-regional design teams; and the development of a design process traceability system that records the creation history and modification trajectory of design works through blockchain.
[0021] Furthermore, it also includes: a design quality assessment module, which builds a multi-dimensional evaluation index system, combines expert knowledge graphs with deep reinforcement learning, and generates explainable design improvement suggestions; an energy consumption optimization unit, which uses neural architecture search to automatically design low-power model architectures.
[0022] Furthermore, it also includes: edge inference acceleration module, introducing innovative formula: Among them, FPS is the number of frames processed per second, C peak It is the peak computing power of the dedicated neural network processor, U util is the hardware utilization, Q bit is the number of quantization bits, N params is the model parameter, C complex To calculate the complexity coefficient; at the same time, the multimodal interaction interface supports voice commands, gesture control, and eye tracking.
[0023] Compared with the existing technology, the beneficial effects of the present invention are:
[0024] The system uses a multi-scale convolutional neural network and Transformer architecture to integrate multimodal data such as images, text, and audio. It can accurately extract texture, semantics, emotion and other features of design works, avoiding the limitations of single modal analysis and making the identification of design features more comprehensive and accurate.
[0025] The cross-modal association analysis module achieves deep fusion and semantic alignment of data from different modalities by constructing a knowledge graph and a dynamic attention mechanism, effectively resolving the problem of information fragmentation between modalities. This enables the system to not only understand the surface features of design works but also to explore their underlying design concepts and stylistic connections, providing designers with more creative and inspiring search results. The multi-dimensional classification engine utilizes a hierarchical cascade architecture and incremental learning technology to maintain classification speed while rapidly adapting to the emergence of new design categories. It supports the detailed classification of over a thousand design categories, significantly improving classification efficiency and accuracy.
[0026] The semantically enhanced retrieval module combines historical user behavior and design intent reasoning to support fuzzy queries, error correction, and cross-modal retrieval, significantly improving retrieval accuracy and flexibility. The intelligent recommendation system, based on reinforcement learning and trend prediction models, provides personalized recommendations based on user preferences and changing design trends, helping designers stay abreast of industry trends and gain inspiration. Furthermore, the system operation and maintenance monitoring module ensures stable system operation through digital twins and intelligent early warning technologies. The application of blockchain technology ensures data security and copyright traceability for design works, providing strong support for efficient collaboration and innovative development in the design industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a schematic diagram of the deep learning-based multimodal design work intelligent classification and retrieval system proposed in the present invention;
[0028] Figure 2 This is a performance comparison chart between the semantic enhancement retrieval module and the traditional method;
[0029] Figure 3 This is a line chart comparing the performance of the multimodal feature extraction module. DETAILED DESCRIPTION
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0031] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0032] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a connection between the two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below with reference to the accompanying drawings.
[0033] Reference Figures 1 to 3 : A multimodal design work intelligent classification and retrieval system based on deep learning, including the following modules:
[0034] Multimodal Feature Extraction Module: This module extracts feature information of design works from multiple dimensions, such as images, text, and audio, providing basic data support for subsequent classification, retrieval, and other operations. This module uses a multi-scale convolutional neural network (MS-CNN) architecture, integrating a visual Transformer with a residual network. The MS-CNN network processes features of different scales in parallel using a three-branch structure. The shallow branch uses a 3×3 convolution kernel to accurately capture the detailed texture of the design work with an accuracy of 0.08mm. The middle branch uses a 5×5 convolution kernel to extract local morphological features, with a measurement error of ≤0.03mm for the radius of the rounded corners of the graphic. The deep branch uses a 7×7 convolution kernel to analyze global structural features. The color space conversion module supports multiple color models such as RGB, HSV, and CIELAB. At the same time, the module integrates a multispectral image analysis unit, fusing imaging technologies of three bands: RGB, near-infrared (700-1100nm), and ultraviolet (200-400nm). By analyzing the spectral information of different bands, it extracts the material characteristics of the design works and supports the recognition of more than 200 materials. The developed text semantic enhancement module combines the BERT pre-training model with the multi-head attention mechanism. The BERT pre-training model processes the design description text to mine deep semantic vectors, and the bidirectional LSTM captures the temporal dependency of the text sequence. The keyword extraction link adopts the fusion of TF-IDF and TextRank algorithms, and the F1 value on the test dataset reaches 0.88. The innovative audio feature extraction algorithm leverages Mel-Frequency Cepstral Coefficients (MFCC) and spectral centroid analysis, employing a dual-channel processing architecture. The MFCC channel extracts 13-dimensional feature parameters to capture the pitch and timbre characteristics of speech, while the spectral centroid channel analyzes the audio energy distribution to identify the emotional state of speech. The innovative gesture recognition module utilizes a skeleton point detection algorithm (an improved version of OpenPose) and a spatiotemporal graph convolutional network. The skeleton point detection algorithm optimizes the network structure, increasing the accuracy of hand key point detection to 0.93 (AP@0.5 indicator). The spatiotemporal graph convolutional network analyzes the dynamic changes of skeleton points in time and space, accurately parsing the semantics of gestures during the design creation process. Through the coordinated operation of these components, the multimodal feature extraction module comprehensively and accurately extracts the multimodal features of design works.
[0035] Cross-modal association analysis module: Constructs a multimodal knowledge graph with a three-layer architecture: the bottom layer is the entity layer, which is used to store basic entity information such as design elements, materials, and styles; the middle layer is the relationship layer, which uses graph neural networks (GNN) to learn the semantic associations between entities; the top layer is the reasoning layer, which implements cross-modal knowledge reasoning based on the rule engine. When modeling the semantic associations between images, text, and audio, a dynamic attention fusion mechanism is developed, and through innovative formulas Adaptively adjust the weight distribution of different modal features. In the formula, W t,i,aRepresent the feature weights of text (t), image (i), and audio (a) respectively; m is the modal variable, and its value range is {t,i,a}; S t,i,a is the semantic association score between the corresponding modality and the target task. The higher the value, the more critical the modality is to the task. t,i,a is a dynamic adjustment coefficient ranging from 0.1 to 1.0, dynamically optimized based on real-time computational complexity. A design style transfer network was constructed based on a generative adversarial network (GAN). This improved upon the StyleGAN2 architecture by introducing a content encoder and decoder to decouple style from content. For example, when migrating a modern minimalist style to a classical design work, the generated image's FID score met the requirements, approaching the creative level of a human designer. A multimodal knowledge distillation mechanism was developed, employing a teacher-student network architecture. The teacher network encodes expert-annotated design knowledge into a graph structure representation, and then transfers this encoded knowledge to the student network through knowledge graph embedding technology. In knowledge transfer experiments in the furniture design field, this mechanism enabled the student network to converge faster than traditional methods, significantly improving classification accuracy. An innovative spatiotemporal attention mechanism adopts a dual-stream network architecture. The spatial stream is mainly responsible for analyzing the static spatial features of design works, while the temporal stream focuses on processing dynamic changes in video sequences. Taking the analysis of industrial design process videos as an example, this mechanism can accurately capture the evolution of designers' creative ideas, provide strong support for the dynamic semantic analysis of video design works, and also improve the accuracy of cross-modal retrieval.
[0036] Multi-dimensional classification engine: The hierarchical cascade classifier adopts a three-level architecture. The first-level MobileNetV3 uses depthwise separable convolution to greatly reduce the number of parameters. When processing images of design works, it first performs operations such as channel-by-channel convolution and point convolution on the input image to extract basic features. Then, through global average pooling and fully connected layers, it quickly screens out major categories such as graphic design and product design at a rate of 1,200 images per second. The second-level ResNet50 uses residual learning units, taking the first-level output features as input. The feature maps are sequentially subjected to operations such as convolution, batch normalization, and ReLU activation functions in multiple residual blocks to continuously extract deep-level features. For example, works initially classified as graphic design are further subdivided into middle categories such as poster design and packaging design. Finally, the classification results are output through a fully connected layer and a Softmax function, with an accuracy rate of 94%. The third-level deep forest model consists of multiple completely random forests and cascade forests. The completely random forest randomly samples feature subsets and sample subsets of the second-level output feature vector to construct multiple decision trees. The obtained classification results are input into the cascade forest. The cascade forest further learns the relationship between these results and iterates multiple times to output fine-grained classification results such as retro posters and minimalist packaging. It supports 1024 design categories. In the ImageNet dataset test, this architecture significantly reduced the amount of computation while ensuring accuracy.
[0037] The incremental learning framework, based on a federated learning mechanism, eliminates the need to retrain the entire model when new data, such as environmental design samples, is added. Instead, it employs a model parameter aggregation strategy. Specifically, local model update parameters from various participants (such as different data sources or computing nodes) are first collected and aggregated through methods such as weighted averaging. These aggregated parameters are then applied to the global model. This completes the model update in just two hours, maintaining high classification accuracy.
[0038] The hierarchical attention classifier uses capsule networks to achieve a hierarchical representation of design elements. Bottom-level capsules represent basic elements like lines and shapes, and their states are determined by extracting and encoding the features of the input design data. Middle-level capsules combine the outputs of bottom-level capsules to form components like icons and text. This process, based on a specific transformation matrix and dynamic routing algorithm, enables middle-level capsules to learn the feature representations of components. High-level capsules are then combined to form complete design units.
[0039] The innovative small-sample learning module uses prototypical networks combined with data augmentation technology. The prototype network calculates a prototype representation for each design category by averaging the feature vectors of all samples in that category. During classification, the distance (such as the Euclidean distance) between the feature vector of the new sample and each prototype is calculated. The category to which the prototype with the closest distance belongs is the predicted category of the new sample. At the same time, data augmentation technology performs operations such as rotation, scaling, and cropping on a small number of samples to expand the number and diversity of samples. With only 5 samples, the recognition accuracy of new design categories can reach a high level, significantly better than traditional small-sample learning methods.
[0040] Semantic Enhanced Retrieval Module: The design intent understanding model adopts a two-stage architecture. First, the BERT model processes the query text entered by the user. Using a self-attention mechanism, the BERT model understands the semantic connections between words in the text and extracts the text's semantic vector. For example, if a user enters "modern office furniture," the BERT model can accurately grasp the semantic meaning. Next, historical user interaction data, such as past search records and operational behavior on different design elements, is collected and processed into a personalized preference vector. The semantic vector and personalized preference vector are then merged, allowing the system to consider both text semantics and user preferences in fuzzy queries, thereby accurately matching relevant design works.
[0041] The intent correction module combines an edit distance algorithm with a language model. The edit distance algorithm calculates the difference between the user's input text and the vocabulary in the vocabulary database, identifying candidate words with the smallest difference. The language model then uses language knowledge and semantic rules to determine the most reasonable correction result from these candidate words. For example, "light furniture" will be corrected to "light furniture."
[0042] The multimodal hybrid retrieval strategy covers six retrieval modes. Taking text-to-image retrieval as an example, the text is first converted into semantic vectors using a designed intent understanding model. The image is then trained using a convolutional neural network to extract feature vectors. The similarity between the two is then calculated to match related images. The other retrieval modes work similarly, enabling cross-modal retrieval.
[0043] Design intent reasoning engine introduces formula Where P(I|U,F) is the probability of inferring design intent I given user search behavior U and design work feature F; P(U|I,F) is the conditional probability of U appearing given I and F; and P(I,F) is the joint probability of I and F. i traverses all design intents, and the denominator serves as a normalizer. Based on this formula, the system analyzes causal relationships based on user search behavior and design work features to infer design intent. For example, if a user frequently clicks on works labeled "eco-friendly materials," the system infers a preference for sustainable design, thereby optimizing recommendations and increasing the relevance of the results.
[0044] The knowledge question-answering system is based on the RAG framework. When a user asks a question, relevant knowledge fragments are first retrieved from the design knowledge base, and then the language generation model integrates and generates answers to meet the designer's needs for professional knowledge.
[0045] Intelligent Recommendation System: This reinforcement learning-based recommendation engine utilizes an actor-critic architecture. The actor network receives current state information, such as the user's browsing history and the features of design works. Through a series of neural network layer operations, it outputs a list of recommendations. The designs in these recommendations are generated by the actor network based on its understanding of user preferences and analysis of current design resources. The critic network evaluates the recommendations generated by the actor network. It uses user feedback (such as clicks, saves, etc.) after displaying the recommendations as input, and combines it with pre-defined evaluation metrics such as click-through rate and user dwell time to calculate a recommendation score. The critic network then provides feedback to the actor network based on this score, guiding it to adjust parameters and optimize its subsequent recommendation list generation strategy. For design trend prediction tasks, the engine collects data from over one million design works on platforms like Pinterest and Behance, as well as Google Trends search data, and performs pre-processing on this data, including data cleaning and feature extraction. The processed data is then fed into a time series forecasting model, which analyzes the time series characteristics and trend patterns in the data to predict design trends six months in advance.
[0046] In addition to collecting design work data and search data from the above-mentioned platforms, the design trend prediction unit also collects a wide range of data such as user discussions, likes, and shares on social media, as well as professional analysis in industry reports, market research data, and other multi-source information. This multi-source information is integrated and feature-engineered to extract key features related to design trends, such as color preferences, frequency of element use, and popularity of style keywords. These features are then used as input to construct a time series prediction model. This model can be based on a recurrent neural network (RNN) and its variants (such as LSTM, GRU), or a time series model based on the Transformer architecture. By learning from historical data, the model captures the evolution of design trends over time.
[0047] The user feedback optimization mechanism employs an online learning strategy, monitoring user interactions in real time. Once 1,000 pieces of user interaction data (such as clicks, favorites, and ratings) have been collected, a model update is triggered. A preference learning algorithm is then used to analyze this interaction data, uncovering user preferences for various design elements, styles, and features. For example, by analyzing the time series of user clicks, we can determine the changing trends in user interest in different design styles; based on user favorites, we can determine the degree of user preference for specific design elements. Based on these analysis results, user profiles are updated in real time, encompassing multiple dimensions, including basic user information, interests, preferences, and behavioral habits.
[0048] The design work value assessment module uses an attention mechanism to analyze various features of a design. For example, in UI design evaluation, the system first extracts interactive experience-related features such as the smoothness of interactive animations, the rationality of interface layout, and the harmony of color schemes. It also extracts visual aesthetic features such as the aesthetics of interface elements and the consistency of the overall visual style. It also extracts functional implementation-related features such as functional completeness and ease of operation. The attention mechanism then quantifies the contribution of these features to the design quality.
[0049] The system operation and maintenance monitoring module uses ANSYS TwinBuilder to build a digital twin model of device performance. By deploying numerous sensors and data acquisition interfaces on the server side, the model continuously collects real-time data on 15 core performance indicators, including CPU usage, memory usage, network bandwidth, and disk I / O, and continuously transmits this data to the model. The model uses this collected data to simulate the physical system with high precision. When system load approaches 80%, the model leverages built-in predictive algorithms and historical data to predict the server's future operational status, predicting risks such as server overheating 24 hours in advance, enabling operations personnel to proactively respond. For abnormal behavior detection, the system combines isolation forests with autoencoders. Isolation forests, based on the principle that outliers are rare and isolated from normal data points, construct multiple binary trees to learn from normal system operation data (such as user logins and data access patterns). The system then calculates the forest path length for each data point to determine the likelihood of anomalies. The autoencoder learns feature representations by encoding and decoding normal data, and identifies new data with large reconstruction errors as likely anomalies. The system successfully detected multiple SQL injection attacks and malicious data scraping with a false positive rate of only 0.15%. The adaptive load balancing algorithm is based on the Q-learning reinforcement learning framework. It continuously monitors server cluster traffic, including the number, type, and data transmission volume of user requests. Based on real-time traffic, it uses reinforcement learning to continuously try different resource allocation strategies, calculates the reward value under each strategy (such as system response time, throughput, and other indicators), and gradually learns the optimal resource allocation method. It dynamically adjusts server cluster resources and controls system response time fluctuations within ±5% during peak traffic periods (such as during design competitions), which is significantly better than traditional algorithms. The innovative system security protection mechanism uses federal zero-knowledge proof to audit user behavior without disclosing user privacy data. When verifying system compliance, users do not need to disclose private information. Verification can be completed in just 1.2 seconds through a specific encrypted interaction protocol.
[0050] The present invention also includes a design collaboration workshop module. For real-time multi-person annotation, the underlying communication architecture utilizes the WebSocket protocol. The server preconfigures a specific listening port. After the client initiates a connection request, a handshake is completed, establishing a full-duplex communication link. When designers annotate their designs on the client side, for example, when drawing a graphic, the client encodes the graphic's coordinates, color, shape, and other attributes, along with the operation timestamp. For text annotations, the client records the text content, font size, and other information. These information is then packaged into a message packet that follows a specific format (including an operation type identifier, operation subject information, and specific content data). The message packet is then transmitted to the server via a WebSocket channel. After verifying and unpacking the packet, the server forwards the message packet based on the target member's information. Messages are processed in an orderly manner using a message queue and distributed caching of commonly used data, ensuring transmission latency of ≤50ms and achieving real-time synchronization of operations. The version control mechanism is deeply customized based on the Git system. When the design team creates a project repository, the system automatically generates an initial version. Whenever a member submits a modification, the system extracts information such as the file list involved, the changed content (such as line numbers of code modifications, added or deleted characters, etc.), the submission time, and the member's account number, generating a new version snapshot and recording the inheritance relationship between versions. It supports branch management. Members can create independent branches from the main branch and carry out creative ideas, function implementation and other work in parallel on the branch. After completion, they will go through the merge request process and be merged into the main branch after team review. Simple conflicts of text files will be automatically resolved during the merge, and complex conflicts will prompt members to handle manually to ensure the orderly evolution of design versions.
[0051] The design process traceability system is built on blockchain technology, employing a consortium chain architecture. All parties involved in a design project (designers, copyright agencies, etc.) serve as nodes and jointly maintain the blockchain network. When a design work is created, the system automatically extracts metadata, including file hash values, format type, creation time, and author information, generating a unique digital identity. This identity is then packaged with the initial creation state data into a genesis block, which is then verified across all nodes and written to the blockchain using a consensus algorithm (such as Practical Byzantine Fault Tolerance (PBFT)). For each modification during the design process, the system captures the modification time, operator account, and file differences before and after the modification (generated through hash comparison). These data are encrypted and created into a new block, which contains the hash value of the previous block, forming a chain structure. When proof of copyright ownership is required, the system synchronizes the complete creation chain data from all blockchain nodes, rapidly retrieves key information, and uses smart contracts to automatically generate an encrypted proof document containing the work's creation history, copyright owner, and other details. The entire process is completed in less than one minute, providing authoritative and immutable evidence of copyright ownership.
[0052] The present invention also includes: a design quality assessment module, which constructs an evaluation index system including 8 dimensions such as innovation, functionality, and aesthetics. First, by crawling 100,000+ high-quality cases from design platforms such as Behance and Dribbble, and combining them with ISO design standard documents, an expert knowledge graph is constructed. The knowledge graph is stored in a graph database, and the nodes include entities such as design elements (such as colors, lines), evaluation indicators (such as contrast, user experience), and edges represent the relationship between entities. The deep reinforcement learning model uses the knowledge graph as prior knowledge. During training, the multimodal features of the design works (ResNet50 features of images, BERT vectors of text, MFCC features of audio) are input into the Actor-Critic network. The Actor network generates preliminary evaluation vectors and improvement suggestions. The Critic network evaluates the suggestions based on the expert experience in the knowledge graph and calculates the reward value (such as the degree of improvement in innovation). For example, for a certain UI design, the system will point out that "adjusting the main color from RGB (255, 255, 255) to RGB (240, 240, 240) can improve the contrast to the WCAGAA standard, based on the association rules between 'color contrast' and 'readability' in the knowledge graph."
[0053] The energy optimization unit uses Neural Architecture Search (NAS) technology to define a search space encompassing 20 operations, including convolutional layers, attention mechanisms, and depthwise separable convolutions. A reinforcement learning controller is used to generate candidate architectures, which are then trained and evaluated for 10 epochs on the CIFAR-10 dataset. To reduce computational overhead, a weight sharing strategy is employed, with different architectures sharing some weight parameters. The controller uses the number of model parameters and FLOPs as negative rewards and accuracy as positive rewards, optimizing the architecture using the REINFORCE algorithm. After 500 search iterations, the optimal architecture was found, maintaining 85% classification accuracy while reducing computational overhead by 60% compared to the traditional ResNet18. This architecture utilizes hybrid depthwise separable convolutions and Ghost modules, reducing inference energy consumption on the ImageNet validation set from 2.3mJ to 0.9mJ. When this low-power architecture is deployed on edge computing devices, it is further optimized using TensorRT, reducing classification latency for a single designed image from 320ms to 110ms, meeting real-time requirements.
[0054] The present invention also includes: an edge-side reasoning acceleration module, which introduces the formula In actual operation, a dedicated neural network processor that supports INT8 quantization is selected, and hardware parameters C are obtained in real time through system monitoring tools. peak with U util , the number of statistical model parameters N params And evaluate the computational complexity coefficient C complexINT8 quantization technology is used to convert model weights and activation values from high-precision to INT8, reducing computational complexity. Resources are dynamically allocated based on hardware utilization, increasing the computational load when utilization is low and optimizing the computational process when near saturation. This results in a frame rate of ≥30 frames per second (FPS), enabling fast and efficient local real-time inference. The multimodal interactive interface's voice commands utilize an end-to-end speech recognition model. The user's voice signal undergoes pre-processing, including noise reduction and normalization. The signal is then input into the model, where a multi-layer neural network extracts and analyzes voice features, converting it into text commands. The intent understanding model then interprets the command semantics and matches it to the design task type. For example, for a design generation task like "generate a minimalist logo," the relevant model is invoked to generate a design solution. Gesture interaction uses a depth camera to capture user hand movements in real time, extracting a sequence of key point coordinates. This is then fed into a spatiotemporal convolutional network to analyze the movement trajectory and classify it into primitives such as translation, scaling, and rotation. Multiple primitives are combined to form a complete gesture command, such as zooming in on a design by opening both hands. The system uses a Kalman filter to predict hand position, reducing latency and improving interaction fluidity. Haptic feedback integrates a piezoelectric ceramic array into the touch screen to generate tactile feedback signals with specific vibration frequencies and amplitudes based on the different material properties of the design work (such as metal and wood). When the user touches the virtual design work, the signal drives the piezoelectric ceramic to vibrate, simulating the touch of different materials. At the same time, the feedback intensity is adjusted in combination with the physical property parameters of the design elements. For example, raised elements produce stronger vibrations, enhancing user immersion and the intuitiveness of design evaluation.
[0055] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A multimodal design work intelligent classification and retrieval system based on deep learning, characterized by: Includes the following modules: Multimodal feature extraction module: This module uses a multi-scale convolutional neural network, integrating a visual Transformer and a residual network, to extract image texture, color, and geometric features. It combines BERT with a multi-head attention mechanism to perform text sentiment analysis and keyword extraction. It also uses Mel-frequency cepstral coefficients and spectral centroids to extract innovative audio features for emotion recognition. Cross-modal association analysis module: Builds a multimodal knowledge graph, uses graph neural networks to model the semantic associations between images, text, and audio, develops a dynamic attention fusion mechanism, and uses innovative formulas to: Adaptively adjust the weight distribution of different modal features, W t,i,a Represent the feature weights of text t, image i, and audio a respectively; m is the modal variable, S t,i,a is the semantic association score between the corresponding modality and the target task, λ t,i,a is the adjustment coefficient, and at the same time innovates the spatiotemporal attention mechanism to capture the correlation between temporal information and spatial features in design works; Multi-dimensional classification engine: Design a hierarchical cascade classifier architecture, using lightweight MobileNetV3 for coarse classification at the bottom layer and integrating Deep Forest and XGBoost algorithms for fine-grained classification at the top layer; Develop an incremental learning framework to continuously optimize classification models through a federated learning mechanism; Semantic Enhanced Retrieval Module: Builds a design intent understanding model, combines historical user interaction data with the current search context, and performs fuzzy matching and intent error correction on query semantics; Innovative multimodal hybrid retrieval strategy, supporting cross-modal retrieval methods such as text retrieval of images and image retrieval of text; Intelligent recommendation system: Develop a recommendation engine based on reinforcement learning, using the Actor-Critic architecture to dynamically adjust the recommendation strategy. Combined with a design trend prediction model, it enables personalized design recommendations. A design work value assessment module uses an attention mechanism to quantify the contribution of each feature to design quality and generate an interpretable rating report. System operation and maintenance monitoring module: Build a digital twin model of device performance to monitor server CPU utilization, memory usage, and network throughput in real time; develop an abnormal behavior detection algorithm and achieve early warning of system failures by combining isolation forests with autoencoders.
2. The multimodal design work intelligent classification and retrieval system based on deep learning according to claim 1 is characterized in that: The multimodal feature extraction module also includes: developing a multispectral image analysis unit that integrates RGB, near-infrared, and ultraviolet spectral imaging technologies to extract the material characteristics of the design works; and innovating a gesture action recognition module that uses a skeleton point detection algorithm and a spatiotemporal graph convolutional network to analyze the semantics of gestures during the design creation process.
3. The multimodal design work intelligent classification and retrieval system based on deep learning according to claim 1 is characterized in that: The cross-modal association analysis module also includes: constructing a design style transfer network to achieve automatic conversion of different design styles through a generative adversarial network; developing a multimodal knowledge distillation mechanism to encode expert knowledge into a graph structure representation to achieve cross-domain design knowledge transfer.
4. The multimodal design work intelligent classification and retrieval system based on deep learning according to claim 1 is characterized in that: The multi-dimensional classification engine also includes: designing a hierarchical attention classifier, realizing hierarchical representation of design elements through capsule networks, supporting step-by-step classification from macro style to micro details; and innovating a small sample learning module, using prototype networks combined with data enhancement technology to classify rare design categories.
5. The multimodal design work intelligent classification and retrieval system based on deep learning according to claim 1 is characterized in that: The semantic enhancement retrieval module also includes: introducing the formula P(I|U,F) refers to the probability of inferring design intent I when the user search behavior U and the design work feature F are known. P(U|I,F) is the conditional probability of U appearing given I and F. P(I,F) is the joint probability of I and F. i traverses all design intents, and the denominator serves as a normalizer. This formula analyzes causal relationships, assists the design intent inference engine in making recommendations, and cooperates with the knowledge question-answering system to serve users.
6. The multimodal design work intelligent classification and retrieval system based on deep learning according to claim 1 is characterized in that: The intelligent recommendation system also includes: a design trend prediction unit, which builds a time series prediction model to predict design trends six months in advance; a user feedback optimization mechanism, which uses a preference learning algorithm to update user portraits in real time.
7. The multimodal design work intelligent classification and retrieval system based on deep learning according to claim 1 is characterized in that: The system operation and maintenance monitoring module also includes: developing an adaptive load balancing algorithm to dynamically adjust server cluster resource allocation based on reinforcement learning; innovating system security protection mechanisms to implement user behavior audits through federated zero-knowledge proofs and verify system compliance without leaking privacy.
8. The multimodal design work intelligent classification and retrieval system based on deep learning according to claim 1 is characterized in that: It also includes: a design collaboration workshop module that integrates real-time multi-person annotation tools and version control mechanisms to support collaborative creation by cross-regional design teams; and the development of a design process traceability system that records the creation history and modification trajectory of design works through blockchain.
9. The multimodal design work intelligent classification and retrieval system based on deep learning according to claim 1 is characterized in that: Also includes: The design quality assessment module builds a multi-dimensional evaluation index system, combines expert knowledge graphs with deep reinforcement learning, and generates explainable design improvement suggestions; Energy optimization unit, which uses neural architecture search to automatically design low-power model architecture.
10. The multimodal design work intelligent classification and retrieval system based on deep learning according to claim 1 is characterized in that: Also includes: The edge inference acceleration module introduces an innovative formula: Among them, FPS is the number of frames processed per second, C peak It is the peak computing power of the dedicated neural network processor, U util is the hardware utilization, Q bit is the number of quantization bits, N params is the model parameter, C complex To calculate the complexity coefficient; at the same time, the multimodal interaction interface supports voice commands, gesture control, and eye tracking.
Citation Information
Cited By
Method for classifying species of migrant birds based on voiceprint recognition
CN121237099A
Multi-modal sensing-based spaceflight exercise equipment man-machine interaction optimization method and system
CN121578893A