A material extraction method and system based on a large model and a multi-storage technology

CN121958577BActive Publication Date: 2026-09-29GOLDEN TIMES CULTURE COMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610016542.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-09-29
Estimated Expiration
2046-01-07

AI Technical Summary

Technical Problem

面对如此庞大且异构的数据流,传统的素材处理方式已难以满足高效性、准确性与可扩展性的需求

Benefits of technology

[0008]第三方面,本申请提供一种计算机可读存储介质,采用如下的技术方案:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958577B_ABST
    Figure CN121958577B_ABST
Patent Text Reader

Abstract

The application relates to a material extraction method and system based on a large model and a multi-storage technology, and belongs to the field of computer information technology. The material extraction method comprises the following steps: receiving heterogeneous data streams from text, image, video and audio sources, and performing protocol analysis to generate standardized data blocks; extracting data volume features, access frequency features and semantic entropy value features, and combining the features into a routing feature vector; calculating a routing decision score according to a weight matrix, distributing the data blocks to a first level in a memory, a graph structure or a compressed storage layer according to the score, and outputting a location identifier; in response to a query, retrieving associated data from the hierarchical storage, aligning multi-modal features through an attention mechanism and a space-time convolution, generating a fusion feature vector and inputting the fusion feature vector into a pre-training base model fine-tuned by a lightweight adapter module, and outputting a prediction result; and converting the prediction result into structured data in a target format according to a request end device type and outputting the structured data. The application improves user experience and system compatibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer information technology, and in particular to a method and system for extracting materials based on large model and multi-storage technology. Background Technology

[0002] With the rapid development of internet and mobile communication technologies, multimedia data is experiencing explosive growth, including various data sources such as text, images, videos, and audio, continuously generated from web pages, sensors, social media platforms, and other digital terminal devices. Faced with such a massive and heterogeneous data stream, traditional material processing methods are struggling to meet the demands for efficiency, accuracy, and scalability.

[0003] In current technological practices, when faced with heterogeneous data streams containing multiple modalities such as text, images, video, and audio, these data streams often have different transmission protocols, encoding formats, and access patterns. Traditional storage architectures, lacking a deep understanding of the characteristics of multimodal data, cannot intelligently allocate storage based on the actual access patterns and semantic features of the data. This results in low storage resource utilization efficiency, with high-frequency and low-frequency access data being stored together, impacting the overall system's responsiveness. Furthermore, when processing spatiotemporal information and textual semantics simultaneously, existing systems often fail to fully capture the complex relationships between multimodal data. They also cannot adaptively adjust data formats based on the device characteristics of the requesting end when facing diverse output requirements from different terminal devices, affecting user experience and system compatibility. Summary of the Invention

[0004] To address the aforementioned technical issues, this application provides a method and system for extracting materials based on large model and multi-storage technology.

[0005] Firstly, this application provides a material extraction method based on large model and multi-storage technology, employing the following technical solution: A method for extracting materials based on large models and multi-storage technology, the method comprising: Receive heterogeneous data streams from text, image, video, and audio sources, perform protocol parsing on the heterogeneous data streams, and generate standardized data blocks; Extract the data volume features, access frequency features, and semantic entropy value features of the standardized data blocks, and combine them into a routing feature vector; The routing feature vector is weighted according to the dynamically updated weight matrix to generate a routing decision score; Based on the routing decision score, the standardized data block is allocated to one of the following levels: memory storage layer, graph structure storage layer, or compressed storage layer, and the corresponding storage location identifier is output. In response to a query request from an external client, the system retrieves associated data from the hierarchical storage layer based on the storage location identifier, aligns multimodal features through an attention mechanism and spatiotemporal convolution operations, and generates a cross-modal fusion feature vector. The cross-modal fusion feature vector is input into the pre-trained base model, and the prediction result is output through the lightweight adapter module; Based on the device type of the external requesting party, the prediction results are converted into structured data in the target format and output.

[0006] By adopting the above technical solutions, the system receives and standardizes multi-source heterogeneous data streams, ensuring data uniqueness and traceability; it uses a dynamic routing mechanism to intelligently allocate data to hierarchical storage layers, optimizing storage efficiency and access performance; it combines attention mechanisms and spatiotemporal convolution to achieve efficient alignment and fusion of multimodal features, improving feature representation capabilities; it leverages pre-trained base models and lightweight adapter modules to output accurate predictions with low computational cost while ensuring generalization ability; and finally, it adaptively converts the result format according to the requesting device, significantly enhancing the system's compatibility, response speed, and overall usability.

[0007] Secondly, this application provides a material extraction system based on large model and multi-storage technology, which adopts the following technical solution: A material extraction system based on large model and multi-storage technology, specifically including: The heterogeneous data stream standardization module is used to receive heterogeneous data streams from text, image, video and audio sources, perform protocol parsing on the heterogeneous data streams, and generate standardized data blocks. The routing feature extraction module is used to extract the data volume features, access frequency features, and semantic entropy value features of the standardized data block, and combine them into a routing feature vector. The intelligent routing decision module is used to perform weighted calculations on the routing feature vectors based on a dynamically updated weight matrix to generate a routing decision score. The hierarchical storage allocation module is used to allocate the standardized data block to a first-level storage layer among the memory storage layer, graph structure storage layer, or compressed storage layer based on the routing decision score, and output the corresponding storage location identifier. The cross-modal retrieval and fusion module is used to respond to query requests from external requesters, retrieve related data from the hierarchical storage layer according to the storage location identifier, align multimodal features through attention mechanism and spatiotemporal convolution operation, and generate cross-modal fusion feature vectors. The multimodal prediction output module is used to input the cross-modal fused feature vector into the pre-trained base model and output the prediction result through the lightweight adapter module. The device adaptation and conversion module is used to convert the prediction result into structured data in the target format and output it according to the device type of the external requesting end.

[0008] Thirdly, this application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect.

[0009] In summary, this application includes at least one of the following beneficial technical effects: by constructing a multi-layered heterogeneous data processing architecture, it achieves unified acquisition, intelligent routing, and efficient storage of multimodal data such as text, images, videos, and audio; by utilizing a large model-driven cross-modal feature fusion and lightweight adapter mechanism, it enhances the accuracy and processing efficiency of material extraction; and by using dynamic weight optimization and hierarchical storage strategies, it effectively reduces system latency and resource consumption, ultimately achieving adaptive output adaptation for different terminal devices, improving user experience and system compatibility. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of the first process of a material extraction method based on large model and multi-storage technology according to one embodiment of this application.

[0011] Figure 2 This is a schematic diagram of the second process of a material extraction method based on large model and multi-storage technology according to one embodiment of this application.

[0012] Figure 3 This is a schematic diagram of the third process of a material extraction method based on large model and multi-storage technology according to one embodiment of this application.

[0013] Figure 4 This is a schematic diagram of the fourth process of a material extraction method based on large model and multi-storage technology according to one embodiment of this application.

[0014] Figure 5 This is a schematic diagram of the fifth process of a material extraction method based on large model and multi-storage technology according to one embodiment of this application.

[0015] Figure 6 This is a schematic diagram of the sixth process of a material extraction method based on large model and multi-storage technology according to one embodiment of this application.

[0016] Figure 7 This is a schematic diagram of the seventh process of a material extraction method based on large model and multi-storage technology according to one embodiment of this application. Detailed Implementation

[0017] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figures 1-7 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.

[0018] This application discloses a material extraction method based on large model and multi-storage technology.

[0019] Reference Figure 1 A method for extracting materials based on large models and multi-storage technology, specifically including: Step S101: Receive heterogeneous data streams from text, image, video and audio sources, perform protocol parsing on the heterogeneous data streams, and generate standardized data blocks; The standardized data blocks carry hash values, timestamps, and modal tags; Specifically, the system needs to handle data streams from different sources, in different formats, and with different transmission protocols. This raw data may come from text content obtained by web crawlers, real-time video frames captured by cameras, audio clips captured by microphones, or unstructured documents returned by third-party APIs. Since they each follow different communication standards (such as HTTP for web requests, MQTT for lightweight messaging in IoT scenarios, and RTMP commonly used for live streaming), it is necessary to use probe modules deployed at the access layer to identify and decode various protocols, thereby completing the information restoration process from the physical layer to the application layer.

[0020] Building upon this foundation, to ensure strong uniqueness and traceability in subsequent operations, each successfully parsed basic unit is appended with a SHA-256 hash value as an identifier, along with a timestamp of the capture time and a modal tag for its media type (e.g., text / image / audio / video). This approach not only helps prevent duplicate writes but also supports rapid retrieval of specific resources by time period or category.

[0021] Step S102: Extract the data volume features, access frequency features, and semantic entropy value features of the standardized data blocks, and combine them into a routing feature vector; Among these features, "data volume characteristic" refers to a quantitative indicator of the space occupied by the data block. It is usually normalized using a natural logarithmic transformation to avoid imbalances in subsequent calculations due to extreme values. "Access frequency characteristic" refers to the number of read requests initiated against the data block per unit of time. It can be continuously tracked and counted by maintaining a sliding time window to reflect whether the data is currently in a hotspot state. The introduction of "semantic entropy value characteristic" is particularly crucial. It is a measure of the complexity of information within a text segment. More random and disordered language expressions tend to have higher entropy values, and vice versa.

[0022] In this embodiment, a lightweight BERT language model is used for context modeling, and the information entropy level of the entire sentence or passage is estimated accordingly. The resulting routing feature vector is a high-dimensional vector representation that includes the comprehensive evaluation results of the above three dimensions, which can comprehensively characterize the behavior patterns and intrinsic characteristics of a certain object to be assigned.

[0023] Step S103: The routing feature vector is weighted according to the dynamically updated weight matrix to generate a routing decision score; "Dynamic updates" means that the system's strategy is not fixed, but will automatically adjust its behavior as the environment changes.

[0024] In this embodiment, a learning mechanism based on a reinforcement learning framework is designed, where the reward function R is defined as 1 minus the ratio between the actual response latency and a preset service level threshold. That is, the closer the latency is to or even lower than the expected target, the higher the corresponding reward; conversely, the penalty increases. Through iterative training, the specific values ​​of each element in the weight matrix are continuously optimized, ensuring that each weighted sum more accurately reflects which level of data should be preferentially cached. This approach effectively addresses the impact of fluctuating business loads, improving overall scheduling efficiency while reducing the probability of resource waste.

[0025] Step S104: Based on the routing decision score, allocate the standardized data blocks to one of the following levels: memory storage layer, graph structure storage layer, or compressed storage layer, and output the corresponding storage location identifier. Specifically, based on the different needs of various application scenarios, the entire storage system can be divided into three layers: a high-speed but limited-capacity memory storage layer is mainly used to store frequently accessed and time-sensitive objects; the middle graph structure storage layer uses a graph database organization method to manage datasets with complex relational links, which is particularly suitable for describing the interactions between entities; and the last layer is specifically used to store historical data that has not been accessed for a long time, and its disk space is significantly reduced by implementing an efficient fractal compression algorithm. Whenever a new standardized data block is generated, the system will place it appropriately among these three layers according to the previously obtained routing score and assign it a globally unique location number for future retrieval.

[0026] Step S105: Respond to the query request from the external requester, retrieve related data from the hierarchical storage layer according to the storage location identifier, align multimodal features through attention mechanism and spatiotemporal convolution operation, and generate cross-modal fusion feature vector; The system will respond to external query requests, requiring not only accurate location of the target file but also extraction of useful signals for further analysis. For text-based materials, the self-attention mechanism in the Transformer architecture can be used to capture long-distance dependencies between words, thereby obtaining a set of embedding vectors rich in contextual meaning. For multimedia materials such as images or videos, a three-dimensional convolutional neural network combining receptive fields in both temporal and spatial directions is employed to extract high-level abstract features at the frame and even sequence levels. These two methods are applicable to different types of information formats and together form the foundation for subsequent cross-modal integration.

[0027] Step S106: Input the cross-modal fusion feature vector into the pre-trained base model and output the prediction result through the lightweight adapter module; The "base model," as a pre-trained general backbone network, is typically initialized with parameters on a large-scale multimodal corpus through alignment tasks or other self-supervised objectives. It is responsible for extracting deep, abstract semantics from primary fusion features. Since the input is already a highly integrated cross-modal representation, the base model does not need to perform low-level feature extraction and can directly proceed to the nonlinear mapping stage.

[0028] Furthermore, this application superimposes a small parameter adjustment component, the so-called "lightweight adapter," onto the original backbone network. Its basic idea is to allow only this newly added connection to participate in gradient backpropagation and parameter updates while keeping the backbone network weights frozen. This ensures the continuation of the strong generalization ability while also enabling small-scale adjustments tailored to specific task objectives. Particularly in the embodiments of this application, since high-quality feature extraction has already been completed, only a few additional parameters are needed to achieve the desired fitting effect, greatly improving deployment flexibility and operational economy.

[0029] Step S107: Based on the device type of the external requesting end, convert the prediction results into structured data in the target format and output them.

[0030] The structured data conversion of the target format includes: generating a JSON key-value pair structure when the requesting device is identified as a PC; generating a Protocol Buffers binary sequence when identified as a mobile terminal; and converting it into a three-dimensional point cloud coordinate array when identified as an AR device.

[0031] Understandably, the same data analysis conclusions might be better presented as an easy-to-read and editable JSON string for desktop computer users, while mobile app developers might prefer a compact and efficient binary stream, and some cutting-edge AR glasses manufacturers might even expect to receive a 3D point cloud data package that can be directly rendered and displayed.

[0032] Therefore, this application embodiment establishes a complete client fingerprint recognition mechanism. By jointly judging the User-Agent field and other environmental variables, it accurately distinguishes various terminal platforms and selects the most appropriate form of expression to release the final result.

[0033] In the above implementation, multi-source heterogeneous data streams are received and standardized to ensure data uniqueness and traceability; a dynamic routing mechanism is used to intelligently allocate data to a hierarchical storage layer to optimize storage efficiency and access performance; attention mechanisms and spatiotemporal convolutions are combined to achieve efficient alignment and fusion of multimodal features, improving feature representation capabilities; a pre-trained base model and a lightweight adapter module are used to output accurate predictions with low computational cost while ensuring generalization ability; and finally, the result format is adaptively converted according to the requesting device, significantly enhancing the system's compatibility, response speed, and overall usability.

[0034] The technical solution of this application can be widely applied to intelligent media content management platforms. For example, in large video websites, the system can simultaneously process heterogeneous data such as user-uploaded videos, audio, bullet screen text, and screenshots. It ensures the uniqueness of content through hash values ​​and modal tags, and uses dynamic routing to store popular video clips in the memory layer, user relationship graphs in the graph database, and historical content in compressed storage. When a user searches for relevant content, the system uses multimodal fusion technology to understand the semantic relationship between video content and text descriptions. Finally, based on the different terminal types used by the user, such as PCs, mobile phones, or VR devices, it intelligently outputs corresponding structured recommendation results, achieving millisecond-level response and personalized content distribution.

[0035] Reference Figure 2 As one implementation of step S103, the step of generating a routing decision score by weighting the routing feature vector according to the dynamically updated weight matrix includes: Step S201: Obtain the routing feature vector, which includes data volume features, access frequency features, and semantic entropy value features. These three key dimensions represent information abstractions at the physical, behavioral, and semantic levels, respectively, and together constitute an input vector in a high-dimensional feature space for subsequent intelligent decision-making.

[0036] Step S202: Read the currently stored weight matrix; The weight matrix is ​​not static; rather, it is obtained by periodically updating it through a reinforcement learning mechanism.

[0037] Step S203: Receive response latency data and service level agreement thresholds collected during system operation in real time, and calculate the reward function value; Specifically, each iteration collects relevant performance data such as response latency and service level agreement (SLA) thresholds generated in the actual operating environment, and calculates the corresponding reward function value based on this data. The purpose of this formula is to enable the system to perceive changes in its service quality. The closer the response speed is to or better than the predetermined target, the higher the corresponding reward, and vice versa.

[0038] Step S204: Based on the reward function value, incrementally modify each element in the weight matrix using the Q-learning algorithm to periodically update the weight matrix; Specifically, the Q-learning algorithm is used to incrementally modify each element in the weight matrix, i.e. ; In the above formula, W_ij(t) represents the value of the element in the i-th row and j-th column of the weight matrix at the current time t; R is the reward function value; the learning rate α controls the aggressiveness of the update magnitude at each step; the discount factor γ affects the importance weight of future returns relative to immediate rewards; and F_j represents the j-th dimension component of the current input routing feature vector; Q(s,a) represents the expected long-term cumulative reward of performing an action a (i.e., making a routing decision) in the current state s (defined by the current system state); max Q(s',a') represents the maximum long-term cumulative reward that can be obtained from all possible actions a' in the next state s' after performing action a. This design ensures that the weight matrix can gradually approach the optimal solution through continuous trial and error, thereby enhancing the robustness and generalization ability of the entire routing strategy.

[0039] In this embodiment, the state space of the Q-learning algorithm is defined as: s= CPU load, memory usage, network throughput The action space is defined as the direction of incremental adjustment of the weight matrix: a∈{increase text weight, increase image weight, increase video weight}.

[0040] Step S205: Perform matrix multiplication on the routing feature vector and the weight matrix to generate a routing decision score.

[0041] This step is essentially a linear combination process, where multiple heterogeneous features are weighted and summed according to a certain priority ratio to ultimately generate a single scalar routing decision score. More importantly, as the number of training rounds increases, the coefficients at each position within the weight matrix tend to reach a reasonable distribution. This means that even the same input sample may achieve drastically different scores at different times, fully demonstrating the algorithm's ability to self-evolve.

[0042] In the above implementation, adaptive routing decision optimization based on Q-learning reinforcement learning is achieved through weighted calculation of dynamic weight matrix and routing feature vector. The system can perceive network performance changes in real time and automatically adjust weight parameters, so that routing decisions simultaneously consider data volume at the physical layer, access frequency at the behavioral layer, and entropy characteristics at the semantic layer. This allows for continuous optimization of routing strategies in complex network environments, improving system response efficiency and service quality, and ensuring the intelligence and adaptability of routing decisions.

[0043] Reference Figure 3 As one implementation of step S104, the step of allocating standardized data blocks to a level among the memory storage layer, graph structure storage layer, or compressed storage layer based on routing decision scores, and outputting the corresponding storage location identifiers, includes: Step S301: Obtain the routing decision score and the corresponding standardized data block; The routing decision score is a quantitative indicator reflecting the probability of a data block being frequently accessed in the current or future period. This score is typically derived from historical call records in upper-layer business scenarios, the output of context-aware models, or inferences about user behavior patterns by machine learning engines. For example, in a video platform, a newly released TV series may receive a high access score due to its popularity; while in the field of scientific data analysis, intermediate results from recent experiments may have higher priority because they are in an active processing phase.

[0044] Meanwhile, the standardized data block contains three key metadata fields: a hash identifier for uniquely identifying content entities (preventing duplicate writing), a timestamp indicating the time of its creation or most recent modification (facilitating timeliness assessment), and modal tags (such as images, audio, and text) to support cross-media type retrieval needs. These three dimensions together constitute a complete information unit, providing necessary semantic support for subsequent hierarchical decisions.

[0045] Step S302: Compare the routing decision score with preset hot storage threshold, warm storage threshold and cold storage threshold; This step essentially involves implementing a data lifecycle segmentation strategy based on a performance-cost trade-off. Here, "hot storage," "warm storage," and "cold storage" are not simply temperature concepts, but rather refer to different levels of storage devices and their corresponding service response capabilities. "Hot storage" is typically deployed in high-speed caches or distributed memory clusters, offering extremely low latency but limited capacity; "warm storage" generally exists in the form of high-performance disk arrays or graph databases, balancing speed and scalability; while "cold storage" refers to object storage services used for long-term archiving, offering near-infinite persistent storage capabilities despite higher read latency.

[0046] Therefore, setting multiple critical thresholds aims to establish a dynamic switching mechanism, enabling each piece of data to automatically adapt to the optimal storage location based on its potential value and usage frequency. This approach avoids the resource waste or access bottlenecks caused by traditional one-size-fits-all storage arrangements and improves the overall system's elasticity and scheduling capabilities.

[0047] Step S303: Determine whether the routing decision score is greater than the hot storage threshold; if yes, proceed to step S304; if no, proceed to step S305. Step S304: Store the standardized data block in the memory storage layer and generate a hot storage layer identifier containing the memory node address; Specifically, when the routing decision score is higher than the hot storage threshold, a write operation to the distributed memory cluster is triggered, and a corresponding hot storage layer identifier is generated.

[0048] In this process, "distributed memory clusters," as a type of high-concurrency, low-latency data storage facility, are often built upon architectures such as Redis Cluster, Apache Ignite, or other in-memory data grids. They can not only handle fast read / write tasks for large amounts of frequently accessed data but also allow cross-node synchronous replication to ensure reliability. The "hotspot identifier" generated after the write operation is a structured string containing the memory node's IP address and specific offset (e.g., MEM_192.168.1.10_0x1A2B3C4D). This design ensures that clients can directly skip the directory search step and locate the physical memory region where the target data is located in one step, thus significantly improving the response efficiency of frequently accessed categories. Furthermore, due to the scarcity of memory resources, the improved LRU-K algorithm mentioned later is needed to reasonably reclaim space and maintain the continuous availability of the system.

[0049] Step S305: Determine whether the routing decision score is between the warm storage threshold and the cold storage threshold; if yes, proceed to step S306; if no, proceed to step S307. Step S306: Store the standardized data blocks in the graph structure database, construct feature vector nodes and connect them with cross-modal hyperedges, and generate warm storage layer identifiers; Among them, "graph structure databases" are different from traditional relational databases. They are a special data management system for complex network topology modeling. Products such as Neo4j and ArangoDB belong to this category.

[0050] In this embodiment, each newly added standard data block undergoes a deep semantic parsing process to extract core vector features (i.e., so-called embedding representations) that can be used to express its inherent meaning. These features can be BERT encoding of natural language sentences, ResNet abstract projection of image content, or MFCC statistical summaries of audio signals, etc.

[0051] The system then calculates the cosine similarity between the new node and its surrounding existing nodes. This is a classic mathematical tool for measuring the angle between two non-zero vectors, widely used in recommendation systems, search engines, and other fields. Once an existing node is found to have a similarity exceeding a preset threshold (e.g., 0.7), a weighted "hyperedge" connection is established, forming a knowledge graph subnetwork with generalized reasoning capabilities. This mechanism not only effectively organizes the implicit connections between massive amounts of heterogeneous materials but also assists in achieving more accurate content retrieval through graph traversal algorithms. The resulting "graph path identifier" carries complete graph ID, node ID, and even hyperedge ID information, providing a highly condensed relation index entry point for downstream search engines.

[0052] Step S307: Determine whether the routing decision score is less than the cold storage threshold; if so, proceed to step S308. Step S308: Perform lossless compression encoding on the standardized data block, store the compressed data in the compressed storage layer, and generate a cold storage layer identifier.

[0053] This step reflects the system's high regard for storage economy. "Lossless compression coding" is different from JPEG and MP4, which sacrifice some quality for size compression. Instead, it reduces the number of redundant bits as much as possible while fully preserving the original information.

[0054] In this embodiment, a "fractal compression algorithm" derived from fractal geometry theory is introduced. This algorithm treats the image to be compressed as the result of iterative transformations of several local regions and attempts to approximate the details of the original image using a set of affine transformation functions. Although this method incurs significant upfront overhead, it maintains good reconstruction accuracy even at high compression ratios.

[0055] More importantly, since the material library may contain a large amount of texture-dense visual data (such as satellite remote sensing images, medical scan slices, etc.), using affine coefficients to describe its periodicity and self-similarity properties perfectly matches the needs of practical application scenarios. The compressed output is then sent to a remote Object Storage Service (OSS), one of the most mainstream large-scale unstructured data hosting platforms in cloud computing environments, possessing inherent cost advantages and off-site disaster recovery mechanisms. To facilitate the future retrieval of these dormant assets, the system will also issue a unique "compression locator," whose composition rules explicitly state that it should include both the bucket name and the encrypted object key, for secure isolation and precise addressing.

[0056] It should be noted that although the various identifiers have different forms, they all follow a unified design principle in terms of naming conventions: the prefixes are respectively given as MEM, GRAPH, and OSS to distinguish their respective levels; the middle part accurately reflects the key parameter combination of their internal storage structure; and additional information such as version number and checksum can be added at the end to enhance robustness. This ensures the consistency of the system's external interfaces and enhances the visualization convenience for maintenance personnel when troubleshooting.

[0057] In the above implementation, automated hierarchical management of data is achieved. The system dynamically allocates data to the memory storage layer, graph structure storage layer, or compressed storage layer based on the frequency of data access, and generates corresponding storage location identifiers. This ensures low-latency response for frequently accessed data, maintains semantic relationships between data through the graph structure database, and optimizes the storage cost of cold data using lossless compression technology. Overall, this achieves optimal allocation of storage resources and a significant improvement in access performance.

[0058] Reference Figure 4 As one implementation of step S105, the steps of responding to a query request from an external requester, retrieving associated data from the hierarchical storage layer based on the storage location identifier, aligning multimodal features through an attention mechanism and spatiotemporal convolution operations, and generating a cross-modal fusion feature vector include: Step S401: Receive a query request from an external requesting party, and parse the semantic keywords and device type of the requesting party contained in the query request; The core objective of this step is to establish an intelligent response mechanism for multi-terminal environments. External query requests can be content retrieval instructions issued from various access methods such as mobile clients, desktop browsers, or edge servers; and the process of parsing these requests involves key aspects of natural language processing, namely semantic keyword extraction and device context recognition.

[0059] Specifically, semantic keywords guide the precise positioning and filtering of information on specific topics. For example, when a user inputs "cat playing," the system needs to accurately convert it into a topic representation vector with generalization capabilities, such as "animal play" or "catbehavior." Simultaneously, it needs to determine hardware constraints such as whether the requesting terminal device supports high-definition video playback and has sufficient memory resources to handle complex model inference results, providing a basis for selecting appropriate feature granularity. This stage is not only the starting point of the entire system's intelligent decision-making chain but also determines the direction and effectiveness of all subsequent operations.

[0060] Step S402: Determine the storage level of the target data based on the storage location identifier, and obtain the corresponding original data block; Specifically, the system adopts a three-tier storage structure: a memory storage layer, a graph structure storage layer, and a compressed storage layer. Each type of data is assigned a unique "storage location identifier" upon initial storage. This identifier not only contains basic location index information but may also include related meta-attribute fields such as access priority, update cycle, and modality category.

[0061] This design enables the system to quickly locate the storage area of ​​the target object when faced with a massive heterogeneous material library, thus avoiding the performance loss caused by a full scan. For example, frequently accessed hot text fragments are usually placed in the memory layer to ensure extremely low latency reading, while historical video data that has not changed for a long time but still has potential value can be archived and saved in the compression layer using a high compression ratio. This layered strategy greatly improves the system's scalability and operating efficiency.

[0062] Step S403: Perform a multi-head attention mechanism on the text modal data in the original data block to extract text feature vectors related to semantic keywords; Specifically, the execution of the multi-head attention mechanism includes: converting the text data into a word embedding matrix after word segmentation; and calculating the dot product similarity between the query vector and the key vector. ; In the above formula, Q is the query vector generated from semantic keywords, which originates from the semantic keywords parsed from the user's query request; K and V are text word embedding matrices, and d K is the dimension of the key vector.

[0063] In this embodiment, the purpose of the formula is to accurately extract feature vectors that are highly correlated with the "semantic keywords" queried by the user from the text modality data. These feature vectors, along with the feature vectors from the visual modality, are then sent to a downstream fusion module for further processing to complete the cross-modal intelligent retrieval task.

[0064] Step S404: Perform a spatiotemporal separation convolution operation on the visual modality data in the original data block to extract spatiotemporally correlated visual feature vectors; The spatiotemporal separation convolution operation includes: first, performing spatial dimension convolution, using a 3×3 convolution kernel to extract keyframe spatial features; then performing temporal dimension convolution, executing a one-dimensional convolution kernel along the time axis to capture motion trajectories; and finally fusing spatial and temporal features through residual connections.

[0065] Specifically, this scheme employs a spatial-temporal decoupled dual-path convolution mode. First, a small two-dimensional filter is applied to each frame individually to extract static visual element feature maps. Then, a one-dimensional kernel function slides along the frame sequence to capture continuous motion pattern transition curves. The intermediate products obtained from these two methods are fused through residual connections and then fed into the nonlinear activation function ReLU for further reinforcement learning. This method retains the advantages of the original structure while significantly reducing computational overhead, and also enhances the model's robustness and generalization ability. For example, in analyzing a video of an athlete running training, this method can simultaneously capture details of foot posture changes and overall rhythm fluctuations, laying a solid foundation for higher-order behavioral understanding tasks.

[0066] Step S405: Tensor fusion is performed on the text feature vector and the visual feature vector to generate a cross-modal fused feature vector.

[0067] Understandably, modern artificial intelligence research increasingly emphasizes the importance of breaking the limitations of single sensory channels, attempting to enable machines to learn to understand the world by comprehensively utilizing multiple sensory channels, just like humans. In this process, how to appropriately combine seemingly unrelated text encoding with image descriptions has become a pressing technical challenge.

[0068] To address this, this application employs a highly versatile and compatible tensor product operation for connection. This involves mapping two sets of features to the same shared semantic space and then performing inner product or other bilinear operations to form a new composite vector representation. This approach not only preserves the advantages of each individual task but also reveals the potential connection between them at a higher level, laying a solid foundation for further improving the performance of downstream tasks.

[0069] In the above implementation, a hierarchical storage architecture and a multimodal fusion mechanism are constructed to achieve efficient and accurate cross-modal retrieval response. The system first quickly locates target data based on storage location identifiers, avoiding the performance loss of full-data scanning; then, it uses a multi-head attention mechanism to extract semantically relevant text features and captures the spatiotemporal correlation information of visual data through spatiotemporal separation convolution operations; finally, it performs tensor fusion of feature vectors from different modalities to generate a unified cross-modal fused feature vector. This technical solution not only significantly improves the retrieval efficiency and accuracy of massive heterogeneous data but also achieves intelligent resource adaptation through device type recognition, providing reliable technical support for intelligent response in multi-terminal environments.

[0070] As one implementation of step S402, the step of determining the storage level where the target data is located based on the storage location identifier and obtaining the corresponding original data block includes: If the target data is located in the memory storage layer, the original data block is read directly through the cache interface; The cache interface is typically a high-efficiency access channel built on LRU (Least Recently Used) or other advanced replacement algorithms. Because memory itself has nanosecond-level random access speeds, once data is confirmed to be active, the complex addressing and loading processes can be bypassed to retrieve the required content directly. This step not only significantly reduces average response time but also alleviates the pressure on the underlying database.

[0071] In addition, to further enhance concurrent service capabilities, actual deployments often involve working in conjunction with distributed cache clusters, which replicate identical data copies across multiple nodes to form a redundant backup system, thereby enhancing fault tolerance and load balancing.

[0072] If the target data is located in the graph structure storage layer, execute the hypergraph traversal algorithm to retrieve the original data blocks connected by associated nodes and hyperedges; In this embodiment, the hypergraph traversal algorithm specifically includes: using the semantic embedding vector of the query keyword as the initial node; iteratively searching for associated nodes with a similarity exceeding the threshold of 0.8 along the hyperedge connection direction; and returning a set of associated data blocks in descending order according to the node weight values.

[0073] In this context, a "hypergraph" is a generalized version of a graph theory concept. Compared to a traditional simple graph, it allows multiple vertices to share an edge (called a hyperedge), making it ideal for modeling many-to-many relationship networks. Within this framework, each independent knowledge unit is abstracted as a node, and the co-occurrence, causal, or semantic relationships between them are characterized by their corresponding hyperedges. Hypergraph traversal algorithms utilize these topological characteristics for navigation and search. Their core idea is to use the semantic embedding vector mapped from the initial query keywords as the starting point, gradually expanding the exploration range along connected hyperedges until a predetermined termination condition is reached.

[0074] Throughout the process, a similarity threshold (e.g., 0.8) is set. Only newly discovered nodes whose cosine distance to the current central node is less than this value will be included in the candidate list for final sorting. In addition, each node is assigned a corresponding weight parameter to reflect its importance. Finally, the most likely combination of data blocks to meet the requirements is returned after being sorted in descending order for upper-layer calls and consumption.

[0075] If the target data is located in the compressed storage layer, the decompression engine is called to restore the data blocks to their original format, thus obtaining the original data blocks.

[0076] In this embodiment, the decompression engine's workflow includes: identifying the fractal encoding identifier in the compression header; reconstructing image / video frames using an iterative function system; and verifying the consistency between the SHA256 hash value of the reconstructed data and the original metadata.

[0077] Among them, the fractal coding technology mentioned above is a unique image compression method developed based on the mathematical theory of IFS (Iterated Function System). It does not rely on the traditional discrete cosine transform, but achieves the ultimate compression ratio through repeated iterations of self-similarity.

[0078] Specifically, considering that some old historical data may be large in size but have low utilization, it is necessary to perform lossless compression before storing it in the database. To this end, the system has a built-in customized decompression engine component responsible for restoring the original form. This engine mainly includes two sub-processes: first, a header parsing module, which analyzes the file header to identify the specific encoding standard (such as JPEG, H.264, Fractal Encoding, etc.); second, a reconstruction engine, which regenerates pixel sequences or character streams according to their respective compression specifications.

[0079] In addition, for security reasons, an integrity verification process is added after each decompression, which compares the fingerprint of the data digest to be recovered (SHA-256 hash value) with the original registration record to prevent the risk of tampering.

[0080] In the above implementation, differentiated data acquisition strategies are adopted for different storage tiers, achieving an optimal balance between storage efficiency and access performance. High-speed cache direct read in the memory storage layer ensures nanosecond-level response for hot data; the hypergraph traversal algorithm in the graph structure storage layer can deeply mine the complex relationships between multimodal data; and the intelligent decompression engine in the compressed storage layer achieves efficient storage and on-demand recovery of massive amounts of historical data while ensuring data integrity. This layered and heterogeneous data management architecture not only significantly improves the overall retrieval efficiency and concurrent processing capabilities of the system but also enhances the system's reliability and security through distributed caching and integrity verification mechanisms.

[0081] Reference Figure 5 As one implementation of step S106, the step of inputting the cross-modal fusion feature vector into the pre-trained base model and outputting the prediction result through the lightweight adapter module includes: Step S501: Input the cross-modal fusion feature vector into the pre-trained base model and perform high-order feature transformation through a fully connected layer; The base model, as a pre-trained general backbone network, is typically initialized with parameters on a large-scale multimodal corpus through alignment tasks or other self-supervised objectives. It is responsible for extracting deep, abstract semantics from primary fused features. Since the input is already a highly integrated cross-modal representation, the base model does not need to perform low-level feature extraction and can directly proceed to the nonlinear mapping stage.

[0082] In this process, a fully connected layer means that every node in the current layer is connected to all nodes in the previous layer. Mathematically, it is a combination of linear transformations and activation functions. Where W is the weight rectangle, b is the bias term, and f() represents a nonlinear activation function (such as ReLU, GELU, etc.). Multiple consecutive fully connected layers stacked together can form complex nonlinear decision boundaries, which helps to discover complex patterns hidden behind high-dimensional features. Especially for multimodal problems, fully connected structures are beneficial for breaking down modal barriers, promoting cross-domain knowledge transfer, and further enhancing the model's generalization ability.

[0083] Step S502: Connect a lightweight adapter module consisting of an updatable parameter matrix to the output layer of the base model; The lightweight adapter module is a plug-in incremental learning mechanism that optimizes only some newly added components without changing the main structure of the original model. Specifically, the "lightweight adapter" mentioned here is essentially a small neural subnetwork, typically containing two fully connected layers (one for dimensionality reduction and one for dimensionality increase) with an activation function sandwiched between them. Its main responsibility is to locally modify the high-level semantic features generated by the base model, making them better suited to the needs of the target task.

[0084] It should be noted that the "updatable parameter matrix" means that this part of the structure can evolve dynamically with sample feedback, while most of the other basis parameters remain frozen. This ensures both model stability and improves flexibility.

[0085] Step S503: Linear projection of high-order features is performed through the lightweight adapter module to generate prediction results.

[0086] The "linear projection" essentially involves performing an affine transformation operation using the parameter matrix defined internally by the adapter. Assume the feature to be processed is h∈R. d Then the new feature after the adapter can be written as h ′ =Ah+c, where A and c represent the transformation matrix and the offset vector, respectively. This step can be seen as a recalibration of the base model output, helping to eliminate potential errors caused by training bias; at the same time, by using low-rank constraints and other methods, computational costs can be significantly reduced and the risk of overfitting can be prevented.

[0087] Furthermore, thanks to its modular design, the entire model does not need to be retrained for new application scenarios. Simply replacing the corresponding adapter components allows for rapid switching of roles, greatly enhancing the system's scalability.

[0088] In the above implementation, the collaborative architecture of the base model and the lightweight adapter enables efficient integration of pre-trained knowledge with the target task. The deep nonlinear transformation of the base model fully exploits the high-order semantic relationships in the cross-modal fusion features, while the incremental learning mechanism of the lightweight adapter achieves task-specific optimization while maintaining model stability, avoiding the computational overhead of full-scale fine-tuning. This "frozen backbone + adapter fine-tuning" strategy ensures the model's generalization ability while improving parameter efficiency and deployment flexibility, enabling the system to quickly adapt to diverse cross-modal task requirements.

[0089] Reference Figure 6 As a further implementation of the material extraction method, after the step of outputting the prediction results, the method further includes: Step S601: Continuously collect the current predicted distribution of the prediction results and store it in a circular buffer to form a historical distribution sequence; The system needs to obtain the probability distribution or classification label statistics from the prediction module and convert them into a standard form (such as a normalized frequency vector) that can be used for subsequent comparisons. This data is sequentially written into a fixed-length circular buffer. Whenever a new prediction is generated, the oldest historical data is replaced according to a first-in, first-out (FIFO) principle. This design not only effectively controls memory usage but also ensures that the historical window used for comparison is always within the most recent time period, avoiding interference from outdated data in judging recent trends. For example, if the buffer size is set to N, only the predicted distribution samples within the most recent N time slices are retained each time, forming a sliding time window sequence. This allows the system to capture short-term change patterns without over-reliance on static benchmarks.

[0090] Step S602: Calculate the statistical difference between the current predicted distribution and the historical baseline distribution, and generate the real-time distribution offset. The system selects a historical period as a reference benchmark (usually the average distribution over a recent period or an empirical distribution on the initial training set), and then uses statistical tools such as KL divergence to measure the deviation of the current predicted distribution from this benchmark. KL divergence, a classic indicator for measuring the asymmetric difference between two probability distributions, is expressed as: ; Where P is the current predicted distribution and Q is the historical baseline distribution. This formula describes the information loss when using distribution Q to approximate the true distribution P. Since KL divergence does not satisfy the commutative law, two-way metrics or other variants (such as JS divergence) are usually considered to obtain more robust results. The numerical values ​​obtained in this way reflect whether potential concept drift or covariate shift has occurred in the current task space, serving as the basis for subsequent decisions.

[0091] Step S603: When the real-time distribution offset exceeds the preset offset threshold, freeze the base model parameters and update the adjustable parameter matrix of the lightweight adapter module to generate the adapter parameter update amount. If the distribution offset exceeds a preset threshold, it means the existing model may no longer accurately depict the relationship between the current input features and the target output. In this case, the system will freeze the base model parameters. "Freezing" means preventing the backbone network weights from continuing to participate in the backpropagation process, thus maintaining the original knowledge system. Simultaneously, the adjustable parameter matrix of the lightweight adapter module is updated to quickly respond to new situations without compromising the overall architecture's stability. The core idea of ​​this strategy is to separate general capabilities from specific adaptation capabilities: the former is obtained through large-scale pre-training and remains effective in most scenarios, while the latter is fine-tuning for local perturbations. These two are independent yet complementary, forming a flexible and efficient incremental learning framework.

[0092] Specifically, lightweight adapter modules are typically nested at intermediate nodes of the original neural network as insertable add-on layers. They contain a small number of trainable parameters and can introduce additional transformation capabilities without affecting the main structure. To improve efficiency and specificity, update operations are not performed globally but are limited to a selected range.

[0093] In this embodiment of the application, the low-rank gradient update matrix can be constructed as follows: ; Where η represents the dynamic learning rate, which automatically adjusts the pace according to the current convergence speed; L CE This represents the cross-entropy loss function, reflecting the current model error level; This refers to the gradient direction of the corresponding parameter; the Hadamard product operation represented by the symbol ⊙, together with the parameter importance mask matrix M, determines which dimensions should be adjusted first. This mask M determines the position of key parameters by accumulating the evaluation of gradient magnitudes over several rounds, allowing only those directions that change frequently and significantly to retain update privileges, while the rest are temporarily ignored, thereby achieving the purpose of dimensionality reduction and efficiency improvement.

[0094] Step S604: Calculate the association weight adjustment coefficient of the graph structure storage layer based on the adapter parameter update amount; Specifically, the system maps changes in internal parameters from the deep learning domain to an external knowledge base structure. To achieve this, the system extracts the Frobenius norm of the aforementioned ΔW matrix. As a comprehensive indicator of the intensity of change, it is further enhanced by the Sigmoid function. This is transformed into a standardized coefficient λ in the interval [0,1], where k is a scaling factor and θ is an offset constant. This function exhibits good nonlinear smoothing properties, effectively suppressing the side effects of extreme fluctuations. Simultaneously, the two control factors k and θ provide users with a degree of freedom to balance sensitivity and stability requirements. The standardized coefficient λ transforms the originally difficult-to-perceive high-dimensional tensor changes into easily understood and manipulated one-dimensional real-valued signals, facilitating their application to heterogeneous system interfaces.

[0095] Step S605: Based on the correlation weight adjustment coefficient, call the graph structure database interface to perform the correlation weight update operation, and output the updated lightweight adapter parameters and correlation weight set.

[0096] The association weight update operation specifically includes: retrieving the set of superedges related to the current prediction task; and updating the superedge weights according to the formula: Where δ is the preset weight step value.

[0097] Specifically, based on the task context, a subset of hyperedges E_old closely related to the current prediction intent is retrieved. Then, incremental adjustments are applied to each hyperedge according to the update rule E_new = E_old + λ·δ, where δ is a pre-defined small step size. This method balances the need for precise matching with the principle of computational economy, avoiding performance bottlenecks caused by blindly refreshing the entire domain. Finally, the new adapter parameter matrix produced in this cycle, along with the updated hyperedge weight list, is packaged and returned to the front-end controller or logging unit for later use in the next cycle or for archiving and auditing.

[0098] In the above implementation, a circular buffer is used to maintain the historical predicted distribution sequence, and KL divergence is combined to monitor distribution shifts in real time and trigger an adaptive parameter update mechanism, thus achieving a dynamic response to concept drift. The system adopts a low-rank gradient update strategy, performing incremental optimization only on the lightweight adapter without affecting the stability of the base model. At the same time, parameter changes are mapped to the adjustment of association weights in the graph-structured database, forming a closed-loop feedback mechanism from the prediction layer to the knowledge layer. This design not only ensures the model's continuous adaptability to distribution changes, but also achieves a balance between computational efficiency and system stability through modular decoupling and standardized coefficient transformation, significantly improving the robustness and scalability of cross-modal tasks in dynamic environments.

[0099] Reference Figure 7As a further implementation of the material extraction method, after the step of converting the prediction results into structured data in the target format and outputting it, the method further includes: Step S701: Obtain user interaction data from the terminal device, including click rate, dwell time, and operation logs; The core objective of this step is to establish a data-driven mapping between users' true intentions and preferred behaviors. The system utilizes multiple terminal interface methods and employs differentiated data collection strategies based on different types of terminal devices (such as PCs, mobile devices, and augmented reality devices) to ensure the completeness and timeliness of behavioral information.

[0100] For example, for PCs, the WebSocket protocol is the preferred transmission channel because it supports bidirectional communication; for mobile devices, the operating system-level behavior monitoring API is used to capture atomic operation events such as swiping and clicking; as for AR devices, since they involve complex gesture recognition tasks in three-dimensional space, they need to access a dedicated spatial interaction event bus to extract high-dimensional motion trajectory data.

[0101] In this embodiment, the collected behavioral data mainly includes three dimensions: click-through rate (CTR), dwell time, and operation logs. CTR reflects the user's interest in the prediction result, specifically the proportion of valid clicks to total exposures. Dwell time reflects the relevance and attractiveness of the content, typically recorded by client-side tracking from the start of display to page closing. Operation logs cover deeper operational behaviors, such as secondary processing actions like saving, forwarding, and editing, and are generally organized in JSON format, including event type, timestamp, and other contextual attribute fields.

[0102] It should be noted that all the above data carries a request source identifier, enabling the system to automatically deduce the corresponding collection rule chain based on the device category, forming a terminal adaptation mechanism with inheritance characteristics.

[0103] Step S702: Analyze the correlation between user interaction data and prediction results, and calculate and generate performance evaluation indicators; The performance evaluation metrics include prediction accuracy, response latency deviation, and semantic matching degree. Specifically, the system designs a set of multi-dimensional evaluation functions, each corresponding to different business focuses. The first is prediction accuracy, calculated based on the classic confusion matrix theory framework. This involves treating actual user clicks as positive samples and non-clicks as negative samples. The system then counts the number of true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN), finally substituting these values ​​into the formula Acc = (TP + TN) / (TP + FP + FN + TN) to obtain a numerical score. This classification accuracy metric helps to intuitively understand whether the model can correctly distinguish content that users are interested in.

[0104] Secondly, there is the response latency deviation ΔT, which measures the system's real-time service capability. It is calculated as the absolute error between the actual response time T_real and the pre-defined standard value T_SLA specified in the Service Level Agreement (SLA): ΔT = |T_real - T_SLA|. If the average latency on a certain type of terminal is found to exceed the tolerable range, it indicates potential network congestion, resource contention, or other performance bottlenecks, requiring further investigation to pinpoint the root cause.

[0105] The third key metric is the semantic matching degree J(A,B), which uses the Jaccard similarity coefficient to characterize the semantic fit between the predicted output and the user's potential interests. A represents the set of result labels provided by the model, and B is the set of implicit demand labels formed by clustering keywords in the user's operation logs. Dividing the intersection of the two by their union yields a decimal between 0 and 1; the closer to 1, the higher the semantic overlap, and vice versa, indicating a significant risk of semantic drift. By combining these three metrics, the system can comprehensively understand the current model performance and make further optimization decisions accordingly.

[0106] Step S703: Based on the comparison results between the performance evaluation index and the preset performance threshold, a parameter optimization instruction set is generated; wherein, the parameter optimization instruction set includes the base model retraining trigger flag, the lightweight adapter parameter update direction, and the storage layer weight adjustment coefficient. Specifically, this process is essentially a typical application of cybernetics, namely, determining whether to activate a specific control loop by setting a series of decision conditions.

[0107] In this embodiment, if the semantic matching degree J is detected to be less than a predetermined threshold θ_J (assumed to be 0.5) multiple times consecutively within a certain time period, it indicates that the existing model cannot accurately capture the user's true semantic needs. In this case, a parameter update operation of the lightweight adapter module should be initiated immediately. To improve convergence efficiency and avoid oscillations, the system constructs the gradient descent direction based on the partial derivative of the objective function with respect to the weights W. η represents the learning rate hyperparameter, which determines the step size of each iteration; This represents the gradient of the loss function J with respect to the weights W.

[0108] On the other hand, if the prediction accuracy acc is observed to be lower than the threshold θacc (e.g., 0.85) and the response delay deviation ΔT exceeds the maximum allowable deviation θdelay (e.g., 50 milliseconds), it indicates that not only has the model's discrimination ability declined significantly, but the response speed has also failed to meet expectations. In this case, a more aggressive response strategy must be adopted, such as activating the base model retraining flag flag_retrain=1, and notifying the background scheduling engine to prepare to reload the large-scale training task.

[0109] In addition, to address the issue of degraded storage layer access performance, a weight adjustment factor that decreases exponentially with increasing latency has been introduced. , where k controls the steepness of the sensitivity curve slope, which can be used to guide the changes in the connection strength of each hyperedge on the hotspot path in the graph structure database.

[0110] The above three subcommands together constitute the complete parameter optimization instruction set I={ W, flag_retrain, λ}, and are arranged and managed by a finite state machine to ensure good causal consistency and controllable execution order among the instructions.

[0111] Step S704: If the base model retraining trigger flag is activated, retrieve the set of historical associated data blocks from the graph structure storage layer and construct the incremental training dataset. This step first requires quickly locating the historical associated data block S_graph at a specified position within the graph structure storage layer. The retrieval is based on the location identifier attached when the prediction request was initially initiated. In addition, two other layers of auxiliary material are needed, taking into account the time series characteristics: first, frequently accessed active data segments S_mem in the memory cache (obtainable through LRU-K filtering); and second, aging data packets S_cold stored in cold storage media for a long time but potentially still valuable (which need to be decompressed and restored to their original form after compression).

[0112] To unify modal differences among data from different sources and eliminate the bias caused by missing samples, a series of preprocessing steps are required, including but not limited to time window alignment (interpolating every 10ms to fill in blank intervals), noise filtering and denoising, feature scale normalization, and imbalanced class distribution correction (such as the SMOTE oversampling algorithm).

[0113] Finally, the resulting incremental dataset is integrated. It not only covers the latest user behavior dynamics but also retains sufficient historical background information, facilitating the integration and transfer of new and old knowledge. More importantly, each sample is accompanied by its original source hash code and its corresponding modality tag metadata, which facilitates later source tracing, auditing, and model version tracking.

[0114] Step S705: Under the constraint of freezing the main parameters of the base model, perform online fine-tuning of the lightweight adapter module using the incremental training dataset to generate updated lightweight adapter parameters. The system employs a low-rank approximation approach to decompose and reconstruct the internal parameter matrix W_adapter of this module, representing it as the product of two smaller matrices A and B: This approach significantly reduces the number of parameters to be trained (theoretically by over 90%) while maintaining the original expressive power. Here, d represents the dimension of the original adapter weight matrix, and r represents the rank of the decomposition.

[0115] In terms of loss function design, the InfoNCE loss L_con under the contrastive learning paradigm is selected to guide the learning direction of the feature embedding space, encouraging the reduction of distance between similar samples and the widening of distance between dissimilar samples. The specific formula is as follows: ; Where sim() is defined as the cosine similarity operation, τ is the temperature scaling factor, and x i For the target sample, x j For its positive sample partner, x k This represents a group of negative sample candidates.

[0116] Furthermore, gradient updates are only allowed to be applied to matrix B throughout the entire fine-tuning process. A remains frozen to maintain the overall stability of the system. After one iteration, it is also necessary to check whether the change in KL divergence ΔDKL of the predicted probability distribution before and after the update is less than a certain minimum threshold. Only when this constraint is met can it be confirmed that there is no risk of semantic mutation or overfitting.

[0117] Step S706: Adjust the coefficients according to the storage layer weights, call the graph structure database interface to dynamically adjust the association weights, and output the system optimization report to the monitoring platform.

[0118] This step primarily focuses on optimizing resource allocation at the underlying infrastructure level, with the aim of improving the data read and write efficiency of high-frequency query paths.

[0119] In this embodiment, the storage layer weight adjustment coefficient is first parsed and applied to the weight w_old update process on the hyperedge E_target in the graph database. The new weight... δ is used as a fixed-step increment factor (default value is 0.1) to gradually guide the weights to evolve in a direction that favors high-frequency access. Simultaneously, the optimization results for this round must be uploaded, including updated lightweight adapter parameters, the latest hyperedge weight mapping table, and a summary of performance evaluation scores.

[0120] The above implementation constructs a complete closed loop for user behavior data collection, analysis, and feedback optimization, enabling the adaptive evolution of the intelligent system. The system collects user interaction data from multiple dimensions and calculates key performance indicators such as prediction accuracy, response latency deviation, and semantic matching degree. Based on preset thresholds, it automatically triggers parameter optimization instruction sets, including online fine-tuning of the lightweight adapter and incremental training of the base model. Furthermore, it optimizes data access efficiency through dynamic weight adjustment of the graph-structured database, ultimately forming a continuously improving intelligent service system that effectively enhances prediction accuracy, response speed, and user experience.

[0121] As one implementation method for dynamically adjusting associated weights, the specific steps include: parsing the hyperedge influence factor in the weight adjustment coefficient of the storage layer; locating the subset of hyperedges related to the current prediction task in the storage layer of the graph structure; linearly scaling the weight values ​​according to the hyperedge influence factor; and recording the weight adjustment history to the distributed log system.

[0122] First, in the stage of resolving the hyperedge influence factor, the system calculates the storage layer weight adjustment coefficient γ through a nonlinear transformation channel. This coefficient is based on the Frobenius norm of the lightweight adapter parameter update and is smoothed using the Sigmoid function in conjunction with the sensitivity scaling factor k and translation threshold θ, compressing the intensity of the original parameter changes to a stable range. Subsequently, the hyperedge influence factor λ is extracted by multiplying by a preset step size δ (e.g., 0.1), and a binary bit pattern resolution mechanism is used to ensure that the influence factor is accurately bound to the corresponding hyperedge, thereby enhancing the robustness and control accuracy of the system.

[0123] Secondly, in the stage of locating relevant hyperedge subsets, the system adopts a semantic-driven topology matching method. It obtains the task description vector V_task from the output layer of the preceding lightweight adapter network, using it as a query benchmark to perform an approximate nearest neighbor search algorithm (such as HNSW) in a large-scale heterogeneous graph structure, quickly identifying a set of candidate data nodes N_k. Based on these nodes, it expands the probe to detect the hyperedges connected to them, forming a candidate hyperedge set E_candidate. Then, by calculating the comprehensive similarity sim(e_i) between each candidate hyperedge e_i and V_task (based on the average cosine distance), it filters out the target hyperedge set E_target whose similarity exceeds a preset threshold (e.g., 0.8). This method optimizes the retrieval time complexity from O(|E|) to close to O(log|E|), significantly improving efficiency.

[0124] Next, in the stage of linearly scaling the weight values ​​according to the hyperedge influence factor, the system transforms the hyperedge influence factor λ into actual weight adjustment actions. This is achieved by adding an incremental term to the original weight w_old, which is jointly determined by λ, the sign function sgn(Δ) (indicating the direction of model update), and the basic adjustment granularity δ. The weight adjustment adopts a linear superposition method to avoid the risk of numerical overflow caused by proportional amplification, and boundary constraints are applied to the new weight w_new to maintain global stability and fairness.

[0125] Finally, during the weight adjustment history recording phase, the system constructs a standardized log structure. Each log entry contains a 5-tuple consisting of a unique identifier for the superedge, the old and new weight values, a timestamp, and the superedge influence factor λ. Persistent storage is achieved using a distributed architecture (such as a Kafka message queue), and the columnar file format Parquet is used to optimize read performance. The MurMurHash algorithm is used for consistent hashing of superedge IDs to ensure the sequential operation of the same superedge, and a minimum change threshold ε (e.g., 0.001) is set to suppress redundant logs.

[0126] The above technical solution achieves adaptive optimization of the prediction model and the graph storage system through a closed-loop framework of "model perception - semantic matching - weight adjustment - behavior archiving".

[0127] This application also discloses a material extraction system based on large model and multi-storage technology.

[0128] A material extraction system based on large model and multi-storage technology, specifically including: The heterogeneous data stream standardization module is used to receive heterogeneous data streams from text, image, video and audio sources, perform protocol parsing on the heterogeneous data streams, and generate standardized data blocks. The routing feature extraction module is used to extract the data volume features, access frequency features, and semantic entropy value features of standardized data blocks, and combine them into a routing feature vector. The intelligent routing decision module is used to perform weighted calculations on the routing feature vectors based on the dynamically updated weight matrix to generate a routing decision score. The hierarchical storage allocation module is used to allocate standardized data blocks to one of the following levels: memory storage layer, graph structure storage layer, or compressed storage layer, based on routing decision scores, and output the corresponding storage location identifier. The cross-modal retrieval and fusion module is used to respond to query requests from external clients, retrieve related data from hierarchical storage layers based on storage location identifiers, align multimodal features through attention mechanisms and spatiotemporal convolution operations, and generate cross-modal fusion feature vectors. The multimodal prediction output module is used to input cross-modal fused feature vectors into a pre-trained base model and output the prediction results through a lightweight adapter module. The device adaptation and conversion module is used to convert the prediction results into structured data in the target format and output them according to the device type of the external requesting party.

[0129] The material extraction system based on large model and multi-storage technology in this application embodiment can implement any of the above methods, and the specific working process of each module in the system can refer to the corresponding process in the above method embodiment.

[0130] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.

[0131] This application also discloses a computer-readable storage medium.

[0132] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in any of the material extraction methods based on large model and multiple storage technologies.

[0133] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0134] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.

Claims

1. A method for extracting materials based on large model and multi-storage technology, characterized in that, The material extraction method includes: Receive heterogeneous data streams from text, image, video, and audio sources, perform protocol parsing on the heterogeneous data streams, and generate standardized data blocks; The data volume features, access frequency features, and semantic entropy value features of the standardized data blocks are extracted and combined into a routing feature vector; among them, the data volume feature is a quantitative indicator of the space occupied by the data block; The routing feature vector is weighted according to the dynamically updated weight matrix to generate a routing decision score; Based on the routing decision score, the standardized data block is allocated to one of the following levels: memory storage layer, graph structure storage layer, or compressed storage layer, and the corresponding storage location identifier is output. In response to a query request from an external client, the system retrieves associated data from the hierarchical storage layer based on the storage location identifier, aligns multimodal features through an attention mechanism and spatiotemporal convolution operations, and generates a cross-modal fusion feature vector. The cross-modal fusion feature vector is input into the pre-trained base model, and the prediction result is output through the lightweight adapter module; Based on the device type of the external requesting party, the prediction results are converted into structured data in the target format and output.

2. The material extraction method based on large model and multi-storage technology according to claim 1, characterized in that, The steps for generating a routing decision score by weighting the routing feature vectors based on a dynamically updated weight matrix include: Obtain a routing feature vector, which includes data volume features, access frequency features, and semantic entropy value features; Read the currently stored weight matrix; It receives response latency data and service level agreement thresholds collected during system operation in real time and calculates the reward function value. Based on the reward function value, the elements in the weight matrix are incrementally modified using the Q-learning algorithm to periodically update the weight matrix; The routing feature vector and the weight matrix are multiplied together to generate a routing decision score.

3. The material extraction method based on large model and multi-storage technology according to claim 1, characterized in that, The steps of allocating the standardized data block to a level one of the memory storage layer, graph structure storage layer, or compressed storage layer based on the routing decision score, and outputting the corresponding storage location identifier, include: Obtain the routing decision score and the corresponding standardized data block; The routing decision score is compared with preset hot storage threshold, warm storage threshold and cold storage threshold; When the routing decision score is greater than the hot storage threshold, the standardized data block is stored in the memory storage layer, and a hot storage layer identifier containing the memory node address is generated. When the routing decision score is between the warm storage threshold and the cold storage threshold, the standardized data block is stored in the graph structure database corresponding to the graph structure storage layer, feature vector nodes are constructed and connected to cross-modal hyperedges, and a warm storage layer identifier is generated. When the routing decision score is less than the cold storage threshold, lossless compression encoding is performed on the standardized data block, the compressed data is stored in the compressed storage layer, and a cold storage layer identifier is generated.

4. The material extraction method based on large model and multi-storage technology according to claim 1, characterized in that, The steps of responding to a query request from an external client, retrieving associated data from the hierarchical storage layer based on the storage location identifier, aligning multimodal features through an attention mechanism and spatiotemporal convolution operations, and generating a cross-modal fusion feature vector include: Receive query requests from external requesters and parse the semantic keywords and requester device types contained in the query requests; The storage level of the target data is determined based on the storage location identifier, and the corresponding original data block is obtained. A multi-head attention mechanism is performed on the text modal data in the original data block to extract text feature vectors related to semantic keywords; Perform a spatiotemporal separation convolution operation on the visual modality data in the original data block to extract spatiotemporally correlated visual feature vectors; The text feature vector and the visual feature vector are fused using tensors to generate a cross-modal fused feature vector.

5. The material extraction method based on large model and multi-storage technology according to claim 4, characterized in that, The steps of determining the storage level of the target data based on the storage location identifier and obtaining the corresponding original data block include: If the target data is located in the memory storage layer, the original data block is read directly through the cache interface; If the target data is located in the graph structure storage layer, execute the hypergraph traversal algorithm to retrieve the original data blocks connected by associated nodes and hyperedges; If the target data is located in the compressed storage layer, the decompression engine is invoked to restore the data block to its original format, thus obtaining the original data block.

6. The material extraction method based on large model and multi-storage technology according to claim 4, characterized in that, The steps of inputting the cross-modal fused feature vector into the pre-trained base model and outputting the prediction result through the lightweight adapter module include: The cross-modal fusion feature vector is input into a pre-trained base model, and high-order feature transformation is performed through a fully connected layer; A lightweight adapter module consisting of an updatable parameter matrix is ​​connected to the output layer of the base model; The high-order features are linearly projected using the lightweight adapter module to generate prediction results.

7. The material extraction method based on large model and multi-storage technology according to claim 6, characterized in that, Following the step of outputting the prediction results is: The current predicted distribution of the prediction results is continuously collected and stored in a circular buffer to form a historical distribution sequence; Calculate the statistical difference between the current predicted distribution and the historical baseline distribution to generate the real-time distribution offset; When the real-time distribution offset exceeds the preset offset threshold, the base model parameters are frozen and the adjustable parameter matrix of the lightweight adapter module is updated to generate the adapter parameter update amount. Calculate the association weight adjustment coefficient of the graph structure storage layer based on the adapter parameter update amount; Based on the association weight adjustment coefficient, the graph structure database interface corresponding to the graph structure storage layer is called to perform the association weight update operation, and the updated lightweight adapter parameters and association weight set are output.

8. A method for extracting materials based on large model and multi-storage technology according to any one of claims 1 to 7, characterized in that, After the step of converting the prediction results into structured data in the target format and outputting it, the method further includes: Acquire user interaction data from terminal devices, including click-through rate, dwell time, and operation logs; Analyze the correlation between the user interaction data and the prediction results, and calculate and generate performance evaluation metrics; Based on the comparison results between the performance evaluation index and the preset performance threshold, a parameter optimization instruction set is generated; wherein, the parameter optimization instruction set includes a base model retraining trigger flag, a lightweight adapter parameter update direction, and a storage layer weight adjustment coefficient; If the base model retraining trigger flag is activated, retrieve the set of historical associated data blocks from the graph structure storage layer and construct an incremental training dataset. Under the constraint of freezing the main parameters of the base model, the online fine-tuning operation of the lightweight adapter module is performed using the incremental training dataset to generate updated lightweight adapter parameters; Based on the storage layer weight adjustment coefficient, the graph structure database interface corresponding to the graph structure storage layer is called to dynamically adjust the association weights, and a system optimization report is output to the monitoring platform.

9. A material extraction system based on large model and multi-storage technology, characterized in that, The material extraction system includes: The heterogeneous data stream standardization module is used to receive heterogeneous data streams from text, image, video and audio sources, perform protocol parsing on the heterogeneous data streams, and generate standardized data blocks. The routing feature extraction module is used to extract the data volume features, access frequency features, and semantic entropy value features of the standardized data blocks and combine them into a routing feature vector; wherein, the data volume feature is a quantitative indicator of the space occupied by the data block; The intelligent routing decision module is used to perform weighted calculations on the routing feature vectors based on a dynamically updated weight matrix to generate a routing decision score. The hierarchical storage allocation module is used to allocate the standardized data block to a first-level storage layer among the memory storage layer, graph structure storage layer, or compressed storage layer based on the routing decision score, and output the corresponding storage location identifier. The cross-modal retrieval and fusion module is used to respond to query requests from external requesters, retrieve related data from the hierarchical storage layer according to the storage location identifier, align multimodal features through attention mechanism and spatiotemporal convolution operation, and generate cross-modal fusion feature vectors. The multimodal prediction output module is used to input the cross-modal fused feature vector into the pre-trained base model and output the prediction result through the lightweight adapter module. The device adaptation and conversion module is used to convert the prediction result into structured data in the target format and output it according to the device type of the external requesting end.

10. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-user data storage docking and secure transmission method based on AI

    CN120547383A

  • Advanced systems and methods for multimodal ai: generative multimodal large language and deep learning models with applications across diverse domains

    US20250272534A1