A file retrieval method and system based on cloud computing

By using context-aware semantic analytical model and dynamic index map in a cloud-based archive search system, combining multi-layer projection and multi-modal semantic fusion algorithm, the problems of low retrieval efficiency and poor user experience in the existing technology are solved, and more efficient and personalized archive search results are achieved.

CN120086414BActive Publication Date: 2025-06-27HANGZHOU YUNJIA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510578284.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-27
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The existing cloud-based archive search methods have problems such as low search efficiency, poor user experience and inability to effectively deal with diversified information needs, especially in aspects such as insufficient user retrieval intention analysis, insufficient historical behavior and context information analysis, low index efficiency, lack of semantic relationship understanding and insufficient multimodal information processing capabilities.

Method used

The context-aware semantic analysis model is used to receive user search requests, combine user historical behavior and domain database analysis search intentions to generate user search semantic vectors. Build a dynamically updated archive data index graph, represent the archive entries and the semantic relationships between them through nodes and edges, and deploy them in the cloud distributed computing framework. Multi-layer projection calculations are used to locate the search domain in the index graph, and candidate archives are filtered and sorted using a matching algorithm for multimodal semantic fusion within the search domain.

Benefits of technology

It improves the accuracy, efficiency and user experience of archive search, can better understand users' search intentions and behaviors, and provides more relevant and personalized search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086414B_ABST
    Figure CN120086414B_ABST
Patent Text Reader

Abstract

The present invention discloses an archive retrieval method and system based on cloud computing. The method includes: receiving the retrieval request information input by a user, using a context-aware semantic parsing model, parsing the retrieval intention according to the retrieval content, the user's historical retrieval behavior and the retrieval domain database in the retrieval request information, and generating a user retrieval semantic vector in combination with the retrieval intention; constructing an archive data index graph for the archive data distributedly stored in a cloud storage environment; dynamically positioning a search domain related to the retrieval semantics in the archive data index graph in a manner of multi-layer projection calculation according to the user retrieval semantic vector; and within the search domain, screening a candidate archive set that meets the query intention by using a multi-modal semantic fusion matching algorithm and sorting and displaying the same. By using the embodiment of the present invention, the accuracy, efficiency and user experience of archive retrieval can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data retrieval, and particularly relates to an archive retrieval method and system based on cloud computing. Background Art

[0002] In the information age, archive management and retrieval systems are widely used in multiple fields such as government, enterprises, and scientific research. However, traditional archive retrieval methods are increasingly showing limitations due to problems such as low retrieval efficiency, poor user experience, and inability to effectively handle diverse information needs. With the rapid development of cloud computing technology, it provides a new solution for archive management.

[0003] However, existing archive retrieval methods based on cloud computing still have many deficiencies. For example, the parsing of users' retrieval intentions is insufficient, and the lack of in-depth analysis of historical behaviors and context information leads to insufficient accuracy and relevance of retrieval results. At the same time, in the cloud storage environment, the indexing efficiency of distributed archive data is low, and existing static index structures cannot respond to data changes in a timely manner. In addition, there is a lack of understanding of the semantic relationships between archives and the inability to effectively process multi-modal information such as text and images, resulting in retrieval results that cannot meet the complex needs of users. Finally, existing systems often ignore the utilization of users' behavioral context, and this information is crucial for optimizing retrieval intentions and results. Summary of the Invention

[0004] The purpose of the present invention is to provide an archive retrieval method and system based on cloud computing to solve the deficiencies in the prior art and improve the accuracy, efficiency, and user experience of archive retrieval.

[0005] An embodiment of the present application provides an archive retrieval method based on cloud computing, and the method includes:

[0006] Receiving the retrieval request information input by the user, using the context-aware semantic parsing model, parsing the retrieval intention according to the retrieval content, the user's historical retrieval behavior, and the retrieval domain database in the retrieval request information, and generating a user retrieval semantic vector in combination with the retrieval intention;

[0007] For the archive data distributedly stored in the cloud storage environment, constructing an archive data index graph, wherein the archive data index graph represents archive entries through nodes and represents the semantic or logical relationships between archives through edges, and based on a dynamic update mechanism, the graph structure and indexing efficiency are optimized in real time, and the archive data index graph is deployed in the cloud distributed computing framework;

[0008] According to the user retrieval semantic vector, dynamically locating the search domain related to the retrieval semantics in the archive data index graph in a multi-layer projection calculation manner;

[0009] Within the search domain, a matching algorithm using multi-modal semantic fusion is employed to screen and sort a candidate file set that meets the query intent for display. Among them, the screening process comprehensively considers file semantic relevance, file storage location characteristics, and user behavior context.

[0010] Optionally, the method of receiving the retrieval request information input by the user, using a context-aware semantic parsing model, parsing the retrieval intent according to the retrieval content, user historical retrieval behavior, and retrieval domain database in the retrieval request information, and generating a user retrieval semantic vector in combination with the retrieval intent includes:

[0011] Mapping the retrieval request information input by the user into a semantic space of a fixed dimension through a BERT or Transformer model to generate a semantic embedding vector;

[0012] Obtaining the user's historical retrieval records, analyzing historical retrieval keywords, document types, and the user's click behavior, constructing a user profile, and converting the user profile data into a user preference embedding vector;

[0013] Accessing the retrieval domain database to determine the domain knowledge related to the retrieval request information, and extracting domain feature embedding vectors;

[0014] Using an algorithm based on the attention mechanism to fuse the semantic embedding vector, user preference embedding vector, and domain feature embedding vector, perform retrieval intent parsing, assign weights to each parsed intent, and generate a user retrieval semantic vector based on the weights, where the user retrieval semantic vector covers the current retrieval intent and reflects the user's individual needs and historical behavior characteristics.

[0015] Optionally, for the archival data stored distributively in a cloud storage environment, an archival data index graph is constructed, where the archival data index graph represents archival entries through nodes and represents the semantic or logical relationships between archives through edges, and the graph structure and indexing efficiency are optimized in real time based on a dynamic update mechanism, and the archival data index graph is deployed in a cloud distributed computing framework, including:

[0016] Using natural language processing technology to extract semantic features of archival entries from archival data, taking each archival entry as a node in the archival data index graph structure, and assigning corresponding semantic features of the archival entry to each node;

[0017] Creating edges in the archival data index graph structure according to the semantic or logical relationships between archival entries, and setting the initial weight of each edge to obtain the archival data index graph, where the weight is initialized by comparing the similarity of two archival entries to reflect the strength of the edge;

[0018] Based on the dynamic update mechanism, track changes in the content, metadata, and their relationships of the archives to ensure the real-time nature of the archive data index graph. Among them, when a new archive is added or an existing archive is modified, an incremental update strategy is used to only update the affected nodes and edges, and the weights of the edges are dynamically adjusted according to time and user behavior to gradually optimize the graph structure and query efficiency;

[0019] Select a distributed computing framework in the cloud computing environment and deploy the real-time archive data index graph in the distributed graph database of the selected distributed computing framework.

[0020] Optionally, according to the user retrieval semantic vector, in the archive data index graph, in a multi-layer projection calculation manner, dynamically locate the search domain related to the retrieval semantics, including:

[0021] Use the graph embedding algorithm to perform pre-training on the archive data index graph to generate node embeddings for each node in the graph. Among them, perform deep feature representation on each archive entry node so that the position of each node in the embedding space can reflect its own attributes and semantic or logical relationships with other nodes at the same time. The embedding space is the space composed of node embeddings;

[0022] Use the deep alignment model to map the user retrieval semantic vector into the embedding space so that the mapped user retrieval semantic vector and the node embeddings in the archive data index graph are in the same space;

[0023] Calculate the similarity between the mapped user retrieval semantic vector and all node embeddings, and select the K nodes in the archive data index graph that are most similar to the user retrieval semantic vector as seed nodes to represent the starting point of the initial search;

[0024] According to the adjacent nodes of each seed node, expand the first-hop neighbor nodes outward, and judge whether the similarity of these neighbor nodes to the user retrieval semantic vector reaches a specific threshold. If the similarity of the neighbor nodes does not reach the specific threshold, stop expanding. If the similarity reaches the specific threshold, add them to the candidate node set. And for the new neighbor nodes, repeat the expansion process and continue to explore the next-hop neighbor nodes outward until the preset layer limit is reached or the condition of accumulating the number of relevant nodes is met;

[0025] Determine the set of all candidate nodes obtained through multi-layer projection expansion as the search domain. The search domain is a subgraph of the archive data index graph and contains graph structure information and the complete semantic embeddings of the candidate nodes.

[0026] Optionally, within the search domain, use a multi-modal semantic fusion matching algorithm to screen and sort the candidate archive sets that meet the query intent, including:

[0027] For each archive entry that may be related to the retrieval semantics within the search domain, extract multi-modal feature vectors;

[0028] Use a multi-modal fusion algorithm to fuse the feature vectors of different modalities. Among them, weighted summation, attention mechanism or deep neural network is used to achieve the fusion of features, and generate the fused feature vectors of each archive entry, which are used to represent the multi-modal semantic information of each archive entry;

[0029] Calculate the similarity between the fused feature vector of each archive entry and the user's retrieval semantic vector, and filter out the archive entries with similarity higher than the preset threshold as the candidate archive set. The candidate archive set is the archive entry most relevant to the user's query intention;

[0030] Sort the candidate archive set according to similarity, user preference, archive importance and freshness, and display the sorted archive entries to the user.

[0031] Another embodiment of the present application provides an archive retrieval system based on cloud computing. The system includes:

[0032] A receiving module for receiving the retrieval request information input by the user, using a context-aware semantic parsing model to parse the retrieval intention according to the retrieval content, user's historical retrieval behavior and retrieval domain database in the retrieval request information, and generating a user's retrieval semantic vector in combination with the retrieval intention;

[0033] A construction module for constructing an archive data index graph for the archive data distributedly stored in the cloud storage environment. Among them, the archive data index graph represents archive entries through nodes and represents the semantic or logical relationship between archives through edges, and optimizes the graph structure and indexing efficiency in real time based on a dynamic update mechanism, and deploys the archive data index graph in the cloud distributed computing framework;

[0034] A positioning module for dynamically positioning the search domain related to the retrieval semantics in the archive data index graph in the form of multi-layer projection calculation according to the user's retrieval semantic vector;

[0035] A screening module for screening and sorting and displaying the candidate archive set that meets the query intention within the search domain by using a matching algorithm for multi-modal semantic fusion, where the screening process comprehensively considers the semantic relevance of the archive, the characteristics of the archive storage location and the user behavior context.

[0036] Another embodiment of the present application provides a storage medium in which a computer program is stored. Among them, the computer program is set to execute the method described in any one of the above when running.

[0037] Another embodiment of the present application provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any one of the above.

[0038] Compared with the prior art, a cloud computing-based file retrieval method provided by the present invention receives retrieval request information input by a user, uses a context-aware semantic parsing model, analyzes the retrieval intention according to the retrieval content, the user's historical retrieval behavior, and the retrieval domain database in the retrieval request information, and generates a user retrieval semantic vector in combination with the retrieval intention; constructs a file data index graph for the file data distributedly stored in the cloud storage environment; dynamically locates a search domain related to the retrieval semantics in the file data index graph in a multi-layer projection calculation manner according to the user retrieval semantic vector; within the search domain, uses a multi-modal semantic fusion matching algorithm to screen and sort and display a candidate file set that meets the query intention, thereby being able to improve the accuracy, efficiency, and user experience of file retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a hardware structure block diagram of a computer terminal for a cloud computing-based file retrieval method provided by an embodiment of the present invention;

[0040] Figure 2 is a schematic flowchart of a cloud computing-based file retrieval method provided by an embodiment of the present invention;

[0041] Figure 3 is a schematic structural diagram of a cloud computing-based file retrieval system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention, and cannot be construed as limiting the present invention.

[0043] An embodiment of the present invention first provides a cloud computing-based file retrieval method, which can be applied to an electronic device, such as a computer terminal, specifically, a general computer, etc.

[0044] The following takes running on a computer terminal as an example for a detailed description. Figure 1 is a hardware structure block diagram of a computer terminal for a cloud computing-based file retrieval method provided by an embodiment of the present invention. As Figure 1 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.

[0045] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions which, when executed, can cause the processor to execute any cloud computing-based file retrieval method.

[0046] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0047] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it can cause the processor to execute any cloud computing-based file retrieval method.

[0048] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 1 the structure shown in

[0049] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0050] See Figure 2 , the embodiments of the present invention provide a cloud computing-based file retrieval method, which may include the following steps:

[0051] S201, receive the retrieval request information input by the user, use the context-aware semantic parsing model to parse the retrieval intention according to the retrieval content, the user's historical retrieval behavior and the retrieval domain database in the retrieval request information, and generate a user retrieval semantic vector in combination with the retrieval intention;

[0052] Receive the retrieval request information input by the user, and utilize the context-aware semantic parsing model, aiming to deeply understand the user's retrieval intention by analyzing the content in the retrieval request, the user's historical retrieval behavior, and the relevant domain database. This process includes interpreting the user's input at the semantic level, not only focusing on the keywords themselves, but also considering the user's past behavior and retrieval habits, so as to generate a retrieval semantic vector that can accurately reflect the user's current needs. Through such parsing, the system can capture the user's true intention and improve the accuracy of retrieval.

[0053] The core role of this step is to enhance the user's retrieval experience and ensure that the system can provide more relevant retrieval results based on the user's personalized needs. Through context-aware semantic parsing, the system can go beyond simple keyword matching and incorporate the analysis of the user's historical behavior. This method not only optimizes the relevance of the retrieval results, but also provides a solid foundation for subsequent file screening and ranking, thus effectively enhancing the user's satisfaction with the retrieval system.

[0054] Specifically, the retrieval request information input by the user can be mapped to a semantic space with a fixed dimension through the BERT or Transformer model to generate a semantic embedding vector;

[0055] The core of this step lies in using advanced deep learning models such as BERT or Transformer to convert the user's retrieval request into a semantic embedding vector. The system first formats the text input by the user, and then applies the model to analyze its semantics, so as to capture the important information and context in the text. By generating a semantic embedding vector, the system can represent the user's retrieval intention in a standardized way. This vectorized representation enables the system to more effectively analyze and compare different text requests, thereby improving the overall accuracy and efficiency of retrieval.

[0056] In the implementation process, first, the system needs to load a pre-trained BERT or Transformer model. This process includes selecting a suitable model version and loading the model and its parameters through a deep learning framework (such as TensorFlow or PyTorch). Then, when the user inputs a retrieval request, the system preprocesses the request, such as cleaning the input text, removing irrelevant characters, tokenizing, and adding special tokens required by the model. The cleaned text is converted into an input format that the model can accept, that is, an index list, and these indexes correspond to the vocabulary of the model. Thereafter, the model projects the input data onto its embedding layer to generate a preliminary semantic representation.

[0057] After obtaining the initial vector representations, BERT or Transformer processes these vectors through its multi-layer self-attention mechanism to capture the context relationships of the input text. During this process, the model weights each word in the input text and re-encodes it based on its importance in the entire sentence, thereby enhancing the understanding of the user's retrieval intent. Through multiple iterations, the model finally outputs a semantic embedding vector of a fixed dimension, representing the deep meaning and context of the input retrieval request. Therefore, the generated vector not only retains the core information of the user input but also strengthens its semantic features, laying the foundation for the subsequent retrieval process.

[0058] Finally, the generated semantic embedding vector will be stored in a cache for quick retrieval and subsequent use. Meanwhile, the system performs similarity calculations based on this vector to provide the most relevant results to the user's retrieval request. This step emphasizes the importance of the BERT or Transformer model in understanding natural language, ensuring the system's accurate capture of the user's intent, and greatly improving the efficiency and relevance of information retrieval.

[0059] Obtain the user's historical retrieval records, analyze the historical retrieval keywords, document types, and the user's click behavior, construct a user profile, and convert the user profile data into user preference embedding vectors;

[0060] This step focuses on collecting and analyzing the user's historical retrieval behavior to construct a user profile. This process deeply analyzes the keywords, document types, and user click behavior in the past retrieval records to identify the user's preference settings. Constructing a user profile helps to achieve personalized services in the system. By understanding the user's historical retrieval behavior, the system can more accurately understand their needs and then provide customized retrieval results that meet the user's preferences. This process can not only enhance the user experience but also effectively reduce the irrelevance of retrieval results and improve the efficiency of information retrieval.

[0061] In this step, the system first extracts the user's historical retrieval records from the user behavior database. These records include information such as the user's input search terms, timestamps, document types clicked, and their feedback. The system will clean and organize this data to ensure that the noise data (such as duplicate records, invalid requests, etc.) is removed. Next, the system uses natural language processing tools to perform word frequency analysis on the extracted keywords to identify the topics and fields that the user often retrieves and summarize the user's interest preferences.

[0062] Once the analysis of historical retrieval behaviors is completed, the system will construct a user profile. This profile will include multi-dimensional information such as the types of documents preferred by the user (e.g., academic articles, blogs, news reports), frequently used keywords (based on frequency and relevance), click-through rate, and user dwell time. The system represents this information in a normalized vector form to ensure that each user profile can effectively capture their unique retrieval behaviors. To further quantify user preferences, the system can also stratify users through cluster analysis, dividing users into different groups and extracting the common characteristics of each group to generate more representative user preference embedding vectors.

[0063] Finally, this user preference embedding vector will be stored in the user profile and used as the input for the personalized recommendation algorithm. So that when the user conducts a retrieval next time, the system can prioritize the display of the most relevant search results based on their habits and preferences. Through this process, the system can not only improve the user's retrieval satisfaction but also effectively increase the user's return rate, forming a good user interaction experience.

[0064] Access the retrieval domain database to identify domain knowledge related to the retrieval request information in order to extract domain feature embedding vectors;

[0065] In this step, the system needs to access a dedicated retrieval domain database to obtain domain knowledge related to the user's retrieval request. This process not only involves querying the database but also includes processing the retrieval results to extract valuable domain features. Extracting domain feature embedding vectors can enhance the professionalism of the retrieval system, enabling it to respond to user needs promptly and accurately. When the user's request is related to a specific domain, the system can improve the accuracy and relevance of the retrieval results through the domain feature embedding vectors, thereby optimizing the user experience.

[0066] In this step, the system first matches the user's retrieval request with the retrieval domain database to search for relevant domain knowledge. The database should contain rich documents on a specific domain, such as research papers, industry standards, technical manuals, and user guides. The system uses keyword matching and similarity search techniques to retrieve the most relevant documents from the database by processing the user's semantic embedding vectors, and these documents will serve as the basis for subsequent domain feature extraction.

[0067] Subsequently, the system performs text analysis and feature extraction on the retrieved documents. By using natural language processing techniques such as named entity recognition, topic models, and concept extraction, the system can extract important domain features from the documents. These features include key terms, related concepts, and the relationships between them. After extraction, the system transforms these domain features into vector form to create domain feature embedding vectors. To ensure that these vectors can effectively capture the context information within the domain, the system can use a weight adjustment mechanism to weight the features according to their importance in the document, making the finally generated domain feature embeddings more accurate.

[0068] Finally, this domain feature embedding vector will be combined with the previously generated user preference embedding vector and the semantic embedding vector of the retrieval request to provide comprehensive retrieval results. This integrated vector form not only improves the accuracy of the retrieval but also enables the system to better understand the user's needs. Through such a comprehensive feature extraction and vector generation strategy, the system ensures that highly relevant and specialized information is provided during the retrieval, significantly enhancing the user's retrieval experience.

[0069] An algorithm based on the attention mechanism is used to fuse the semantic embedding vector, the user preference embedding vector, and the domain feature embedding vector for retrieval intention parsing. Weights are assigned to each parsed intention, and a user retrieval semantic vector is generated based on the weights, where the user retrieval semantic vector covers the current retrieval intention and reflects the user's personality needs and historical behavior characteristics.

[0070] In this step, the system comprehensively considers the semantic embedding vector of the user's retrieval request, the user's preference embedding vector, and the domain feature embedding vector, and uses an algorithm based on the attention mechanism for fusion processing. By weighted combination of these three types of embedding vectors, the system can refine the user's current retrieval intention. This process not only emphasizes the complementarity of each embedding vector but also reflects its importance in a specific retrieval context by assigning different weights. In addition, the finally generated user retrieval semantic vector will fully cover the current user's retrieval intention and combine the user's personality needs and historical behavior characteristics, thus providing more accurate retrieval results for the user.

[0071] The core of this step is to improve the system's ability to understand and parse the user's retrieval intention by fusing vector information from different sources. Through weighted fusion based on the attention mechanism, the system can dynamically adjust the contribution degrees of various embedding vectors, making the finally generated user retrieval semantic vector not only highly reflect the user's current needs but also take into account their historical behavior and personalized preferences. In this way, the system can more effectively understand the user's complex retrieval intention, thereby providing more relevant search results, further enhancing the user experience, and increasing the user's satisfaction and the system's usage frequency.

[0072] The system first standardizes the previously generated semantic embedding vectors, user preference embedding vectors, and domain feature embedding vectors to ensure they can be compared and fused in the same dimension. Next, the system uses an algorithm based on the attention mechanism to perform weighted fusion on these three types of vectors. In this process, the system dynamically generates attention weights by calculating the similarity between each embedding vector and the other two embedding vectors. Specifically, a similarity metric method, such as cosine similarity or dot product, is used to evaluate the similarity degree between each pair of embedding vectors. In this way, the system can quantify the similarity between different embedding vectors and obtain a set of representative similarity scores. Once the similarity scores are obtained, the system then needs to convert these scores into attention weights. To achieve this, the system can use the softmax function to normalize the similarity scores into weights such that their sum is 1. Through the softmax function, the system can make the embedding vectors with higher similarity obtain larger weights, while the vectors with lower similarity obtain smaller weights, ensuring that strongly relevant vectors play a greater role in the subsequent fusion process.

[0073] After obtaining the attention weights, the system will apply these weights to each embedding vector respectively and sum them up with weights to generate a comprehensive vector representation. This comprehensive vector not only contains the current retrieval intention information but also integrates the user's preference features and domain knowledge, thus forming a user retrieval semantic vector. Through the generation process of this vector, the system can effectively integrate information from different sources and create a retrieval semantic vector with higher information density and semantic depth.

[0074] Finally, the generated user retrieval semantic vector will be used in the subsequent retrieval and recommendation processes. The system will use this vector to compare the similarity with the document vectors in the database to quickly match the retrieval results most relevant to the user's needs. By organically combining the user's individual needs, historical behaviors, and current retrieval intentions, the system can not only improve the precision and efficiency of retrieval but also create a personalized retrieval experience for users, promoting users to use the system more frequently in future retrievals.

[0075] S202. For the archival data stored distributively in the cloud storage environment, construct an archival data index graph, where the archival data index graph represents archival entries through nodes and represents the semantic or logical relationships between archives through edges, and based on a dynamic update mechanism, optimize the graph structure and indexing efficiency in real time, and deploy the archival data index graph in the cloud distributed computing framework;

[0076] In a cloud storage environment, the distributed storage characteristic of archival data makes the retrieval efficiency and flexibility of data particularly important. To this end, constructing an archival data index graph is a crucial step. In this index graph, each archival entry is abstracted as a node, while the semantic or logical relationships between archives are represented by edges. Such a structure can not only effectively organize and index a vast amount of archival data, but also facilitate the implementation of complex query and update operations through the graph structure. In addition, to enhance the flexibility and response speed of the system, the archival data index graph will optimize the graph structure and indexing efficiency in real time through a dynamic update mechanism. This means that when new archives are added or existing archives change, the system will immediately adjust the graph structure to ensure that users can obtain the latest information when retrieving. This index graph is deployed in a cloud distributed computing framework, ensuring high reliability of data access and high concurrent processing capabilities, enabling users to quickly and accurately access the required archival information.

[0077] The process of constructing an archival data index graph plays an important role in the entire retrieval method. By organizing archival data in the form of a graph, the system not only improves the retrieval efficiency and semantic relevance of data, but also achieves more flexible query processing capabilities. The design of the graph structure provides a clear visualization of the relationships between various archives, and different types of relationships such as semantic similarity and logical association can be represented by edges, making the query of complex relationships more intuitive and efficient. In addition, by using the dynamic update mechanism, the system can respond to changes in archival content in real time, ensuring that the information obtained by users during retrieval is the latest and most relevant. This real-time nature is particularly important in the field of archival management because users often need to find accurate information within a short period of time, enhancing the user experience and satisfaction.

[0078] Specifically, natural language processing technology can be used to extract the semantic features of archival entries from archival data, take each archival entry as a node in the archival data index graph structure, and assign corresponding archival entry semantic features to each node;

[0079] In this step, the system uses natural language processing (NLP) technology to extract semantic features from archival data. Each archival entry extracts valuable information from text, metadata, and related documents to form its unique semantic features. These semantic features provide basic data support for the nodes of archival entries in the graph structure, enabling subsequent indexing and retrieval activities to understand and match the user's query intention at a deep level. At the same time, the system assigns corresponding semantic features to each archival entry, making these features the cornerstone of constructing the archival data index graph and ensuring that the graph can effectively reflect the diversity and complexity of archival data.

[0080] Using NLP technology to extract semantic features from archive entries can significantly improve the accuracy and efficiency of archive retrieval. By assigning semantic features to each archive entry, the system can retrieve based on the true meaning of the content rather than surface keywords. This semantic-level understanding greatly enhances the user's ability to obtain information during retrieval, enabling them to find the archives most relevant to their query more quickly. At the same time, in a complex archive relationship network, this feature extraction provides an effective basis for subsequent graph structure optimization and update, further improving the overall performance of the database.

[0081] In this step, first, it is necessary to use natural language processing (NLP) technology to effectively extract semantic features from archive data. Specifically, the system can use pre-trained language models (such as BERT or Word2Vec) to encode the text content of each archive entry, and extract the semantic features of the text through these models. This process includes tokenizing the archive entries, removing stop words, and vectorizing the words, and finally transforming each archive entry into a vector of a fixed dimension to capture its inherent semantic information. At this time, these vectors will be used as the node representations in the archive data index graph.

[0082] Next, the system refines the extracted semantic features to ensure that each node not only contains basic text information but also reflects the context of the archive entry. For example, the system can further enrich the semantic representation of the node according to metadata such as the creation time, modification record, and classification information of the archive. In addition, the system can introduce domain-specific knowledge bases and enhance the accuracy of the semantic features of archive entries by comparing with these knowledge bases to make them more suitable for the needs of specific retrieval domains. Each node thus obtains a feature vector containing rich semantic information, improving the subsequent retrieval effect.

[0083] Finally, during the process of constructing nodes, the system also needs to consider the similarities and differences between various archive entries. By calculating the similarity between different archive entries, such as using methods like cosine similarity or Euclidean distance, the system can further optimize the structure of the nodes to ensure efficient and accurate semantic retrieval in subsequent steps. The semantic features of these nodes will provide a more intelligent retrieval experience for users in subsequent queries and matches.

[0084] Create edges in the archive data index graph structure according to the semantic or logical relationships between archive entries, and set the initial weight of each edge to obtain the archive data index graph, where the weight is initialized by comparing the similarity of two archive entries to reflect the strength of the edge;

[0085] The goal of this step is to create edges in the archival data index graph based on the semantic or logical relationships between archival entries and set initial weights for each edge. Using the extracted archival semantic features, the system analyzes the similarity between each pair of archival entries and establishes and initializes the edge weights by comparing the distance or correlation of their semantic features. This process ensures that the strength of the edges in the graph can truly reflect the connections between archival entries, making subsequent queries and retrievals more accurate and efficient.

[0086] Establishing the edge weights is crucial for the overall performance of the archival data index graph. The weights not only represent the strength of the relationship between archival entries but also affect the node selection and expansion strategies during subsequent retrieval processes. By accurately initializing the weights of each edge, the system can ensure that archival entries with strong semantic relevance are prioritized during retrieval, improving the relevance and accuracy of retrieval results. At the same time, this setting of edge weights based on similarity calculation can provide users with a richer retrieval experience and enhance the intelligence level of information retrieval.

[0087] In the process of creating the archival data index graph, the second step mainly focuses on constructing the edges between archival entries, aiming to reflect the semantic or logical relationships between them. This process first needs to identify the relationships between individual archival entries, such as topic similarity, citation relationships, time series, etc. The system can utilize text similarity calculation to determine the similarity between two archival entries by comparing their semantic feature vectors, thereby providing a basis for establishing edges. For example, if two archival entries have the same topic or keywords, then their correlation is relatively high, and thus an edge can be established in the index graph.

[0088] On this basis, it is crucial to set an initial weight for each edge. The system can use the similarity score as the weight of the edge and, through normalization, ensure that the weight value is between 0 and 1, reflecting the strength and importance of the edge. The magnitude of the edge weight not only represents the closeness between two nodes but also means that in the retrieval results, archival entries with higher relevance will be prioritized. This process of setting edge weights will greatly affect the precision and efficiency of subsequent retrieval and matching, and the system needs to ensure the scientificity and rationality of this process.

[0089] Furthermore, the system should also consider dynamically updating the edge weights. Since the content and context of archival entries may change over time, the system can set up a dynamic weight adjustment mechanism to optimize the edge weights regularly or in real-time according to user behavior. This can ensure the timeliness and accuracy of the archival data index graph, enabling the knowledge graph to always reflect the latest archival information and user needs, and enhancing the retrieval experience of users in the cloud computing environment.

[0090] Based on the dynamic update mechanism, track the changes in the content, metadata of the archives and the relationships between them to ensure the real-time nature of the archive data index graph. Among them, when a new archive is added or an existing archive is modified, an incremental update strategy is used to only update the affected nodes and edges, and dynamically adjust the weights of the edges according to time and user behavior to gradually optimize the graph structure and query efficiency;

[0091] This step involves implementing a dynamic update mechanism in the archive data index graph to ensure that the graph structure always reflects the changes in the archive content, metadata and their relationships. When a new archive entry is added or an existing archive is modified, the system uses an incremental update strategy to only update the affected nodes and edges. This strategy optimizes the real-time nature of the graph, enabling the system to quickly respond to user requests with minimal time and resource consumption. The system dynamically adjusts the weights of the edges by tracking time and user behavior markers to maintain the accuracy and efficiency of the graph and keep it always suitable for the user's retrieval requests.

[0092] The dynamic update mechanism plays an important role in maintaining the real-time nature and effectiveness of the archive data index graph. By only updating the affected nodes and edges, the system can save resources, improve the overall performance, adopt user feedback and behavioral changes, thereby enhancing user satisfaction. In addition, the graph structure that reflects changes in a timely manner helps users quickly obtain the latest information, avoid retrieval errors caused by old data, and improve the credibility and usability of the system. This flexibility enables the archive retrieval system to adapt to the ever-changing archive data environment and strengthens its application value in the field of information management.

[0093] The third step focuses on ensuring the real-time nature of the archive data index graph. Over time, the content of the archives and their metadata may change, such as adding new archives, modifying the content of existing archives, or updating the classification information of archives, etc. Therefore, the system needs to have the ability to track these changes and adjust the index graph in a timely manner to ensure that it always reflects the latest archive status.

[0094] In the process of implementing dynamic updates, the system can adopt an incremental update strategy. This means that when a new archive is added or an existing archive is modified, the system only needs to update the nodes and edges directly related to these changes, rather than reconstructing the entire index graph. This method not only improves the update efficiency, saves computing resources, but also effectively reduces the impact on the user retrieval experience, thus achieving more efficient real-time data processing.

[0095] In addition, the system should also dynamically adjust the weights of the edges according to time and user behavior. For example, when a user frequently retrieves a specific type of file, the system can automatically increase the edge weights between these files and other related files, so as to better reflect the user's needs and preferences in subsequent retrievals. This adaptive mechanism can gradually optimize the structure of the graph and the query efficiency, ensuring that users can quickly and accurately obtain the required file information in the cloud computing environment and improving the overall user experience.

[0096] Select a distributed computing framework in the cloud computing environment and deploy the real-time file data index graph in the distributed graph database of the selected distributed computing framework.

[0097] In this step, the system deploys the real-time file data index graph in the distributed graph database of the selected distributed computing framework. The purpose of selecting the distributed computing framework is to utilize its powerful parallel processing ability and elastic scalability to support the efficient storage and processing of massive file data. By deploying the index graph in the distributed graph database, the system can ensure good performance and response speed when processing high-concurrency retrieval requests. In addition, this deployment method also provides a smooth path for future data expansion and system upgrade, enabling the system to still have sufficient flexibility and scalability in the face of increasing file data volume.

[0098] Deploying the real-time file data index graph in the distributed graph database not only improves the storage and computing efficiency, but also enhances the high availability and fault tolerance of the system. Through this deployment, the system can keep other nodes running normally when a single node fails, ensuring the continuity of the entire file retrieval service. In the face of large-scale user requests, the distributed architecture can evenly distribute the requests to each processing node through load balancing technology, thus significantly improving the system throughput and response speed. In addition, the design of the distributed database also makes data backup and recovery easier, enhancing the system security and data integrity, and ultimately improving the user's retrieval experience and satisfaction.

[0099] In the process of implementing the selection of the distributed computing framework and the deployment of the file data index graph, the first step of the system is to evaluate different distributed computing frameworks such as Apache Spark, Apache Flink or Hadoop, and select the most suitable framework according to specific requirements. The evaluation and selection criteria include data processing speed, scalability, fault tolerance and community support, etc. The system will evaluate the performance of each framework in processing large-scale graph data through benchmark tests, so as to ensure that the selected framework can efficiently support real-time data processing and query. After detailed evaluation, the system will select an optimal distributed computing framework and design the architecture according to the technical characteristics of this framework.

[0100] Next, the system will convert the real-time archive data index graph into the data structure required by the distributed database in a suitable format. In this process, the system first needs to store information such as embedded vectors and edge weights in a serialized manner to enable fast access in a distributed environment. At the same time, the system will determine the data distribution strategy of the index graph as hash partitioning or range partitioning to optimize the data read / write speed and storage efficiency. In this process, the system will ensure that each node can independently process part of the data while maintaining consistency globally. After this step is completed, the system will load the data into the distributed graph database and, based on this, implement the update and query of the real-time index graph.

[0101] Finally, the system will implement a series of API interfaces on the distributed graph database to support fast archive retrieval and graph query operations. These interfaces will be designed to be efficient and easy to call, allowing users to obtain relevant archive information through simple requests. At the same time, the system will also set up a monitoring mechanism to track data access situations and performance metrics in real time to ensure that the system can respond to users' retrieval needs in a timely manner. In this way, during the architecture design and implementation process, the system not only ensures performance and scalability but also lays a solid foundation for future function expansion and technological innovation. This enables the archive data index graph to be fully utilized in a distributed computing environment, significantly improving the overall performance of the system and the user experience.

[0102] S203, according to the user retrieval semantic vector, dynamically locate the search domain related to the retrieval semantics in the archive data index graph in the way of multi-layer projection calculation;

[0103] In this stage, we use the graph embedding algorithm to pre-train the archive data index graph to generate the vector representation of each node. By this method, the structure of the graph and the feature information of the nodes can be converted into low-dimensional dense vectors. These vectors not only retain the semantic relationships between nodes but also can reflect the attribute characteristics of nodes, making the subsequent similarity calculation more efficient and accurate. The generation of node embeddings is the basis of information retrieval, which allows the system to have richer semantic information during retrieval. By capturing the complex relationships between nodes, the system can effectively distinguish similar and dissimilar archive data, thus quickly locating the content related to the user's retrieval intention in a wide range of datasets and improving the accuracy and efficiency of retrieval.

[0104] Specifically, the graph embedding algorithm can be used to pre-train the graph embedding of the archive data index graph to generate the node embeddings of each node in the graph. Among them, deep feature representation is performed on each archive entry node so that the position of each node in the embedding space can reflect both its own attributes and the semantic or logical relationships with other nodes. The embedding space is the space composed of node embeddings.

[0105] At this stage, we use graph embedding algorithms to pre-train the archival data index graph to generate vector representations for each node. Through this method, the structure of the graph and the feature information of the nodes can be transformed into low-dimensional dense vectors. These vectors not only preserve the semantic relationships between nodes but also can reflect the attribute characteristics of nodes, making subsequent similarity calculations more efficient and accurate. The generation of node embeddings is the basis of information retrieval, which allows the system to have richer semantic information during retrieval. By capturing the complex relationships between nodes, the system can effectively distinguish between similar and dissimilar archival data, thereby quickly locating content related to the user's retrieval intent in a wide range of datasets and improving the accuracy and efficiency of retrieval.

[0106] First, perform data preprocessing on the archival data index graph, including cleaning invalid data, filling in missing values, and normalizing node features to ensure data consistency and high quality. This preprocessing step is crucial because it lays the foundation for subsequent graph embedding and ensures the reliability of the input data during model training.

[0107] Subsequently, select a suitable graph embedding algorithm, such as Node2Vec, GraphSAGE, or DeepWalk, etc. These algorithms capture graph structure information through methods such as random walks and node sampling and transform it into vector representations. Depending on the application scenario, select the algorithm that best meets the requirements to optimally extract node features and graph structure information.

[0108] Use the selected graph embedding algorithm to train the nodes, regularly check the model performance and perform parameter tuning, and update the node vectors through the backpropagation algorithm, so that the vector distances between similar nodes are reduced, while the vector distances between dissimilar nodes are increased. After training is completed, store the generated node embedding vectors in the database for subsequent use.

[0109] Use a deep alignment model to map the user retrieval semantic vector into the embedding space so that the mapped user retrieval semantic vector and the node embeddings in the archival data index graph are in the same space;

[0110] In this step, the use of the deep alignment model can transform the user's retrieval intent into a vector representation that matches the node embeddings in the archival data index graph. By designing a neural network structure, the system can effectively map the input user retrieval semantic vector into the embedding space, aligning the user's intent with the corresponding feature information in the dataset. Mapping the user retrieval semantic vector into the same space can accurately capture the user's true intent, thereby improving the relevance and accuracy of information retrieval. This process not only enhances the intelligence level of the system but also ensures that users can quickly locate the information they need during retrieval, improving the user experience.

[0111] The retrieval request of the user is parsed through natural language processing technology to extract the key information and semantic features therein. A pre-trained language model (such as BERT) can be used to convert the user input into a vector to form a preliminary user retrieval semantic vector. In this stage, the user intention is converted into a data-driven representation form, laying a foundation for the subsequent mapping.

[0112] Subsequently, a deep alignment model is constructed, which is usually composed of multiple layers of neural networks. Its input layer receives the user retrieval semantic vector, and the output layer generates the corresponding node embeddings. By designing a reasonable network structure, it can help the model learn the mapping relationship between the user intention and the node embeddings.

[0113] After the model design is completed, the deep alignment model is trained with training data to optimize the weights to minimize the distance between the user retrieval vector and the node embeddings. After training, the system can map the new user retrieval semantic vector into the embedding space in real time, ensuring that when the user conducts a retrieval each time, a semantic representation matching the archival data can be obtained quickly.

[0114] Calculate the similarity between the mapped user retrieval semantic vector and all node embeddings, and select the K nodes most similar to the user retrieval semantic vector from the archival data index graph as seed nodes to represent the starting point of the initial search;

[0115] In this step, the system calculates the similarity between the mapped user retrieval semantic vector and the embedding vectors of all nodes in the archival data index graph, and selects the K nodes that most conform to the user intention as seed nodes according to the calculation results. This process can effectively narrow the search scope and help the user quickly find the relevant archival data. Selecting the K nodes similar to the user retrieval intention as seed nodes ensures that the retrieval results of the system can accurately reflect the user's needs. This mechanism not only makes the retrieval process more efficient, but also provides a reliable basis for the subsequent expansion of archival information, improving the relevance of information retrieval and user satisfaction.

[0116] Select an appropriate similarity calculation scheme, such as cosine similarity or Euclidean distance, etc., to measure the similarity between the user's retrieval semantic vector and all node embedding vectors. Cosine similarity is widely used because it can effectively capture the directional similarity between vectors, especially suitable for processing scenarios of high-dimensional data. Using the calculated similarity values, the system stores the similarity between the user's retrieval vector and each node embedding in the similarity matrix. To ensure computational efficiency, parallel computing techniques can be adopted, using multi-core processors or distributed systems to accelerate this process to meet the needs of large-scale data. Finally, the system sorts the similarity matrix and selects the top K nodes with the highest similarity as seed nodes. This selection can not only be customized based on the user's retrieval intention but also combined with the user's historical behavior and preferences to improve the accuracy and practicality of the retrieval results.

[0117] According to the adjacent nodes of each seed node, expand the first-hop neighbor nodes outward, and judge whether the similarity between these neighbor nodes and the user's retrieval semantic vector reaches a specific threshold. If the similarity of the neighbor nodes does not reach the specific threshold, stop the expansion. If the similarity reaches the specific threshold, add it to the candidate node set. Moreover, for the new neighbor nodes, repeat the expansion process and continue to explore the next-hop neighbor nodes outward until the preset layer limit is reached or the condition of accumulating the number of relevant nodes is met.

[0118] The core of this step is to expand the adjacent nodes outward based on the initially selected seed nodes to form a larger search domain. By dynamically evaluating the similarity of the adjacent nodes, the system can effectively determine which nodes are relevant to the user's intention and expand layer by layer to obtain more information. By gradually expanding the adjacent nodes, the system can not only greatly enrich the retrieval results but also ensure the relevance of the results. Being able to continuously explore the nodes closely related to the user's retrieval intention improves the flexibility and intelligence level of the system, enabling the user to obtain a more comprehensive information retrieval experience.

[0119] First, determine the first-hop adjacent nodes of each seed node, and calculate the similarity between their embedding vectors and the user's retrieval semantic vector. This process can utilize the previously constructed similarity matrix to ensure rapid acquisition of similarity values. According to the set threshold, evaluate whether each adjacent node meets the expansion condition. For adjacent nodes that reach the similarity threshold, add them to the candidate node set; for nodes that do not meet the conditions, stop the expansion. For the newly added candidate nodes, repeat the above process, and use the same method to explore their adjacent nodes outward, continuously constructing the candidate set until the preset layer limit is reached or the condition of accumulating the number of relevant nodes is met. When the expansion process ends, determine the obtained candidate node set as the search domain. This search domain not only contains the detailed information of all relevant nodes but also has the graph structure information, ensuring that users can obtain diverse results closely related to the query intention.

[0120] Determine the set of candidate nodes obtained through multi-layer projection expansion as the search domain. The search domain is a subgraph of the archival data index graph and contains the graph structure information and the complete semantic embedding of the candidate nodes.

[0121] In the final step, the system summarizes all the candidate node sets obtained through the aforementioned expansion process into the search domain. This search domain can be regarded as a subgraph of the archival data index graph, with complete node information and the structural relationships between them, forming an information network that is more focused on the user's retrieval needs. By clearly setting the search domain, the system can improve the accuracy of information retrieval and ensure that the results returned to the user are the most relevant. This subgraph not only provides the basis for subsequent matching and ranking but also ensures that users can find specific and valuable information in a large amount of data, significantly enhancing the user's retrieval experience.

[0122] Organize the candidate node set obtained through expansion, determine it as the search domain, and mark it as a subgraph in the archival data index graph. This process requires integrating each node and its corresponding edge weight information together to form an independent structured representation for subsequent data processing. Further, integrate the graph structure information within the search domain, including the edges and their weights between each node, to ensure that this search domain is not just a set of nodes but can also retain the semantic relationships and logical coherence between the nodes. This integration provides the necessary context information for subsequent candidate archival screening and ranking. Finally, output the finally constructed search domain in a data format available to the system for subsequent multi-modal fusion and matching algorithms to process. At the same time, the creation of this search domain makes the entire retrieval process more flexible, capable of quickly adjusting the search strategy according to the user's query intention, and improving the acquisition efficiency and accuracy of the retrieval results.

[0123] S204. Within the search domain, use a matching algorithm for multi-modal semantic fusion to screen and sort the candidate file sets that meet the query intent and display them. Among them, the screening process comprehensively considers the semantic relevance of the files, the characteristics of the file storage locations, and the user behavior context.

[0124] In this stage, the system uses a matching algorithm for multi-modal semantic fusion to screen and sort the candidate files according to the retrieval intent of the user within the search domain. The multi-modal semantic fusion technology ensures that it can effectively capture the multi-dimensional information of the file entries by integrating features from different sources (such as text, images, and metadata). In this process, the system not only focuses on the semantic relevance of the file content but also combines the location characteristics of the files in the storage system to ensure the practicality and accessibility of the query results. At the same time, it incorporates the context information of the user's historical behavior to achieve a personalized retrieval experience.

[0125] Through the matching algorithm for multi-modal semantic fusion, the system can improve the accuracy and relevance of the retrieval results. This process not only ensures the consistency between the file entries and the user's query intent but also optimizes the information display by combining the storage location characteristics and the user behavior context, ensuring that the user can quickly find the files that best meet their needs. This comprehensive consideration greatly improves the user experience and makes the retrieval ability of the system more intelligent and personalized.

[0126] Specifically, for each file entry that may be related to the retrieval semantics within the search domain, extract multi-modal feature vectors.

[0127] In this step, the system comprehensively extracts the multi-modal feature vectors for each candidate file entry. These feature vectors not only include text features but also cover other forms of information such as images, audio, and video. Using natural language processing technology, the system first performs semantic analysis on the text information to extract keywords and topic features. At the same time, for files containing images or videos, computer vision technology is applied to extract visual features to ensure that each file entry is comprehensively represented to reflect its multi-dimensional information characteristics. By extracting multi-modal feature vectors, the system can comprehensively capture the information of the file entries, thereby providing a richer semantic background. This diverse data representation enhances the system's ability to understand the user's query, making the subsequent matching and screening processes more accurate and better meeting the user's retrieval needs in different scenarios.

[0128] When processing text information, the system uses natural language processing (NLP) technology for semantic analysis, uses text embedding technology (such as BERT or Word2Vec) to vectorize the archive content, and extracts key information and thematic features. At the same time, the most representative words in the archive are identified using a keyword extraction algorithm to construct a text feature vector. This process ensures that the text information is highly concentrated and effectively expressed. For the image and video content contained in the archive, the system uses deep learning technologies such as convolutional neural networks (CNN) and temporal convolutional networks (TCN) to extract the corresponding features. By analyzing the color distribution and edge features in the image and the motion pattern in the video, a visual feature vector that meets the requirements of multimodal features is generated. This process supports the system to better understand the content characteristics when processing multimedia information. After the extraction is completed, the system integrates various feature vectors such as text, images and videos to form a multimodal feature representation. This integration can be done by splicing or weighted averaging to ensure that different types of information can be effectively integrated to form the final multimodal feature vector, providing a comprehensive foundation for subsequent matching algorithms.

[0129] Using a multimodal fusion algorithm to fuse feature vectors of different modalities, wherein a weighted summation, an attention mechanism or a deep neural network is used to achieve feature fusion, and a fused feature vector of each archive entry is generated to represent the multimodal semantic information of each archive entry;

[0130] In this step, the system effectively fuses the feature vectors extracted from each modal information to form a unified feature representation. By applying technologies such as weighted summation, attention mechanism or deep neural network, the system can dynamically adjust the contribution of each modal information, thereby greatly improving the expressive power of the fused feature vector. Feature fusion not only considers the importance of different modalities, but also enhances the understanding of the semantic relationship within the archive. Through multimodal fusion, the system can generate richer and more comprehensive archive feature vectors, improving the intelligence and accuracy of matching. This process lays the foundation for subsequent similarity calculations and candidate archive screening, enabling the system to respond more accurately to complex queries and optimize the user's retrieval experience.

[0131] The extracted feature vectors can be combined using a weighted summation method, and appropriate weights are designed to reflect the importance of different modality features. This weight can be dynamically adjusted based on prior knowledge, user feedback, or the learning results during model training, so as to ensure that the fused feature vectors can accurately reflect the multi-modal nature of the archive entries. The attention mechanism can also be introduced to weight the features of different modalities. In this process, the system dynamically adjusts the weights according to the importance of the feature vectors, so that the features highly relevant to the user's retrieval intention are enhanced. Using deep learning models, the system can automatically learn the relationships between the features of each modality and generate more accurate fused features. To further enhance the effect of feature fusion, the system can also achieve this process by constructing a deep neural network (DNN). The input of the network is the feature vectors of each modality, and through multiple layers of non-linear transformations, a unified fused feature is output. By optimizing the model parameters using the backpropagation algorithm, this method ensures that the fused feature vectors have the representation ability in the high-dimensional space and comprehensively cover the multi-modal information of the archive entries.

[0132] Calculate the similarity between the fused feature vector of each archive entry and the user's retrieval semantic vector, and filter out the archive entries with a similarity higher than the preset threshold as the candidate archive set, where the candidate archive set is the archive entries most relevant to the user's query intention;

[0133] In this step, the system calculates the similarity between the fused feature vector of each archive entry and the user's retrieval semantic vector to identify the archive set that meets the user's query intention. By setting a preset similarity threshold, the system can effectively distinguish the archive entries closely related to the user's needs from the irrelevant ones, ensuring that the selected candidate set meets certain quality standards. The implementation of this step ensures the accuracy and relevance of the retrieval results, so that the archive entries finally presented to the user can truly reflect their query needs. This process effectively avoids the interference of irrelevant information, improves the efficiency and convenience for the user to find the required information, and enhances the overall user experience.

[0134] When calculating the similarity, the system can choose various similarity measurement methods, such as cosine similarity, Euclidean distance, etc. Cosine similarity is widely used because it can effectively capture the directional similarity of vectors. The system can select according to the specific application scenario and the nature of the feature vectors to ensure the accuracy and efficiency of the calculation.

[0135] After the calculation is completed, the system stores the similarity between the fused feature vector of each archive entry and the user's retrieval semantic vector in the similarity matrix for subsequent screening and sorting. To improve the calculation efficiency, the system can parallelize the similarity calculation and quickly complete the matrix construction through a multi-threaded or distributed computing framework to meet the requirements of large-scale data processing.

[0136] Finally, through the set threshold, the system marks the archive entries in the similarity matrix that exceed the threshold as candidate archives. This process ensures that the candidate archive set only contains entries that are most relevant to the user's query intent, and prepares a high-quality data basis for subsequent sorting and display, enabling the user to quickly obtain the required information.

[0137] Sort the candidate archive set according to similarity, user preferences, archive importance, and freshness, and display the sorted archive entries to the user.

[0138] In this step, the system comprehensively sorts the candidate archive set based on multiple factors for final display to the user. The sorting basis includes similarity (i.e., the degree of match between the archive and the user's retrieval semantics), user preferences (based on the user's historical behavior and preference model), the importance of the archive (such as document type, creation time, etc.), and the freshness of the archive (i.e., the update situation of the content). By weighting these metrics, the system can generate a comprehensive score, thus achieving a more intelligent sorting. A reasonable sorting mechanism can significantly improve the user experience in information retrieval, making the finally displayed archive entries meet the user's real needs. By comprehensively considering multiple dimensions such as similarity and user preferences, the system not only improves the relevance of the retrieval results but also enhances the user's trust and satisfaction with the retrieval system, thereby improving the overall usage experience.

[0139] The system first needs to define the scoring criteria for different sorting factors. For example, for similarity, it can be directly reflected by the similarity value; user preferences can be calculated by analyzing the user's historical click-through rate and preference characteristics; the importance of the archive may be obtained through an internal scoring system (such as expert review), and freshness can be quantified based on the creation and modification dates of the document. These scoring criteria can help the system establish a comprehensive evaluation model for each archive.

[0140] Next, the system assigns weights to each scoring factor to reflect its relative importance in the sorting process. Through a linear weighting model, weighted average, or other more complex machine learning models, the system comprehensively calculates the scores of each archive entry to generate a comprehensive score. This process needs to be optimized through experiments according to the actual situation to ensure that the final sorting can better meet the user's needs.

[0141] Finally, the system sorts the candidate archive set according to the comprehensive score. After the sorting is completed, the top-ranked archive entries are displayed to the user to ensure that the user can quickly find the information that best matches their query intent. In addition, the system can attach relevant metadata (such as creation time, document type, etc.) during the display to help the user better judge the practicality and relevance of the archive. This display form helps to improve the user's satisfaction and operation efficiency during the retrieval process.

[0142] It can be seen that by receiving the retrieval request information input by the user, using the context-aware semantic parsing model, parsing the retrieval intention according to the retrieval content, the user's historical retrieval behavior and the retrieval domain database in the retrieval request information, and generating a user retrieval semantic vector in combination with the retrieval intention; for the archival data distributedly stored in the cloud storage environment, constructing an archival data index graph; according to the user retrieval semantic vector, dynamically locating the search domain related to the retrieval semantics in the archival data index graph in the manner of multi-layer projection calculation; within the search domain, using the matching algorithm of multi-modal semantic fusion to screen and sort and display the candidate archival data sets that meet the query intention, thereby improving the accuracy, efficiency and user experience of archival retrieval.

[0143] Another embodiment of the present invention provides an archival retrieval system based on cloud computing. Refer to Figure 3 , the system may include:

[0144] The first acquisition module 301 is used to acquire the initial archival data, and establish an archival data index network including the archival data according to the initial archival data, wherein the archival data index network includes archival data index nodes;

[0145] The receiving module 301 is used to receive the retrieval request information input by the user, use the context-aware semantic parsing model, parse the retrieval intention according to the retrieval content, the user's historical retrieval behavior and the retrieval domain database in the retrieval request information, and generate a user retrieval semantic vector in combination with the retrieval intention;

[0146] The construction module 302 is used to construct an archival data index graph for the archival data distributedly stored in the cloud storage environment, wherein the archival data index graph represents archival entries by nodes, and represents the semantic or logical relationship between archives by edges, and based on the dynamic update mechanism, the graph structure and index efficiency are optimized in real time, and the archival data index graph is deployed in the cloud distributed computing framework;

[0147] The positioning module 303 is used to dynamically locate the search domain related to the retrieval semantics in the archival data index graph in the manner of multi-layer projection calculation according to the user retrieval semantic vector;

[0148] The screening module 304 is used to screen, sort and display the candidate archival data sets that meet the query intention within the search domain by using the matching algorithm of multi-modal semantic fusion, wherein the screening process comprehensively considers the archival semantic relevance, the archival storage location characteristics and the user behavior context.

[0149] It can be seen that by receiving the retrieval request information input by the user, using the context-aware semantic parsing model, parsing the retrieval intention according to the retrieval content, the user's historical retrieval behavior and the retrieval domain database in the retrieval request information, and generating a user retrieval semantic vector in combination with the retrieval intention; for the archival data distributedly stored in the cloud storage environment, constructing an archival data index graph; according to the user retrieval semantic vector, dynamically locating the search domain related to the retrieval semantics in the archival data index graph in the way of multi-layer projection calculation; within the search domain, using a multi-modal semantic fusion matching algorithm to screen and sort and display the candidate archival data sets that meet the query intention, so as to improve the accuracy, efficiency and user experience of archival retrieval.

[0150] An embodiment of the present invention further provides a storage medium, in which a computer program is stored, and wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0151] Specifically, in this embodiment, the above storage medium may be configured to store a computer program for executing the following steps:

[0152] S201, receiving the retrieval request information input by the user, using the context-aware semantic parsing model, parsing the retrieval intention according to the retrieval content, the user's historical retrieval behavior and the retrieval domain database in the retrieval request information, and generating a user retrieval semantic vector in combination with the retrieval intention;

[0153] S202, for the archival data distributedly stored in the cloud storage environment, constructing an archival data index graph, wherein the archival data index graph represents archival entries through nodes, and represents the semantic or logical relationship between archives, and based on a dynamic update mechanism, the graph structure and index efficiency are optimized in real time, and the archival data index graph is deployed in a cloud distributed computing framework;

[0154] S203, according to the user retrieval semantic vector, dynamically locating the search domain related to the retrieval semantics in the archival data index graph in the way of multi-layer projection calculation;

[0155] S204, within the search domain, using a multi-modal semantic fusion matching algorithm to screen and sort and display the candidate archival data sets that meet the query intention, wherein the screening process comprehensively considers the archival semantic relevance, the archival storage location characteristics and the user behavior context.

[0156] It can be seen that by receiving the retrieval request information input by the user, using the context-aware semantic parsing model, parsing the retrieval intention according to the retrieval content, the user's historical retrieval behavior and the retrieval domain database in the retrieval request information, and generating a user retrieval semantic vector in combination with the retrieval intention; for the archival data distributedly stored in the cloud storage environment, constructing an archival data index graph; according to the user retrieval semantic vector, dynamically locating the search domain related to the retrieval semantics in the archival data index graph in the way of multi-layer projection calculation; within the search domain, using the matching algorithm of multi-modal semantic fusion to screen and sort the candidate archival sets that meet the query intention, so as to improve the accuracy, efficiency and user experience of archival retrieval.

[0157] An embodiment of the present invention also provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0158] Specifically, the above electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0159] Specifically, in this embodiment, the above processor may be configured to execute the following steps through a computer program:

[0160] S201, Receive the retrieval request information input by the user, use the context-aware semantic parsing model, parse the retrieval intention according to the retrieval content, the user's historical retrieval behavior and the retrieval domain database in the retrieval request information, and generate a user retrieval semantic vector in combination with the retrieval intention;

[0161] S202, For the archival data distributedly stored in the cloud storage environment, construct an archival data index graph, wherein the archival data index graph represents archival entries through nodes, and represents the semantic or logical relationship between archives, and based on the dynamic update mechanism, the graph structure and index efficiency are optimized in real time, and the archival data index graph is deployed in the cloud distributed computing framework;

[0162] S203, According to the user retrieval semantic vector, dynamically locate the search domain related to the retrieval semantics in the archival data index graph in the way of multi-layer projection calculation;

[0163] S204, Within the search domain, use the matching algorithm of multi-modal semantic fusion to screen and sort the candidate archival sets that meet the query intention, wherein the screening process comprehensively considers the archival semantic relevance, the archival storage location characteristics and the user behavior context.

[0164] It can be seen that by receiving the retrieval request information input by the user, using the context-aware semantic parsing model, parsing the retrieval intention according to the retrieval content, the user's historical retrieval behavior and the retrieval domain database in the retrieval request information, and generating a user retrieval semantic vector in combination with the retrieval intention; constructing an archive data index graph for the archive data distributedly stored in the cloud storage environment; dynamically positioning the search domain related to the retrieval semantics in the archive data index graph in a multi-layer projection calculation manner according to the user retrieval semantic vector; and screening and sorting and displaying the candidate archive set that meets the query intention by using the matching algorithm of multi-modal semantic fusion within the search domain, so as to improve the accuracy, efficiency and user experience of archive retrieval.

[0165] The structure, features and function effects of the present invention have been described in detail based on the embodiments shown in the drawings. The above are only the preferred embodiments of the present invention, but the present invention is not limited to the scope defined by the drawings. Any changes made according to the concept of the present invention, or equivalent embodiments modified into equivalent changes, should still be within the protection scope of the present invention when they do not exceed the spirit covered by the description and the drawings.

Claims

1. A cloud computing-based archive retrieval method, characterized in that: The method comprises: Receiving search request information input by a user, using a context-aware semantic parsing model to parse the search intent according to the search content in the search request information, the user's historical search behavior, and the search domain database, and generating a user search semantic vector in combination with the search intent; For the archive data distributedly stored in the cloud storage environment, an archive data index graph is constructed, wherein the archive data index graph represents archive entries through nodes, and edges represent semantic or logical relationships between archives, and the graph structure and index efficiency are optimized in real time based on a dynamic update mechanism, and the archive data index graph is deployed in a cloud distributed computing framework; According to the user retrieval semantic vector, the search domain related to the retrieval semantics is dynamically located in the archive data index graph by means of multi-layer projection calculation; wherein, the archive data index graph is pre-trained for graph embedding using a graph embedding algorithm to generate a node embedding of each node in the graph, wherein each archive entry node is represented by a deep feature so that the position of each node in the embedding space can simultaneously reflect its own attributes and the semantic or logical relationship with other nodes, and the embedding space is a space composed of node embeddings; the user retrieval semantic vector is mapped to the embedding space using a deep alignment model so that the mapped user retrieval semantic vector and the node embedding in the archive data index graph are in the same space; Calculate the similarity between the mapped user retrieval semantic vector and all node embeddings, and select the K nodes most similar to the user retrieval semantic vector from the archive data index graph as seed nodes to indicate the starting point of the initial search; According to the adjacent nodes of each seed node, the first-hop neighbor nodes are expanded outwards to determine whether the similarity between these neighbor nodes and the user's search semantic vector reaches a specific threshold. If the similarity of the neighbor nodes does not reach the specific threshold, the expansion is stopped. If the similarity reaches the specific threshold, it is added to the candidate node set, and for the new neighbor nodes, the expansion process is repeated to continue to explore the next-hop neighbor nodes outwards until the preset layer limit is reached or the condition of the cumulative number of related nodes is met; all candidate node sets obtained by multi-layer projection expansion are determined as the search domain, and the search domain is a subgraph of the archive data index graph, and contains graph structure information and complete semantic embedding of candidate nodes; In the search domain, a matching algorithm of multimodal semantic fusion is used to screen candidate archive sets that meet the search intent and sort and display them, wherein the screening process comprehensively considers archive semantic relevance, archive storage location characteristics and user behavior context.

2. The method according to claim 1, characterized in that The receiving of the search request information input by the user, using the context-aware semantic parsing model, parsing the search intent according to the search content in the search request information, the user's historical search behavior and the search domain database, and generating the user search semantic vector in combination with the search intent, including: Map the search request information entered by the user to a fixed-dimensional semantic space through the BERT or Transformer model to generate a semantic embedding vector; Obtain the user's historical search records, analyze historical search keywords, document types, and user click behaviors, build user portraits, and convert user portrait data into user preference embedding vectors; Accessing the retrieval domain database to determine the domain knowledge related to the retrieval request information to extract the domain feature embedding vector; An algorithm based on the attention mechanism is used to fuse the semantic embedding vector, user preference embedding vector and domain feature embedding vector to perform retrieval intent analysis, assign a weight to each analyzed intent, and generate a user retrieval semantic vector based on the weight. The user retrieval semantic vector covers the current retrieval intent and reflects the user's individual needs and historical behavior characteristics.

3. The method according to claim 2, characterized in that The archival data index graph is constructed for the archival data distributedly stored in the cloud storage environment, wherein the archival data index graph represents archival items through nodes, and edges represent semantic or logical relationships between archives, and optimizes the graph structure and index efficiency in real time based on a dynamic update mechanism, and deploys the archival data index graph in a cloud distributed computing framework, including: The archival data is processed using natural language processing technology to extract the semantic features of archival items, each archival item is used as a node in the archival data index graph structure, and the corresponding archival item semantic features are assigned to each node; Creating edges in the archival data index graph structure according to the semantic or logical relationship between archival items, and setting the initial weight of each edge to obtain the archival data index graph, wherein the weight is initialized by comparing the similarity of two archival items to reflect the strength of the edge; Based on the dynamic update mechanism, the changes of archive content, metadata and their relationships are tracked to ensure the real-time performance of the archive data index graph. When adding new archives or modifying existing archives, an incremental update strategy is used to update only the affected nodes and edges, and the weights of the edges are dynamically adjusted according to time and user behavior to gradually optimize the graph structure and query efficiency. A distributed computing framework is selected in a cloud computing environment, and a real-time archive data index graph is deployed in a distributed graph database of the selected distributed computing framework.

4. The method according to claim 3, characterized in that In the search domain, a matching algorithm of multimodal semantic fusion is used to screen candidate archive sets that meet the search intent and sort and display them, including: For each archive entry contained in the search domain that may be related to the search semantics, extracting a multimodal feature vector; Using a multimodal fusion algorithm to fuse feature vectors of different modalities, wherein a weighted summation, an attention mechanism or a deep neural network is used to achieve feature fusion, and a fused feature vector of each archive entry is generated to represent the multimodal semantic information of each archive entry; Calculate the similarity between the fusion feature vector of each archive item and the user search semantic vector, and select archive items with similarity higher than a preset threshold as a candidate archive set, wherein the candidate archive set is the archive items most relevant to the user's search intention; The candidate profile collection is sorted according to similarity, user preference, profile importance and freshness, and the sorted profile entries are displayed to the user.

5. A cloud computing-based archive retrieval system, characterized in that: The system comprises: A receiving module, used to receive search request information input by a user, analyze the search intent according to the search content in the search request information, the user's historical search behavior and the search domain database using a context-aware semantic analysis model, and generate a user search semantic vector in combination with the search intent; A construction module is used to construct an archive data index graph for archive data distributedly stored in a cloud storage environment, wherein the archive data index graph represents archive entries through nodes, and edges represent semantic or logical relationships between archives, and optimizes the graph structure and index efficiency in real time based on a dynamic update mechanism, and deploys the archive data index graph in a cloud distributed computing framework; A positioning module is used to dynamically locate the search domain related to the retrieval semantics in the archive data index graph in a multi-layer projection calculation manner according to the user retrieval semantic vector; wherein, the archive data index graph is pre-trained for graph embedding using a graph embedding algorithm to generate a node embedding of each node in the graph, wherein each archive entry node is represented by a deep feature so that the position of each node in the embedding space can simultaneously reflect its own attributes and the semantic or logical relationship with other nodes, and the embedding space is a space composed of node embeddings; and the user retrieval semantic vector is mapped to the embedding space using a deep alignment model so that the mapped user retrieval semantic vector and the node embedding in the archive data index graph are in the same space; Calculate the similarity between the mapped user retrieval semantic vector and all node embeddings, and select the K nodes most similar to the user retrieval semantic vector from the archive data index graph as seed nodes to indicate the starting point of the initial search; According to the adjacent nodes of each seed node, the first-hop neighbor nodes are expanded outwards to determine whether the similarity between these neighbor nodes and the user's search semantic vector reaches a specific threshold. If the similarity of the neighbor nodes does not reach the specific threshold, the expansion is stopped. If the similarity reaches the specific threshold, it is added to the candidate node set, and for the new neighbor nodes, the expansion process is repeated to continue to explore the next-hop neighbor nodes outwards until the preset layer limit is reached or the condition of the cumulative number of related nodes is met; all candidate node sets obtained by multi-layer projection expansion are determined as the search domain, and the search domain is a subgraph of the archive data index graph, and contains graph structure information and complete semantic embedding of candidate nodes; The screening module is used to screen the candidate archive sets that meet the retrieval intention within the search domain by using the matching algorithm of multimodal semantic fusion and sort and display them, wherein the screening process comprehensively considers the archive semantic relevance, archive storage location characteristics and user behavior context.

6. The system according to claim 5, characterized in that The receiving module is specifically used for: Map the search request information entered by the user to a fixed-dimensional semantic space through the BERT or Transformer model to generate a semantic embedding vector; Obtain the user's historical search records, analyze historical search keywords, document types, and user click behaviors, build user portraits, and convert user portrait data into user preference embedding vectors; Accessing the retrieval domain database to determine the domain knowledge related to the retrieval request information to extract the domain feature embedding vector; An algorithm based on the attention mechanism is used to fuse the semantic embedding vector, user preference embedding vector and domain feature embedding vector to perform retrieval intent analysis, assign a weight to each analyzed intent, and generate a user retrieval semantic vector based on the weight. The user retrieval semantic vector covers the current retrieval intent and reflects the user's individual needs and historical behavior characteristics.

7. The system according to claim 6, characterized in that The building blocks are specifically used for: The archival data is processed using natural language processing technology to extract the semantic features of archival items, each archival item is used as a node in the archival data index graph structure, and the corresponding archival item semantic features are assigned to each node; Creating edges in the archival data index graph structure according to the semantic or logical relationship between archival items, and setting the initial weight of each edge to obtain the archival data index graph, wherein the weight is initialized by comparing the similarity of two archival items to reflect the strength of the edge; Based on the dynamic update mechanism, the changes of archive content, metadata and their relationships are tracked to ensure the real-time performance of the archive data index graph. When adding new archives or modifying existing archives, an incremental update strategy is used to update only the affected nodes and edges, and the weights of the edges are dynamically adjusted according to time and user behavior to gradually optimize the graph structure and query efficiency. A distributed computing framework is selected in a cloud computing environment, and a real-time archive data index graph is deployed in a distributed graph database of the selected distributed computing framework.

8. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 4 when executed.

9. An electronic device, comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Financial product pushing method and device

    CN111292171A

  • Artificial intelligence data search and distribution method and system

    CN118467851A