An intelligent file classification and retrieval method and system
Through multimodal data fusion and deep learning technology, semantic correlation maps are generated, which solves the accuracy and relevance of classification and retrieval in traditional archive management, and realizes the intelligence and modernization of archive management.
Patent Information
- Application Number
- CN202510578193.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-07
AI Technical Summary
In the existing archive management system, the traditional classification method relies on manual settings and is highly subjective, making it difficult to adapt to diversity and dynamics, resulting in low accuracy and correlation of search results. The existing search methods are difficult to effectively match user intentions, reducing the user experience and mining of archive value.
A multimodal data fusion algorithm is used to convert text, image and audio data in a unified format and extract feature. A deep learning adaptive classification model is used to generate multi-level classification labels, and a semantic correlation map is constructed through an enhanced semantic network, and optimized it. Semantic matching and correlation sorting are performed based on the search requests input by users.
It improves the relevance and accuracy of archive retrieval, promotes the modernization and intelligence of archive management, reduces the burden of manual classification, and improves the efficiency and accuracy of information acquisition.
Smart Images

Figure CN120086390B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of file management, and particularly relates to a method and system for intelligent classification and retrieval of files. Background Art
[0002] In the information age, the importance of file management has become increasingly prominent. With the popularization of electronic files, various file information is stored and managed in the form of multi-modal data such as text, images, and audio. The traditional methods of file classification and retrieval have gradually revealed their limitations. Most of the existing classification methods rely on manual annotation and simple keyword matching, which is not only inefficient but also difficult to adapt to the diversity and dynamics of file data, resulting in low accuracy and relevance of retrieval results.
[0003] First of all, the traditional file classification method often relies on manually set classification rules. This method is highly subjective and easily affected by personal experience, making it difficult to ensure consistency and comprehensiveness. At the same time, problems such as information omission and redundancy in the classification process are widespread. Especially when faced with large and complex file data, manual processing will greatly increase the workload. Secondly, most of the existing retrieval methods are based on keyword search. The retrieval requests input by users often fail to effectively match the file content. The retrieval intentions, semantics, and context of users are often ignored, often resulting in the inability to retrieve relevant file information in a timely and effective manner. This method not only reduces the user's retrieval experience but also fails to fully explore the potential value of files. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for intelligent classification and retrieval of files to solve the deficiencies in the prior art, improve the relevance and accuracy of retrieval, and promote the modernization and intelligence of file management.
[0005] An embodiment of the present application provides a method for intelligent classification and retrieval of files, and the method includes:
[0006] According to the original data format of the file, using a multi-modal data fusion algorithm, perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized file feature data set;
[0007] Based on a deep learning-based adaptive classification model, perform classification processing on the standardized file feature data set to generate multi-level file classification labels;
[0008] Through enhanced semantic network construction technology, analyze the semantic relationships between file classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph;
[0009] According to the retrieval request input by the user, semantic matching and relevance ranking are performed using the optimized semantic association graph, and the file information most relevant to the retrieval request is output.
[0010] Optionally, according to the original data format of the file, a multi-modal data fusion algorithm is used to perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized file feature data set, including:
[0011] Perform semantic parsing on the text data through the BERT model to extract text feature vectors, perform feature extraction on the image data through the convolutional neural network CNN to generate image feature vectors, and perform spectral analysis on the audio data through the short-time Fourier transform STFT and adaptive filters to extract audio feature vectors;
[0012] Normalize the obtained text feature vectors, image feature vectors, and audio feature vectors, and use the multi-modal deep belief network DBN for feature fusion. Among them, input the normalized text feature vectors, image feature vectors, and audio feature vectors into the multi-layer DBN, and through layer-by-layer pre-training and backpropagation fine-tuning, learn the deep associations between the features of each modality, and use the attention mechanism to weight the fused multi-modal feature vectors to highlight key features and suppress redundant information;
[0013] Perform unified format conversion on the fused multi-modal feature vectors, use high-dimensional feature mapping technology to map them into a unified feature space, and use the kernel function to enhance the non-linear expression ability of the feature space to ensure the effective representation of different modality features in the unified space, and integrate the feature vectors in unified format into the standardized file feature data set.
[0014] Optionally, the adaptive classification model based on deep learning is used to classify the standardized file feature data set to generate multi-level file classification labels, including:
[0015] Construct a hybrid model based on the convolutional neural network CNN and the long short-term memory network LSTM. CNN is used to extract the spatial information of features, and LSTM is used to process the time series dynamic changes of features. Among them, when constructing the hybrid model, design multiple parallel CNN branches, each branch uses different convolutional kernel sizes and strides to extract features from different scales. At the same time, to prevent feature interference, use the attention mechanism to assign different weights to each branch;
[0016] At the output layer of the model, a hierarchical Softmax mechanism is adopted to construct multi-level classification labels. First, the classification results of the first layer are output. Then, according to the classification results of the first layer, subclass labels for the second-layer classification are correspondingly generated. Among them, for each category, by introducing label relationship modeling based on GCN, the dependency relationship between each label is recorded, and the inference ability of the model is enhanced by using the graph structure information of the labels;
[0017] The constructed hybrid model is used as an adaptive classification model to infer the standardized archive feature dataset and output multi-level archive classification labels.
[0018] Optionally, through the enhanced semantic network construction technology, analyze the semantic relationship between archive classification labels, generate a semantic association graph, and use a graph neural network to optimize the semantic association graph, including:
[0019] Conduct co-occurrence analysis of the archive classification labels, and construct a co-occurrence matrix by calculating the co-occurrence frequency of different classification labels in the same archive;
[0020] Based on the co-occurrence matrix, use the non-negative matrix factorization (NMF) technique to decompose the co-occurrence matrix into two matrices W and H. Among them, W represents the latent semantic features of the labels, and H represents the expression of the labels on these features, so as to map the high-dimensional co-occurrence features to a low-dimensional semantic feature space and extract the latent semantic relationship of the labels;
[0021] According to the latent semantic feature matrix W, use the broadcast mechanism to construct the semantic relationship between the labels to obtain a semantic association graph, where nodes are created for each label, and the edge weights are defined by the similarity between the labels;
[0022] Adopt graph embedding technology to generate the embedding vector of each node based on the context within the semantic association graph, and use the embedding vector to update the edge weights of the semantic association graph to optimize the semantic association graph, so that labels with similar semantic relationships have higher connectivity in the graph.
[0023] Optionally, according to the retrieval request input by the user, use the optimized semantic association graph for semantic matching and relevance ranking, and output the archive information most relevant to the retrieval request, including:
[0024] For the retrieval request input by the user, first perform multimodal parsing on the retrieval content, and convert the text query, voice input, or image content in the retrieval content into embedding features of different modalities. Among them, if it is a text input query, use a pre-trained language model to parse the retrieval text and extract semantic embedding features to ensure that the extracted semantic embedding features can capture the user's intention; if it is a voice input query, convert the voice content into text through speech recognition and further use the language model to extract semantic embedding features; if it is an image input, use a convolutional neural network to extract image embedding features;
[0025] Fuse the embedding features of different modalities into a unified retrieval request embedding feature through a multimodal fusion mechanism to map it into a feature space consistent with the archive feature dataset;
[0026] Query the node closest to the retrieval request in the optimized semantic association graph, use graph embedding technology to generate semantic embeddings of each node, and calculate the similarity score between the retrieval request embedding feature and the node embedding feature in the graph;
[0027] According to the similarity score, select several initial nodes and their direct adjacent nodes most relevant to the retrieval request as preliminary candidate nodes to obtain a corresponding preliminary candidate archive set;
[0028] Use a graph neural network to perform message propagation on the preliminary candidate nodes, and propagate the semantic information of the adjacent nodes into the retrieval matching nodes to enrich the context representation of the nodes;
[0029] Optimize the semantic embedding of each node through multiple rounds of propagation iteration to better represent the semantic features in the current retrieval context;
[0030] Calculate the similarity score between the propagated node embedding and the retrieval request embedding again, sort the preliminary candidate archive set, and sequentially output the candidate archive information most relevant to the retrieval request.
[0031] Another embodiment of the present application provides an archive intelligent classification and retrieval system, and the system includes:
[0032] A conversion module for performing unified format conversion and feature extraction on text, image, and audio data according to the original data format of the archive by using a multimodal data fusion algorithm to obtain a standardized archive feature dataset;
[0033] A classification module for performing classification processing on the standardized archive feature dataset based on an adaptive classification model of deep learning to generate multi-level archive classification labels;
[0034] A generation module, configured to analyze the semantic relationships between archive classification tags through enhanced semantic network construction technology, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph.
[0035] An output module, configured to perform semantic matching and relevance ranking using the optimized semantic association graph according to a retrieval request input by a user, and output the archive information most relevant to the retrieval request.
[0036] Another embodiment of the present application provides a storage medium in which a computer program is stored, wherein the computer program is configured to execute the method described in any one of the above when running.
[0037] Another embodiment of the present application provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any one of the above.
[0038] Compared with the prior art, an intelligent archive classification and retrieval method provided by the present invention, according to the original data format of archives, uses a multi-modal data fusion algorithm to perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized archive feature data set; based on an adaptive classification model of deep learning, classifies the standardized archive feature data set to generate multi-level archive classification tags; analyzes the semantic relationships between archive classification tags through enhanced semantic network construction technology, generates a semantic association graph, and optimizes the semantic association graph to obtain an optimized semantic association graph; according to a retrieval request input by a user, performs semantic matching and relevance ranking using the optimized semantic association graph, and outputs the archive information most relevant to the retrieval request, thereby being able to improve the relevance and accuracy of retrieval and promote the modernization and intelligence of archive management. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a hardware structure block diagram of a computer terminal for an intelligent archive classification and retrieval method provided by an embodiment of the present invention;
[0040] Figure 2 It is a flowchart of an intelligent archive classification and retrieval method provided by an embodiment of the present invention;
[0041] Figure 3 It is a structural diagram of an intelligent archive classification and retrieval system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.
[0043] An intelligent retrieval system for electronic archives constructed based on natural language processing (NLP) technology and machine learning algorithms is one of the solutions born to solve the above problems. Such systems use advanced NLP technology to parse text content and extract key information points; at the same time, powerful machine learning models are used to classify and cluster this information, enabling the computer to "understand" the meaning behind the documents like a human. In this way, when a user enters a query request, the system can quickly identify the most relevant results based on the pre-trained model and return them to the user, greatly improving the speed and quality of the search.
[0044] In addition, with the development of emerging technologies such as deep learning, future intelligent retrieval of electronic archives will be even more intelligent. For example, it can continuously learn and optimize its own knowledge base to adapt to changes in specific terms in different fields; or integrate cross-media information by combining functions such as image recognition to provide users with a more comprehensive and rich service experience. In short, with the power of natural language processing and machine learning, we are gradually moving towards a more convenient and efficient digital information era.
[0045] The embodiment of the present invention first provides a method for intelligent classification and retrieval of archives, which can be applied to electronic devices such as computer terminals, specifically ordinary computers.
[0046] The following takes running on a computer terminal as an example to illustrate it in detail. Figure 1 It is a hardware structure block diagram of a computer terminal for a method for intelligent classification and retrieval of archives provided by an embodiment of the present invention. As Figure 1 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.
[0047] The non-volatile storage medium can store an operating system and computer programs. The computer programs include program instructions, which when executed, can cause the processor to execute any method for intelligent classification and retrieval of archives.
[0048] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0049] The internal memory provides an environment for the operation of the computer programs in the non-volatile storage medium, and when the computer programs are executed by the processor, the processor can execute any method for intelligent classification and retrieval of archives.
[0050] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand, Figure 1The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0051] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0052] See Figure 2 , an embodiment of the present invention provides an intelligent file classification and retrieval method, which may include the following steps:
[0053] S201, according to the original data format of the file, use a multimodal data fusion algorithm to perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized file feature data set;
[0054] This method uses a multimodal data fusion algorithm to perform unified format conversion and feature extraction on text, image, and audio data according to the original data format of the file, forming a standardized file feature data set. This process first identifies the data characteristics of different modalities, such as the semantic information of text content, the visual characteristics of images, and the acoustic characteristics of audio. By comprehensively using deep learning techniques and feature extraction methods, the system integrates multiple data types into a unified feature set, ensuring that different types of data can cooperate efficiently during the processing and providing a more comprehensive information basis.
[0055] The implementation of this method significantly improves the intelligent level of file management and retrieval. By effectively fusing and standardizing different types of data, the system not only simplifies the subsequent analysis process of the data, but also greatly improves the processing ability of complex information. This file feature data set in a unified format lays a good foundation for subsequent classification and retrieval, enabling users to obtain relevant information faster and more accurately when searching for and managing files, improving work efficiency and the scientific nature of decision-making.
[0056] Specifically, the BERT model can be used to perform semantic parsing on text data, extract text feature vectors, use the convolutional neural network (CNN) to extract features from image data, generate image feature vectors, and perform spectral analysis on audio data through the short-time Fourier transform (STFT) and adaptive filters to extract audio feature vectors.
[0057] In this stage, the system uses specific techniques for feature extraction for different data modalities. Text data is semantically parsed through the BERT model to extract feature vectors with deep semantic information; image data uses a convolutional neural network (CNN) to extract spatial features, reflecting the visual information of the image; audio data is spectrally analyzed through the short-time Fourier transform (STFT) and adaptive filters to obtain the frequency-domain features of the audio signal. The extraction of these features not only ensures the accuracy of the information but also facilitates more complex subsequent data processing. This step enables the system to extract valuable information from different types of data, enhancing the flexibility and comprehensiveness of data processing. By using advanced feature extraction techniques, the features of each modality can be quantified and standardized, thus laying a solid foundation for subsequent feature fusion and intelligent classification and ensuring the accuracy and efficiency of subsequent analysis.
[0058] In this step, feature extraction is first performed on text data. The preprocessing of text data is a crucial step, including word segmentation and stop-word removal to ensure that important information is focused on during feature extraction. Next, the system adopts the BERT (Bidirectional Encoder Representations from Transformers) model, which is well-known for its bidirectional context understanding ability. In this stage, the input text is converted into corresponding word vectors, which can capture the deep semantic features of the text. For example, for the sentence "The cat is sitting on the mat", BERT will generate context-related embedding vectors for each word, enabling the relationship between "cat" and "mat" to be reflected semantically.
[0059] After completing the text feature extraction, the system turns to the processing of image data. Images usually contain rich visual information, and it is crucial to retain as much of this information as possible for subsequent classification and retrieval. To this end, a convolutional neural network (CNN) is used to process the image data. In this process, the image undergoes operations in multiple convolutional layers and pooling layers to extract local features and reduce the dimension. In specific operations, the system can adopt classic network architectures such as ResNet or VGG to extract the visual feature vectors of the image. These feature vectors can capture information such as the shape, color, and texture of the image. For example, after processing a photo containing a "cat", the resulting feature vector can present key visual features related to the cat, such as the shape of the ears and the fur color.
[0060] Finally, this step processes the audio data. Audio signals have time-domain and frequency-domain characteristics, so different processing methods are required. The system uses the Short-Time Fourier Transform (STFT) to convert the audio signal into a spectrogram, which can display the changes in time and frequency simultaneously. After STFT processing, the system further applies an adaptive filter to extract audio feature vectors, such as the pitch, volume, and frequency components of speech. These feature vectors can reflect important information in the audio data, such as the speaker's mood or speech rate. When the text, image, and audio data have all completed feature extraction, the system collates and summarizes the feature vectors of each modality for the upcoming fusion stage.
[0061] Normalize the obtained text feature vectors, image feature vectors, and audio feature vectors, and use a multi-modal Deep Belief Network (DBN) for feature fusion. Specifically, input the normalized text feature vectors, image feature vectors, and audio feature vectors into a multi-layer DBN. Through layer-by-layer pre-training and backpropagation fine-tuning, learn the deep associations between the features of each modality, and use an attention mechanism to weight the fused multi-modal feature vectors to highlight key features and suppress redundant information.
[0062] After completing feature extraction, the system normalizes the text feature vectors, image feature vectors, and audio feature vectors, and then uses a multi-modal Deep Belief Network (DBN) to fuse them. Normalization is to ensure the consistency of different features within the numerical range for easy fusion operations. DBN learns the deep associations between the features of each modality through layer-by-layer pre-training and backpropagation fine-tuning mechanisms, thus fusing comprehensive feature vectors. In addition, an attention mechanism is used to weight the fused multi-modal features to significantly enhance the importance of key features and reduce the impact of redundant information. The significance of this step is that through feature fusion, the system can form a more comprehensive feature representation, reflecting the multi-dimensional characteristics of the archival information. Such fused feature vectors can better represent the overall information of the archives, making subsequent classification and retrieval more accurate and efficient. At the same time, the attention mechanism helps the system focus on more important features and optimizes the performance of the model.
[0063] In this stage, the system first normalizes the feature vectors extracted from multiple modalities. The purpose of normalization is to eliminate the differences in the numerical ranges between different modalities and ensure that each feature has the same weight in the subsequent fusion process. This can be achieved through methods such as min-max normalization or z-score standardization. After normalization, the text feature vectors, image feature vectors, and audio feature vectors will be input into a multi-modal Deep Belief Network (DBN) for fusion. In each layer of the DBN, the network uses unsupervised learning to extract high-level abstract information from the features of each modality and gradually learn the potential associations between the features of different modalities.
[0064] During the fusion process, the training of DBN adopts a mechanism of layer-by-layer pre-training and backpropagation fine-tuning. In the pre-training stage, DBN trains the weights of each layer based on the given features, enabling the effective capture of the relationships between features. In the backpropagation fine-tuning stage, the network comprehensively considers the information of all modal features and optimizes the overall performance of the network through gradient descent. It is worth noting that an attention mechanism is adopted in this process, which dynamically adjusts the weights of different modal features, enabling more important features to occupy a larger proportion in the final fused feature vector. For example, when analyzing a video data, the color features of the image may be more important than the features of the voice, and the system will strengthen the performance of the image features in the fusion result through the attention mechanism.
[0065] After the fusion is completed, the system will obtain a feature vector in a unified format. This feature vector synthesizes multi-modal information and can comprehensively reflect the characteristics of the archives. This vector will be converted into a standardized format for subsequent processing steps. Finally, this fused feature vector will form a standardized archive feature dataset, providing stable basic data for subsequent classification and retrieval. This feature fusion not only improves the integrity of information but also lays a solid foundation for subsequent intelligent classification and efficient retrieval.
[0066] Perform a unified format conversion on the fused multi-modal feature vector. Adopt high-dimensional feature mapping technology to map it into a unified feature space, and use kernel functions to enhance the non-linear expression ability of the feature space to ensure the effective representation of different modal features in the unified space. Integrate the unified format feature vector into the standardized archive feature dataset.
[0067] After the feature fusion is completed, the fused multi-modal feature vector will further undergo a unified format conversion. Adopt high-dimensional feature mapping technology to map it into a unified feature space. Use kernel function technology to enhance the non-linear expression ability of the feature space, enabling different modal features to be effectively represented in the same feature space. This process ensures that all features are compared and processed in the same dimension, thus simplifying subsequent classification and retrieval operations. Through non-linear mapping, the system can overcome the problem of dimensional inconsistency that may occur during the processing of multi-modal features, ensuring that all features can be analyzed in the same context. This unified feature representation form provides stable input for subsequent classification models, helping to improve the accuracy and efficiency of classification and retrieval.
[0068] In this step, the fused multi-modal feature vectors will undergo a unified format conversion to ensure that all data can be compared in the same feature space. First, the system invokes high-dimensional feature mapping technology to map the fused feature vectors into a high-dimensional space. Here, kernel functions such as the Gaussian kernel function are usually used to enhance the non-linear expression ability of the feature space. This mapping enables effective comparison and analysis of features from different modalities at the same dimension. This means that non-linear relationships that cannot be fully expressed by traditional linear features can also be distinguished in the new feature space.
[0069] When performing feature mapping, the system needs to consider the distribution of features to ensure that all feature vectors can perform well in the new feature space. Therefore, the system will calculate the distance for each feature vector to confirm the similarity between features. This step can use measurement methods such as Euclidean distance or cosine similarity to help the system evaluate the relationship between different features. For example, the similarity evaluation of text features and image features can reveal the aggregation of the two modalities of "cat" in the feature space, thus better understanding its semantics.
[0070] Finally, the feature vectors mapped by the kernel function will be integrated into a standardized archive feature dataset, which is necessary for subsequent classification and retrieval steps. The feature set in a unified format provides clear input for the intelligent classification model and ensures the preservation of the relevance between all modal data. When a user makes a retrieval request, the system can quickly and accurately provide relevant archive information based on this standardized dataset. This undoubtedly improves the efficiency of data processing and the accuracy of retrieval, making the archive management system more intelligent in operation.
[0071] S202, Based on the deep learning-based adaptive classification model, classify the standardized archive feature dataset to generate multi-level archive classification labels;
[0072] The deep learning-based adaptive classification model aims to classify the standardized archive feature dataset to generate multi-level archive classification labels. The model design combines a convolutional neural network (CNN) and a long short-term memory network (LSTM) to give full play to the advantages of these two algorithms. CNN is responsible for extracting the spatial features of the input data, while LSTM focuses on the changes in time series data. Such a combination enables the model to fully understand complex archive data. At the same time, to prevent interference between features, an attention mechanism is adopted in the model. This mechanism assigns different weights to different features, thus emphasizing the influence of important features and ensuring that the generated classification labels have high accuracy and robustness.
[0073] By using an adaptive classification model based on deep learning to classify standardized archive features, not only the classification efficiency is improved, but also the burden of manual classification is reduced. The multi-level archive classification labels can provide a clearer and more accurate basis for subsequent retrieval and analysis, enabling users to find relevant information more quickly when querying archives. This intelligent processing method helps to improve the intelligence level of the archive management system, greatly shortening the time from data acquisition to information extraction and providing users with a better experience.
[0074] Specifically, a hybrid model based on the Convolutional Neural Network (CNN) and Long Short-Term Memory Network (LSTM) can be constructed. The CNN is used to extract the spatial information of features, and the LSTM is used to process the dynamic changes of the feature time series. When constructing the hybrid model, multiple parallel CNN branches are designed, and each branch uses different convolutional kernel sizes and strides to extract features from different scales. At the same time, to prevent feature interference, an attention mechanism is adopted to assign different weights to each branch.
[0075] At this stage, a hybrid model is designed by combining the advantages of CNN and LSTM to effectively extract the spatial and temporal information of archive features. The CNN uses multiple parallel branches, each branch with different convolutional kernel sizes and strides, and the extracted features include static information such as edges and textures. The LSTM is used to capture the dynamic changes of sequential data (such as time series). This combination allows the model to deeply understand the data from multiple dimensions and obtain a more comprehensive feature representation. The design of the hybrid model not only improves the effectiveness and accuracy of feature extraction, but also takes into account the diversity and complexity of the data, making the final generated classification labels more reliable when dealing with real-world data. Such a design provides a strong foundation for subsequent classification and information retrieval, and also enables the intelligent archive system to have higher flexibility and adaptability.
[0076] In the construction of the hybrid model, first select a suitable framework, such as TensorFlow or PyTorch, for model construction and training. The first part of the model uses a Convolutional Neural Network (CNN) to extract spatial features. We design multiple parallel convolutional layers with different kernel sizes (such as 3x3, 5x5, 7x7, etc.) to capture multi-scale feature information from the input data. After each convolutional layer, an activation function (such as ReLU) is connected, and then a pooling layer (such as max pooling or average pooling) is used for dimensionality reduction to reduce the size of the feature map while retaining key information. The final feature map is transformed into a one-dimensional vector through a Flatten layer to prepare for the subsequent fully connected layer. In addition, the outputs of multiple convolutional branches are concatenated in the network to form a comprehensive feature representation, laying the foundation for subsequent LSTM processing. To enhance the feature extraction ability, Batch Normalization technology is adopted to improve the training stability.
[0077] Next, construct a Long Short-Term Memory (LSTM) layer to capture the time series information of the input data. LSTM can effectively preserve past information and dynamically adjust memory, solving the problem of gradient disappearance that may occur in traditional RNNs when dealing with long sequences. In the model, the input to the LSTM layer is the feature vector output from the CNN layer, which is processed by several LSTM units to extract the temporal features in the data. At the same time, to improve the robustness and generalization ability of the model, a Dropout layer can be set to avoid overfitting. Specifically, a certain proportion of nodes can be randomly discarded after each LSTM unit. Through this randomization method, the performance of the model on new data is enhanced. Finally, the features output by the LSTM are concatenated with the features of the CNN part to form an integrated feature vector as the final representation of the model.
[0078] Finally, an attention mechanism is added to the model to improve the model's ability to focus on key information. The attention mechanism assigns different weights to different features, enabling the model to focus on important features during inference. Specifically, the weighted average method can be used to weight the concatenated feature vector to generate an optimized feature representation. The calculation of the weights can be achieved through a simple fully connected network that maps the feature vector and then obtains the weight values through the softmax function. Finally, the optimized feature vector is fed into the output layer to prepare for the subsequent classification task. This hybrid model combining CNN, LSTM, and attention mechanism can efficiently extract the spatial and temporal features of the input data, laying a good foundation for the classification task.
[0079] At the output layer of the model, a hierarchical Softmax mechanism is adopted to construct multi-level classification labels. First, the classification results of the first layer are output, and then, according to the classification results of the first layer, the subclass labels for the second-layer classification are correspondingly generated. Among them, for each category, by introducing label relationship modeling based on GCN, the dependency relationships between each label are recorded, and the inference ability of the model is enhanced by using the graph structure information of the labels;
[0080] The main purpose of this stage is to output classification labels layer by layer through the hierarchical Softmax mechanism. First, the model outputs the classification results of the first layer, and then generates the subclass labels for the second-layer classification according to the classification results of the first layer. By introducing a graph convolutional network (GCN) to model the dependency relationships between labels, the inference ability of the model is enhanced. The hierarchical Softmax mechanism allows the model to flexibly manage labels at different levels, thus achieving more refined classification. This method not only improves the classification accuracy but also enables the hierarchical relationship of data to be more intuitively reflected during retrieval, ensuring that users can quickly understand and obtain the required information.
[0081] In the design of the model output layer, a hierarchical Softmax mechanism is adopted to construct multi-level classification labels. First, a multi-level label structure needs to be defined to clarify the hierarchical relationship between each layer of labels. For example, the top-level category can be "document type", and the lower-level categories can be subdivided into "report", "invoice", "letter", etc. In the specific implementation, first create a graph structure in the model to represent the relationships between different classification labels. This graph structure is generated by GCN (graph convolutional network) and can capture the dependencies between labels. For each category, record its similarity and relevance with other labels through the graph structure so that this information can be utilized during model inference to generate more accurate labels.
[0082] In actual operation, the implementation of the hierarchical Softmax mechanism involves performing independent Softmax calculations on the labels at each level. First, the feature vector output by the model is linearly transformed to map to the dimension of the top-level labels, and the softmax function is applied to generate the probability distribution for each top-level label. Then, based on the top-level label with the highest probability, the corresponding subclass labels are further selected, and this process is repeated until the final classification labels are generated. This way of deconstructing layer by layer not only effectively reduces the computational burden but also ensures that the hierarchical relationship between labels is effectively utilized, improving the accuracy and precision of classification.
[0083] By introducing a loss function to guide the training process of the model, the hierarchical Softmax mechanism can optimize the classification effect. During the training process, an appropriate loss function (such as cross-entropy loss) and optimization algorithm (such as Adam or SGD) are selected to adjust the model parameters. In each training iteration, the loss is calculated based on the prediction results of the model and the actual labels, and the parameters of the network are updated according to the gradient descent method. This way of optimizing layer by layer ensures that the model can gradually improve the classification ability of each layer while retaining the overall hierarchical structure, thereby improving the overall classification performance and efficiency.
[0084] The constructed hybrid model is used as an adaptive classification model to infer the standardized archive feature dataset, and multi-level archive classification labels are output.
[0085] The last step is to apply the constructed hybrid model to the standardized archive feature dataset for actual classification inference and output multi-level archive classification labels. This model can automatically adapt to different data features and generate classification results that meet user needs. This inference process marks the transition of the model from theory to practice, being able to process real data and produce effective classification labels. Through automated classification, the efficiency of archive management can be greatly improved, providing necessary support for subsequent information retrieval and analysis, and thus achieving the goal of intelligent archive management.
[0086] In this step, the previously constructed hybrid model is applied to the standardized archive feature dataset for classification inference. First, the input data is preprocessed, including different modalities such as text, images, and audio, to ensure it conforms to the input format of the model. After these feature data are converted into a unified format, a standardized feature set is formed, which is convenient for the model to process effectively. Specifically, during the preprocessing process, the text will be tokenized and denoised, the images will be standardized and scaled, and the audio data will be denoised and feature extracted, such as MFCC (Mel Frequency Cepstral Coefficients), etc., to ensure that each data type is in a relatively unified feature space.
[0087] Next, the preprocessed data is input into the hybrid model for inference. The model will gradually extract and integrate the input features through the previously trained CNN and LSTM layers. In this process, the CNN is responsible for extracting spatial features, while the LSTM analyzes the time series features. After feature extraction, combined with the output of the attention mechanism, the model generates a complete feature vector, which will be used for subsequent classification. According to the architecture and training effect of the model, the inference process will be automatic. Through the established hierarchical Softmax mechanism, the model can generate corresponding multi-level classification labels to ensure the accuracy and relevance of the output.
[0088] Finally, the inferred classification labels can be used for subsequent archival information processing and retrieval. In practical applications, the user's retrieval request can be directly matched with the classification labels of the model to retrieve relevant archival information accordingly. This efficient classification inference method can significantly improve the efficiency of archival retrieval, enabling users to quickly locate the required documents based on multi-level labels. In an example application, in a large archival management system, when the user enters keywords, the system can immediately present a list of relevant documents according to the inference results, and the user can then make a selection, meeting the immediacy and accuracy requirements for information acquisition. This ability of automated classification inference not only improves the intelligence level of archival management but also provides a more user-friendly operation experience.
[0089] S203, through enhanced semantic network construction technology, analyze the semantic relationships between archival classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph;
[0090] In this stage, first perform co-occurrence analysis on the archival classification labels to calculate the co-occurrence frequency of different classification labels in the same archive. By constructing a co-occurrence matrix, represent the relationships between labels in numerical form, laying the foundation for subsequent semantic extraction. Subsequently, use non-negative matrix factorization (NMF) technology to decompose the co-occurrence matrix into two matrices, one representing the latent semantic features of the labels and the other representing the expression of the labels on these features. This process can map high-dimensional co-occurrence features to a low-dimensional semantic feature space, extracting the hidden semantic relationships between labels. Then, based on the latent semantic feature matrix, use the broadcast mechanism to construct a semantic relationship graph between labels. By defining nodes and edges, form a complete semantic association graph. Finally, apply graph neural network technology to optimize the graph, improving the connectivity and semantic relevance between labels.
[0091] Through this process, the generated optimized semantic association graph can effectively reveal the deep semantic relationships between archival classification labels, providing a more accurate semantic basis for subsequent classification and retrieval. The optimized graph not only makes the connections between labels closer but also improves the semantic matching efficiency during the retrieval process, ensuring that users can obtain more relevant information when querying. This directly enhances the intelligence level of the system, improving the user experience and facilitating the quick finding of required materials in complex archival data.
[0092] Specifically, co-occurrence analysis of the archival classification labels can be performed by calculating the co-occurrence frequency of different classification labels in the same archive to construct a co-occurrence matrix;
[0093] At this stage, first collect the classification label data of all files, and then analyze the occurrences of the labels in each file. By traversing each file and recording the co-occurrence times of the labels, a co-occurrence matrix is constructed, with the labels as rows and columns, and the elements of the matrix are filled with frequency information. For example, if label "A" co-occurs with label "B" in 10 files, then the value at the corresponding position in the co-occurrence matrix will be updated to 10. Continue this process until all combinations between the labels are counted to form a complete co-occurrence matrix, thus laying the foundation for the subsequent non-negative matrix factorization.
[0094] This process provides the basic data support for the subsequent extraction of potential semantic associations. The co-occurrence matrix effectively shows the relationships between the labels, which can help identify which labels often co-occur and further reveal the potential connections and similarities between them. Through this data-driven method, the subsequent analysis and optimization work will be more accurate, laying a good foundation for constructing an effective semantic graph.
[0095] First, collect the classification label data of all files and organize it into a structured database. The Pandas library in Python can be used to create a data frame, with each file as a row and each label as a column. Then, by traversing each file, record the number of times each pair of labels coexist in the same file. For each file, read the list of labels it contains and use a double loop to traverse its labels, updating the corresponding positions in the co-occurrence matrix. For example, if file A contains labels X and Y, then the connection between X and Y in the co-occurrence matrix will increase by 1.
[0096] After the co-occurrence analysis is completed, the co-occurrence matrix can be normalized. For example, each count value can be divided by the sum of the row to calculate the relative frequency and form a probability distribution, which can eliminate the influence of the different numbers of labels and make the analysis more fair and effective. At the same time, it is necessary to consider the occurrence frequencies of the labels to ensure that the co-occurrence matrix can not only reflect the simple occurrence situations but also reveal the frequently co-occurring label combinations.
[0097] Finally, the constructed co-occurrence matrix will serve as the basis for subsequent analysis. This matrix can not only clearly show the relationships between the labels but also lay a solid data foundation for further potential semantic analysis. By using visualization tools (such as Matplotlib or Seaborn), the co-occurrence matrix can be graphically displayed to facilitate the team members' intuitive understanding of the label relationships.
[0098] Based on the co-occurrence matrix, use the non-negative matrix factorization (NMF) technique to decompose the co-occurrence matrix into two matrices W and H. Among them, W represents the potential semantic features of the labels, and H represents the expressions of the labels on these features, in order to map the high-dimensional co-occurrence features to a low-dimensional semantic feature space and extract the potential semantic relationships of the labels;
[0099] In the process of decomposing the co-occurrence matrix using non-negative matrix factorization (NMF), the goal is first set to decompose the co-occurrence matrix into two non-negative matrices: W and H. The W matrix represents the relationship between labels and latent semantic features, while the H matrix represents the specific expression of labels on these features. In specific implementation, the NMF algorithm adjusts the values of W and H iteratively to make their product as close as possible to the original co-occurrence matrix. This process not only preserves the original information of the labels but also extracts the latent semantic features, enabling the relationship between labels to be expressed in a more concise way.
[0100] Through NMF decomposition, the data dimension can be effectively reduced and the latent semantic features between labels can be extracted, thus enhancing the learning ability of the model. This mapping is crucial because it allows the model to focus on the deep relationships between labels rather than simple surface co-occurrences, thereby promoting subsequent semantic association analysis and graph construction. Through this method, it can be clearly identified which labels share similar features and their semantic positions in the entire dataset can be established.
[0101] Before preparing to perform non-negative matrix factorization (NMF), it is first necessary to import relevant numerical calculation libraries such as NumPy, SciPy, or sklearn. Using the co-occurrence matrix as the input, set the parameters of NMF, including the dimensions of W and H after decomposition, as well as the number of iterations, etc. Enter the execution stage of the NMF algorithm, and use the method of random initialization to generate the initial W and H matrices. In each iteration, the algorithm adjusts the elements of these matrices according to the multiplicative update rule to gradually approximate the original input co-occurrence matrix. This process will continue until the preset convergence criterion is reached. For example, when the reconstruction error is lower than a certain threshold, the iteration stops. It should be noted that NMF itself requires that the elements of the input matrix must be non-negative. For this reason, it may be necessary to preprocess the co-occurrence matrix to ensure that all count values are positive. Through appropriate regularization techniques, NMF can be prevented from falling into local optima, ensuring that the finally obtained W and H matrices have good generalization ability. At the same time, in order to better understand the latent semantic relationships between labels, each column vector of the W matrix will correspond to a latent semantic feature, while the H matrix can reflect the specific expression of each label on these features. The final output of NMF will provide the necessary basic data for subsequent construction of the semantic association graph, ensuring that the latent connections between labels are effectively extracted. Through subsequent analysis, the feature matrix W can be visualized, for example, using the clustering method, to help intuitively understand the relationship between label groups and semantic features.
[0102] According to the latent semantic feature matrix W, use the broadcasting mechanism to construct the semantic relationships between labels to obtain the semantic association graph. Among them, nodes are created for each label, and the edge weights are defined by the similarity between labels;
[0103] At this stage, the potential semantic feature matrix W is used to construct a semantic association graph. First, a node is created for each classification label and embedded into the graph structure. Then, the weights of the edges are defined by calculating the similarity between the labels. Usually, cosine similarity or Euclidean distance can be used to quantify the degree of similarity. Using the broadcast mechanism, the feature vectors in the matrix W can be extended so that each label node can be paired with other label nodes, and the similarity scores between them can be calculated, thus generating a complete semantic association graph that clearly shows the relationships between the labels.
[0104] This construction process can intuitively express the semantic relationships between the labels, enabling each label node to form corresponding connections in the graph. This not only enhances the visualization effect of the data but also lays a foundation for subsequent semantic matching and information retrieval. Through the optimized graph, the intelligence level of the model in processing complex queries can be improved, effectively enhancing the accuracy and efficiency of retrieval.
[0105] When constructing the semantic association graph, first, a node needs to be created for each label. In this process, a graph theory-related toolkit (such as NetworkX) can be used to construct the graph structure. The creation of each node can be achieved by assigning a unique identifier to the label while saving the label information, including its name and relevant data on semantic features. Next, using the data in the potential semantic feature matrix W, the similarity between the labels is calculated, usually evaluated using cosine similarity. After calculating the similarity of all labels, the similarity values can be stored as edge weights. Using the broadcast mechanism, the row vectors of the W matrix can be extended to all combinations of each label with other labels, thus simplifying the process of calculating similarity. By setting a threshold, the edges between labels with low similarity can be filtered out, only retaining relatively meaningful connections, thus ensuring the simplicity and information content of the semantic association graph. The semantic association graph will be able to intuitively display the relationships between each label, helping to understand the similarities and interconnections between the labels. Further, the graph structure can be used for visualization, enabling team members to quickly understand the relationships between the labels and providing clues for later optimization of the retrieval system.
[0106] Using graph embedding technology, an embedding vector for each node is generated based on the context within the semantic association graph, and the edge weights of the semantic association graph are updated using the embedding vectors to optimize the semantic association graph, so that labels with similar semantic relationships have higher connectivity in the graph.
[0107] At this stage, the semantic association graph is further optimized through graph embedding technology. The core idea of graph embedding technology is to map the structural information and attribute information of nodes in the graph into a low-dimensional dense space, so that similar node features are closer in the embedding space. Specifically, when operating, algorithms (such as DeepWalk or Node2Vec) are used to perform random walks on each node to capture the local relationships between nodes and generate embedding vectors for each node. These embedding vectors not only reflect the features of each label but also consider its position and context information in the graph. Then, the edge weights in the graph are adjusted and optimized through these embedding vectors, so that labels with similar semantics have higher connectivity in the graph.
[0108] Through graph embedding technology, the expressive ability of the semantic association graph can be effectively improved, making labels with strong relevance more connected in the graph. The optimized graph further enhances the reasoning ability of the model, thus achieving more accurate and efficient semantic matching in the actual retrieval process. When retrieving, users can more easily find the labels related to their input, thereby improving the sensitivity and practicality of the entire file retrieval system.
[0109] At this stage, first, a suitable graph embedding technology is selected. Commonly used algorithms include DeepWalk, Node2Vec, or GraphSAGE, etc. These algorithms generate feature vectors of nodes by performing random walks or neighbor sampling on the graph. In actual implementation, the parameter settings of the random walk are defined, such as the number of walks, the length of the walk, and the step size of each walk. By flexibly adjusting these parameters, the context relationships between nodes can be better captured. After generating the node feature vectors, the similarity between each pair of nodes needs to be calculated next and applied to update the edge weights in the graph. For example, cosine similarity or L2 distance can be used to quantify the similarity between nodes. In this process, the update of the edge weights helps to reflect the changes in the current semantic information and enhance the connectivity between nodes with similar semantics. It should be noted that the update of the edge weights may need to be iterated multiple times to ensure that the final graph fully reflects the latest semantic relationships. After completing graph embedding and edge weight update, the optimized semantic association graph will provide strong support for subsequent retrieval. When users perform a retrieval, the system can more accurately understand and match the user's query intention and return the most relevant label and file information. Finally, the optimized graph not only improves the intelligence level of the system but also greatly enhances the user experience, making the information retrieval process more efficient and accurate.
[0110] S204, according to the retrieval request input by the user, use the optimized semantic association graph for semantic matching and relevance ranking, and output the file information most relevant to the retrieval request.
[0111] According to the retrieval request input by the user, semantic matching and relevance ranking are performed using the optimized semantic association graph. First, multimodal parsing is performed on the retrieval content to convert inputs in different formats such as text, speech, or images into corresponding embedding features. For text input, the system uses a pre-trained language model (such as BERT) to parse the text and extract semantic embedding features to ensure that the user's query intent can be accurately captured. When processing speech input, the system first converts the speech content into text through speech recognition technology and then uses the language model to extract semantic embedding features. For image input, convolutional neural networks are used to extract image embedding features. After processing all the embedding features of different modalities, the system combines these features into a unified retrieval request embedding feature through a multimodal fusion mechanism for querying in the feature space consistent with the archive feature dataset. Then, the system queries the nodes closest to the retrieval request in the optimized semantic association graph and uses graph embedding technology to generate the semantic embedding of each node. The system calculates the similarity score between the retrieval request embedding feature and the node embedding features in the graph, selects several initial nodes and their direct adjacent nodes most relevant to the retrieval request as preliminary candidate nodes, and further obtains the corresponding preliminary candidate archive set.
[0112] The significance of this process is that through multimodal parsing and the optimized semantic association graph, the system can more accurately understand the user's retrieval intent, no longer limited to simple keyword matching, but deeply analyzing the deep meaning of the user's query. This semantic matching ability enables the system to quickly discover the information most relevant to the user's needs in complex archive data, significantly improving the accuracy and efficiency of retrieval. In addition, through the use of graph embedding technology, the system can effectively utilize the correlation between tags, associate the user's retrieval request with potentially similar archives, and provide richer and more accurate retrieval results, thus enhancing the user experience.
[0113] Specifically, for the retrieval request input by the user, first perform multimodal parsing on the retrieval content, convert the text query, speech input, or image content in the retrieval content into embedding features of different modalities. Among them, if it is a text input query, use a pre-trained language model to parse the retrieval text and extract semantic embedding features to ensure that the extracted semantic embedding features can capture the user's intent; if it is a speech input query, convert the speech content into text through speech recognition and further use the language model to extract semantic embedding features; if it is an image input, use convolutional neural networks to extract image embedding features;
[0114] At this stage, the system needs to perform multimodal parsing on the retrieval content input by the user, identify the format of the input, and extract corresponding features. For text queries, the system usually uses a pre-trained language model (such as BERT) to perform semantic parsing on the input text and extract the semantic embedding features of the text, so as to accurately capture the user's intention. When processing voice input, first, the voice information needs to be converted into text through speech recognition technology, and then the text parsing model is used to extract its semantic embedding features. For image input, the system uses a convolutional neural network to extract the feature vectors of the image, and these feature vectors can reflect the semantic information of the image content. Through this process, the system can uniformly convert user inputs from different modalities into available feature representations, thus laying a foundation for subsequent semantic matching. Such a processing method greatly enhances the flexibility and adaptability of the system, can meet the scenarios where users express their needs through different media, and improves the accuracy and comprehensiveness of retrieval.
[0115] In this step, the system first needs to identify the type of retrieval content input by the user and determine whether it is text, voice, or image. For text input, the system will call a pre-trained language model (such as BERT or GPT, etc.) and use natural language processing (NLP) technology to analyze the text. This includes word segmentation, stop word removal, word vector representation, etc., to convert the text into computable semantic embedding features. Each word is mapped to a high-dimensional vector, and its context information is also taken into account, so as to capture the intention and nuances of the user's query to the greatest extent.
[0116] For voice input, the system will first use speech recognition (ASR) technology to convert the audio signal into text in real time. In this process, the system needs to ensure the recognition accuracy and use a deep learning model (such as a sequence-to-sequence model with CTC loss function) to handle different accents and noise interferences. The converted text will be processed through the same NLP technology to generate semantic embedding features. In addition, the system also needs to consider the processing delay to ensure a quick response after the user input and improve the user experience.
[0117] When processing image input, the system will use a convolutional neural network (CNN) to extract features from the image. Specifically, after the image is uploaded, the system first performs preprocessing, including image scaling, normalization, and enhancement, etc., to ensure that the model can better understand the image content. Then, the extracted feature vectors will be transformed into low-dimensional embeddings to reflect the key information of the image. After these processes are completed, the system integrates the embedding features of the three modalities for subsequent use.
[0118] Fuse the embedding features of different modalities through a multimodal fusion mechanism into a unified retrieval request embedding feature to map to the feature space consistent with the archive feature dataset;
[0119] Multimodal fusion is an important process of integrating features from different sources into a unified representation. In this step, the system will effectively combine the embedded features of text, speech, and images extracted in Step 1 through a multimodal fusion mechanism to generate a unified retrieval request embedded feature. This process requires designing appropriate fusion strategies, such as weighted average method, concatenation method, or other deep learning fusion models, to ensure that the features of each modality can complement each other and enhance the overall expressive ability. By fusing the embedded features of different modalities, the system can comprehensively capture the user's query intention and reduce information loss caused by insufficient single-modal information. The fused feature vector can better reflect the user's needs, improve the understanding of complex queries, and thus enhance the relevance and accuracy of the retrieval.
[0120] In this stage, to achieve effective fusion of multimodal features, the system designs a modular fusion mechanism. First, the system will introduce a data preprocessing step to normalize the features of different modalities to ensure that the features of each modality are fused on the same scale. Then, the system will assign different weights to each modality according to its characteristics. The weights can be set based on previous experimental results or optimized through dynamic learning strategies. Through these steps, the system ensures that each modality occupies an appropriate proportion in the final feature representation.
[0121] Then, the system can adopt various fusion strategies, such as weighted average, concatenation, or more complex deep learning integration methods. Taking weighted average as an example, the system will multiply the features of each modality by their weights and then add them up to generate a unified retrieval request embedded feature. If the concatenation method is used, the system will concatenate the feature vectors of each modality by dimension to form a higher-dimensional feature representation. Through such strategies, the system can make full use of the advantages of different modalities, organically combine various types of information, and improve the accuracy of subsequent matching.
[0122] Finally, the unified retrieval request embedded feature after fusion will be mapped to a feature space consistent with the archive feature dataset. This mapping can be achieved through various methods such as linear transformation or non-linear transformation to ensure that the fused features can be effectively compared with the features of the existing archive data. This process not only lays a solid foundation for subsequent similarity calculation but also ensures the sensitivity and adaptability of the system to multimodal queries.
[0123] Query the node in the optimized semantic association graph that is closest to the retrieval request, generate the semantic embedding of each node using graph embedding technology, and calculate the similarity score between the retrieval request embedded feature and the node embedded feature in the graph;
[0124] Once the unified retrieval request embedding feature is obtained, the system will search in the optimized semantic association graph to find the node that is closest to this embedding feature. In this process, the system uses graph embedding technology to generate the semantic embedding of each node to capture the context relationship of the node in the graph. Then, by calculating the similarity scores between the retrieval request embedding feature and the embedding features of each node in the graph, nodes with higher similarity are found. In this way, the system can accurately identify the archival information most relevant to the user's retrieval request, and then provide higher-quality search results. This process uses the semantic association graph to strengthen the semantic understanding ability of retrieval through similarity matching, ensuring that users can quickly obtain the relevant materials they need.
[0125] In this step, the system first needs to construct and optimize the semantic association graph to ensure that the nodes and edges in the graph can accurately express the relationships between each file. The system will use a graph database (such as Neo4j) to design the attributes of the nodes and the weights of the edges, so that the graph can be dynamically updated and reflect the latest semantic relationships. Then, the retrieval request embedding feature is used as a query vector to match the node features in the semantic association graph.
[0126] Next, the system will generate semantic embeddings for each node through graph embedding technology. This process usually uses graph neural networks (GNNs) or node embedding algorithms (such as Node2Vec or GraphSAGE) to dynamically generate vector representations of nodes. The semantic embedding of each node will consider the feature information of its surrounding adjacent nodes to more comprehensively reflect the semantic context of the node in the graph. The system will continuously update the node features during this process to ensure that each node in the graph has the latest context semantic information.
[0127] Finally, the system will calculate the similarity scores between the retrieval request embedding feature and the embedding features of each node. Similarity calculation can be performed through methods such as cosine similarity, Euclidean distance, or Manhattan distance. The system will screen out the nodes closest to the user's retrieval request according to the similarity scores to ensure that the subsequent candidate file set has the highest relevance. This process not only improves the accuracy of retrieval but also lays a foundation for the subsequent selection of candidate nodes.
[0128] According to the similarity scores, several initial nodes and their direct adjacent nodes that are most relevant to the retrieval request are selected as preliminary candidate nodes to obtain the corresponding preliminary candidate file set;
[0129] After calculating the similarity scores between the retrieval request embedding features and the node embedding features in the graph, the system will select a number of initial nodes that are most relevant to the retrieval request based on these scores. The initial nodes include not only the nodes with the highest similarity, but also their direct adjacent nodes in the graph to expand the scope of the retrieval. This strategy enables the system to further enrich the number of candidate nodes based on the relevance ranking to ensure the comprehensiveness and diversity of the provided archival information. Through this method, the system provides a set of preliminary candidate archives, ensuring that the retrieval results cover both relevance and diversity, increasing the user's chance of obtaining information. At the same time, this method also reduces the risk of information deviation caused by a single node, making the system's recommendation results more reliable.
[0130] In this step, the system first sorts all the nodes according to the similarity scores calculated in the previous step and selects a number of initial nodes with the highest scores. These nodes represent the archival information that is semantically closest to the user's query request. Then, the system will search for the adjacent nodes directly connected to these initial nodes according to the topological structure of the initial nodes. These adjacent nodes may also contain information relevant to the user's request, so their addition can enrich the diversity of the candidate archives.
[0131] Next, the system will integrate the selected initial nodes and their adjacent nodes to create a candidate archive set. To ensure the quality of the candidate set, the system will also set a threshold. For example, the similarity score must be higher than a certain set value to be included in the candidate set. This process will ensure that the generated candidate archive set is guaranteed in terms of relevance and avoid redundant information caused by the increase in the number of candidate nodes.
[0132] Finally, the system will conduct a preliminary retrieval on the obtained candidate archive set and generate a list for subsequent information display. This candidate set includes not only the nodes with the highest scores, but also integrates the information of adjacent nodes with strong relevance so that the system can comprehensively cover the user's retrieval needs. This step provides a solid foundation for subsequent more in-depth retrieval and information presentation.
[0133] Use a graph neural network to perform message propagation on the preliminary candidate nodes, and propagate the semantic information of the adjacent nodes into the retrieval matching nodes to enrich the context representation of the nodes;
[0134] After obtaining the preliminary candidate nodes, the system will use a Graph Neural Network (GNN) for message propagation. Through the architecture of the GNN, nodes can communicate information with each other, and the semantic features of adjacent nodes will be propagated to the target retrieval matching nodes. This process is crucial because it can significantly enhance the context representation of the nodes, making the information of the nodes in the retrieval context more abundant. Through this message propagation mechanism, the system can better capture the deep semantic relationships between nodes, thereby improving the relevance and accuracy of the retrieval results. This dynamic information update also enables the system to adapt to different context changes and improve the flexibility and intelligence level of the retrieval.
[0135] In this stage, the system inputs the preliminary candidate nodes into the Graph Neural Network (GNN) and utilizes its powerful message propagation ability to enhance the feature representation of each node. First, the system initializes the features of each candidate node and prepares for multiple rounds of message propagation in the GNN. In each round, the candidate node receives information from its adjacent nodes and then aggregates this information with its own features to enhance its context representation.
[0136] The message propagation mechanism of the GNN usually includes two main steps: neighborhood aggregation and feature update. In the neighborhood aggregation stage, the system summarizes the features of adjacent nodes according to the connection relationship of the nodes, possibly by means of averaging, summing, or weighted summing, etc. Then, the system merges these aggregated adjacent node features with the features of the current node to update the feature representation of the node. This process is completed through multiple rounds of iteration, and the system continuously improves the context information of the node in each round, making the final node representation more abundant and accurate.
[0137] Finally, the system normalizes the node embeddings after multiple rounds of message propagation for subsequent similarity calculation. This process will ensure that the features of each node have a unified scale, facilitating effective comparison with the retrieval request. Through this message propagation mechanism, the system can better capture the complex relationships between nodes and extend the context of information to a broader semantic level, significantly enhancing the retrieval relevance of candidate archives.
[0138] Optimize the semantic embedding of each node through multiple rounds of propagation iteration to better represent the semantic features in the current retrieval context;
[0139] After the initial message propagation is completed, the system will continue with multiple rounds of propagation iterations to further optimize the semantic embedding of each node. In each iteration, the system will update the node features using the previous node feature updates and the feature information of adjacent nodes to continuously improve the context representation of the nodes. This process ensures that the final node features not only reflect the characteristics of the nodes themselves but also fully consider the relevant context information. Through multiple rounds of propagation, the system can more accurately capture the implicit semantics in the retrieval request and improve the quality of the retrieval results. This iterative optimization approach enables the nodes to adapt to different retrieval contexts, thus providing users with more relevant archival information.
[0140] In this step, the system will perform multiple rounds of iterations to further optimize the semantic embedding of each node. In each iteration, the feature vector of the node will be updated. The system will set an upper limit on the number of iterations (such as 3 to 5 times) and will also monitor the convergence of the node features. If the change in the node features in a certain iteration is less than the set threshold, the system can stop further iterations to avoid unnecessary computational overhead.
[0141] During the iteration process, the system will use the same message propagation mechanism to update the node features. In each iteration, the node will not only collect the features from adjacent nodes but also adjust the aggregation and update strategies to better adapt to the current retrieval context. For example, the system may introduce a context-sensitive weight adjustment mechanism to dynamically adjust the influence of adjacent nodes according to the features of the current retrieval request, thus ensuring that the most relevant information can be better incorporated into the context representation of the current node.
[0142] Finally, the system will ensure that the feature representation of each node is significantly improved after multiple rounds of iterations, forming an embedding that can fully reflect its semantic information in the current context. When the iteration ends, the system will store the optimized node embeddings for subsequent similarity calculations and candidate archival matching. This process not only enhances the context expression ability of the nodes but also lays a solid foundation for the final retrieval results.
[0143] Calculate the similarity score again between the propagated node embeddings and the retrieval request embedding, sort the initial candidate archival set, and sequentially output the candidate archival information most relevant to the retrieval request.
[0144] After multiple rounds of propagation and feature optimization, the system compares the final node embeddings with the initially generated retrieval request embeddings and recalculates the similarity scores. Using the same similarity measurement method (such as cosine similarity), it re-evaluates the relevance of each candidate node in the current retrieval context. Finally, the system sorts the initial candidate file set based on these similarity scores and sequentially outputs the candidate file information that is most relevant to the retrieval request. This process ensures that the search results finally returned to the user are based on the latest and most relevant context information, thereby improving the user's retrieval satisfaction and information acquisition efficiency. Through such platform-based content analysis and dynamic matching, the information retrieval system can meet the complex and ever-changing user needs.
[0145] In the last step, the system compares the node vectors after multiple rounds of message propagation and semantic embedding optimization with the retrieval request embedding features. To calculate the similarity scores, the system will use the same similarity measurement methods as before, such as cosine similarity, Euclidean distance, etc. By comparing the retrieval request embedding with the embedding features of each candidate node, the system will generate a new list of similarity scores to determine which candidate file information is most relevant to the user's retrieval request.
[0146] Next, the system sorts the obtained similarity scores and arranges the candidate nodes from high to low. To enhance the user experience, the system sets a threshold to ensure that only candidate files with scores exceeding a certain standard are returned, which can avoid returning results that are not relevant enough to the request. After sorting, the system generates a clear list of candidate files, and these files will be ready to be presented to the user.
[0147] Finally, the system sequentially outputs the candidate file information that is most relevant to the retrieval request according to the sorting result, including the key content, creation time, abstract, and other relevant metadata of the file. This process not only ensures that users can quickly find the information they need, but also provides rich background information to help users understand and select the most suitable file. At the same time, the system records the user feedback for each retrieval for future further optimization and upgrade. This series of steps ensures the accuracy of information retrieval and user satisfaction.
[0148] It can be seen that according to the original data format of the archives, using the multimodal data fusion algorithm, the text, image and audio data are subjected to unified format conversion and feature extraction to obtain a standardized archive feature data set; based on the deep learning-based adaptive classification model, the standardized archive feature data set is classified to generate multi-level archive classification labels; through the enhanced semantic network construction technology, the semantic relationships between the archive classification labels are analyzed to generate a semantic association graph, and the semantic association graph is optimized to obtain an optimized semantic association graph; according to the retrieval request input by the user, the optimized semantic association graph is used for semantic matching and relevance ranking, and the archive information most relevant to the retrieval request is output, so as to improve the relevance and accuracy of the retrieval and promote the modernization and intelligentization of archive management.
[0149] Another embodiment of the present invention provides an intelligent archive classification and retrieval system. Refer to Figure 3 , the system may include:
[0150] A conversion module 301, configured to perform unified format conversion and feature extraction on text, image and audio data according to the original data format of the archives by using a multimodal data fusion algorithm to obtain a standardized archive feature data set;
[0151] A classification module 302, configured to classify the standardized archive feature data set based on a deep learning-based adaptive classification model to generate multi-level archive classification labels;
[0152] A generation module 303, configured to analyze the semantic relationships between the archive classification labels by using the enhanced semantic network construction technology to generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph;
[0153] An output module 304, configured to perform semantic matching and relevance ranking according to the retrieval request input by the user by using the optimized semantic association graph, and output the archive information most relevant to the retrieval request.
[0154] It can be seen that according to the original data format of the archives, using the multimodal data fusion algorithm, the text, image and audio data are subjected to unified format conversion and feature extraction to obtain a standardized archive feature data set; based on the deep learning-based adaptive classification model, the standardized archive feature data set is classified to generate multi-level archive classification labels; through the enhanced semantic network construction technology, the semantic relationships between the archive classification labels are analyzed to generate a semantic association graph, and the semantic association graph is optimized to obtain an optimized semantic association graph; according to the retrieval request input by the user, the optimized semantic association graph is used for semantic matching and relevance ranking, and the archive information most relevant to the retrieval request is output, so as to improve the relevance and accuracy of the retrieval and promote the modernization and intelligentization of archive management.
[0155] An embodiment of the present invention also provides a storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0156] Specifically, in this embodiment, the above storage medium can be configured to store a computer program for executing the following steps:
[0157] S201, according to the original data format of the file, using a multimodal data fusion algorithm, perform unified format conversion and feature extraction on text, image and audio data to obtain a standardized file feature data set;
[0158] S202, based on an adaptive classification model of deep learning, perform classification processing on the standardized file feature data set to generate multi-level file classification labels;
[0159] S203, through an enhanced semantic network construction technology, analyze the semantic relationships between file classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph;
[0160] S204, according to the retrieval request input by the user, use the optimized semantic association graph for semantic matching and relevance ranking, and output the file information most relevant to the retrieval request.
[0161] It can be seen that according to the original data format of the file, using a multimodal data fusion algorithm, perform unified format conversion and feature extraction on text, image and audio data to obtain a standardized file feature data set; based on an adaptive classification model of deep learning, perform classification processing on the standardized file feature data set to generate multi-level file classification labels; through an enhanced semantic network construction technology, analyze the semantic relationships between file classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph; according to the retrieval request input by the user, use the optimized semantic association graph for semantic matching and relevance ranking, and output the file information most relevant to the retrieval request, so as to improve the relevance and accuracy of retrieval and promote the modernization and intelligentization of file management.
[0162] An embodiment of the present invention also provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0163] Specifically, the above electronic device may further include a transmission device and an input / output device. Wherein, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0164] Specifically, in this embodiment, the above-mentioned processor can be set to execute the following steps through a computer program:
[0165] S201, according to the original data format of the file, use the multi-modal data fusion algorithm to perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized file feature data set;
[0166] S202, based on the deep learning-based adaptive classification model, classify the standardized file feature data set to generate multi-level file classification labels;
[0167] S203, through the enhanced semantic network construction technology, analyze the semantic relationships between file classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph;
[0168] S204, according to the retrieval request input by the user, use the optimized semantic association graph to perform semantic matching and relevance ranking, and output the file information most relevant to the retrieval request.
[0169] It can be seen that according to the original data format of the file, use the multi-modal data fusion algorithm to perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized file feature data set; based on the deep learning-based adaptive classification model, classify the standardized file feature data set to generate multi-level file classification labels; through the enhanced semantic network construction technology, analyze the semantic relationships between file classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph; according to the retrieval request input by the user, use the optimized semantic association graph to perform semantic matching and relevance ranking, and output the file information most relevant to the retrieval request, so as to improve the relevance and accuracy of retrieval and promote the modernization and intelligentization of file management.
[0170] The above has detailed the structure, features, and function effects of the present invention according to the illustrated embodiments. The above is only the preferred embodiment of the present invention, but the present invention is not limited to the scope defined by the drawings. Any changes made according to the concept of the present invention, or modified into equivalent embodiments with equivalent changes, still fall within the spirit covered by the specification and the drawings, and should be within the protection scope of the present invention.
Claims
1. An intelligent file classification and retrieval method, characterized in that The method includes: According to the original data format of the archives, using a multi-modal data fusion algorithm, perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized archive feature dataset; Based on a deep learning-based adaptive classification model, perform classification processing on the standardized archive feature dataset to generate multi-level archive classification labels; among them, construct a hybrid model based on a convolutional neural network (CNN) and a long short-term memory network (LSTM). The CNN is used to extract the spatial information of features, and the LSTM is used to process the temporal sequence dynamic changes of features. When constructing the hybrid model, design multiple parallel CNN branches, each branch using different convolutional kernel sizes and strides to extract features from different scales. At the same time, to prevent feature interference, adopt an attention mechanism to assign different weights to each branch; at the output layer of the model, adopt a hierarchical Softmax mechanism to construct multi-level classification labels. First, output the classification results of the first layer, and then, according to the classification results of the first layer, correspondingly generate subclass labels for the second-layer classification. Among them, for each category, by introducing a label relationship modeling based on a graph convolutional network (GCN), record the dependency relationship between each label, and utilize the graph structure information of the labels to enhance the inference ability of the model; use the constructed hybrid model as the adaptive classification model to infer the standardized archive feature dataset and output multi-level archive classification labels; Through enhanced semantic network construction technology, analyze the semantic relationships between archive classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph; among them, perform co-occurrence analysis of the archive classification labels. By calculating the co-occurrence frequency of different classification labels in the same archive, construct a co-occurrence matrix; based on the co-occurrence matrix, use non-negative matrix factorization (NMF) technology to decompose the co-occurrence matrix into two matrices W and H, where W represents the latent semantic features of the labels, and H represents the expression of the labels on these features, so as to map the high-dimensional co-occurrence features to a low-dimensional semantic feature space and extract the latent semantic relationships between the labels; according to the latent semantic feature matrix W, use a broadcast mechanism to construct the semantic relationships between the labels to obtain a semantic association graph, where nodes are created for each label, and the edge weights are defined by the similarity between the labels; adopt graph embedding technology to generate embedding vectors for each node based on the context within the semantic association graph, and use the embedding vectors to update the edge weights of the semantic association graph to optimize the semantic association graph, so that labels with similar semantic relationships have higher connectivity in the graph; According to the retrieval request input by the user, use the optimized semantic association graph for semantic matching and relevance ranking, and output the archive information most relevant to the retrieval request.
2. The method according to claim 1, characterized in that The step of according to the original data format of the archives, using a multi-modal data fusion algorithm, performing unified format conversion and feature extraction on text, image, and audio data to obtain a standardized archive feature dataset includes: Semantically parse the text data through the BERT model to extract text feature vectors, extract features from the image data through the convolutional neural network CNN to generate image feature vectors, and perform spectral analysis on the audio data through the short-time Fourier transform STFT and adaptive filters to extract audio feature vectors; Normalize the obtained text feature vectors, image feature vectors, and audio feature vectors, and use the multi-modal deep belief network DBN for feature fusion. Among them, input the normalized text feature vectors, image feature vectors, and audio feature vectors into the multi-layer DBN, and through layer-by-layer pre-training and backpropagation fine-tuning, learn the deep associations between the features of each modality, and use the attention mechanism to weight the fused multi-modal feature vectors to highlight key features and suppress redundant information; Convert the fused multi-modal feature vectors into a unified format, adopt the high-dimensional feature mapping technology to map them into a unified feature space, and use the kernel function to enhance the non-linear expression ability of the feature space to ensure the effective representation of different modality features in the unified space, and integrate the feature vectors in the unified format into the standardized archive feature dataset.
3. The method according to claim 2, wherein According to the retrieval request input by the user, use the optimized semantic association graph for semantic matching and relevance ranking, and output the archive information most relevant to the retrieval request, including: For the retrieval request input by the user, first perform multi-modal parsing on the retrieval content, and convert the text query, voice input, or image content in the retrieval content into embedded features of different modalities. Among them, if it is a text input query, use the pre-trained language model to parse the retrieval text and extract semantic embedded features to ensure that the extracted semantic embedded features can capture the user's intention; if it is a voice input query, convert the voice content into text through speech recognition and further use the language model to extract semantic embedded features; if it is an image input, use the convolutional neural network to extract image embedded features; Fuse the embedded features of different modalities into a unified retrieval request embedded feature through the multi-modal fusion mechanism to map into the feature space consistent with the archive feature dataset; Query the node closest to the retrieval request in the optimized semantic association graph, use the graph embedding technology to generate the semantic embedding of each node, and calculate the similarity score between the retrieval request embedded feature and the node embedded feature in the graph; According to the similarity score, select several initial nodes and their direct adjacent nodes most relevant to the retrieval request as preliminary candidate nodes to obtain the corresponding preliminary candidate archive set; Use the graph neural network to perform message propagation on the preliminary candidate nodes, and propagate the semantic information of the adjacent nodes into the retrieval matching nodes to enrich the context representation of the nodes; Optimize the semantic embedding of each node through multiple rounds of propagation iteration to better represent the semantic features in the current retrieval context; Calculate the similarity score between the propagated node embedding and the retrieval request embedding again, sort the preliminary candidate archive set, and sequentially output the candidate archive information most relevant to the retrieval request.
4. An intelligent file classification and retrieval system, characterized in that, The system includes: A conversion module for performing unified format conversion and feature extraction on text, image, and audio data according to the original data format of the file by using a multi-modal data fusion algorithm to obtain a standardized file feature data set; A classification module for classifying the standardized file feature data set based on an adaptive classification model of deep learning to generate multi-level file classification labels; among them, a hybrid model based on a convolutional neural network CNN and a long short-term memory network LSTM is constructed. CNN is used to extract the spatial information of features, and LSTM is used to process the time series dynamic changes of features. When constructing the hybrid model, multiple parallel CNN branches are designed, and each branch uses different convolutional kernel sizes and strides to extract features from different scales. At the same time, to prevent feature interference, an attention mechanism is used to assign different weights to each branch; at the output layer of the model, a hierarchical Softmax mechanism is adopted to construct multi-level classification labels. First, the classification results of the first layer are output, and then according to the classification results of the first layer, the subclass labels of the second layer classification are generated correspondingly. Among them, for each category, by introducing label relationship modeling based on GCN, the dependency relationship between each label is recorded, and the inference ability of the model is enhanced by using the graph structure information of the labels; the constructed hybrid model is used as the adaptive classification model to infer the standardized file feature data set and output multi-level file classification labels; A generation module for analyzing the semantic relationship between file classification labels through enhanced semantic network construction technology, generating a semantic association graph, and optimizing the semantic association graph to obtain an optimized semantic association graph; among them, co-occurrence analysis of the file classification labels is performed, and a co-occurrence matrix is constructed by calculating the co-occurrence frequency of different classification labels in the same file; based on the co-occurrence matrix, the non-negative matrix factorization NMF technology is used to decompose the co-occurrence matrix into two matrices W and H, where W represents the latent semantic features of the labels, and H represents the expression of the labels on these features, so as to map the high-dimensional co-occurrence features to a low-dimensional semantic feature space and extract the latent semantic relationship between the labels; according to the latent semantic feature matrix W, the semantic relationship between the labels is constructed by using the broadcast mechanism to obtain a semantic association graph, where nodes are created for each label, and the edge weights are defined by the similarity between the labels; graph embedding technology is adopted to generate the embedding vector of each node based on the context in the semantic association graph, and the edge weights of the semantic association graph are updated by using the embedding vector to optimize the semantic association graph, so that the labels with similar semantic relationships have higher connectivity in the graph; An output module for performing semantic matching and relevance ranking by using the optimized semantic association graph according to the retrieval request input by the user, and outputting the file information most relevant to the retrieval request.
5. The system according to claim 4, wherein The conversion module is specifically used for: Semantic parsing is performed on text data through a BERT model to extract text feature vectors. Feature extraction is performed on image data through a convolutional neural network (CNN) to generate image feature vectors. Spectral analysis is performed on audio data through short-time Fourier transform (STFT) and an adaptive filter to extract audio feature vectors. The obtained text feature vectors, image feature vectors, and audio feature vectors are normalized, and a multi-modal deep belief network (DBN) is used for feature fusion. Specifically, the normalized text feature vectors, image feature vectors, and audio feature vectors are input into a multi-layer DBN. Through layer-by-layer pre-training and backpropagation fine-tuning, the deep associations between features of each modality are learned. An attention mechanism is used to weight the fused multi-modal feature vectors to highlight key features and suppress redundant information. The fused multi-modal feature vectors are subjected to unified format conversion. Using high-dimensional feature mapping technology, they are mapped into a unified feature space, and a kernel function is used to enhance the non-linear expression ability of the feature space to ensure the effective representation of different modality features in the unified space. The feature vectors in unified format are integrated into a standardized archive feature dataset.
6. A storage medium, characterized in that, A computer program is stored in the storage medium, where the computer program is configured to execute the method according to any one of claims 1-3 when running.
7. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method according to any one of claims 1-3.
Citation Information
Patent Citations
Regional air quality prediction method based on deep belief network
CN117973583A
Tourism resource hierarchical multi-label classification method and system
CN118312833A