Intelligent file classification and retrieval method and system
Through multimodal data fusion, deep learning classification models and semantic network technology, semantic correlation maps are generated and optimized, and the shortcomings in classification and retrieval in existing archive management technologies are solved, and efficient and intelligent archive management and retrieval are achieved.
Patent Information
- Application Number
- CN202510578193.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-07
AI Technical Summary
In the existing archive management technology, the classification method is highly subjective and difficult to ensure consistency and comprehensiveness. The search method is based on keyword matching and is difficult to capture the user's intentions and semantics, resulting in low accuracy and correlation of the search results.
A multimodal data fusion algorithm is used to convert text, image and audio data in a unified format and extract feature to generate a standardized archival feature data set. Then, the adaptive classification model based on deep learning is classified and processed to generate multi-level archival classification labels. Through enhanced semantic network construction technology, the semantic relationships between classification labels are analyzed, and the semantic correlation maps are generated and optimized for semantic matching and correlation sorting.
It improves the relevance and accuracy of archive retrieval, reduces the burden of manual classification, and enhances the modernization and intelligence level of archive management.
Smart Images

Figure CN120086390A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of file management, and particularly relates to a method and system for intelligent classification and retrieval of files. Background Art
[0002] In the information age, the importance of file management has become increasingly prominent. With the popularization of electronic files, various file information is stored and managed in the form of multi-modal data such as text, images, and audio. The traditional methods of file classification and retrieval have gradually revealed their limitations. Most of the existing classification methods rely on manual annotation and simple keyword matching, which is not only inefficient but also difficult to adapt to the diversity and dynamics of file data, resulting in low accuracy and relevance of retrieval results.
[0003] First of all, the traditional file classification method often relies on manually set classification rules. This method is highly subjective and vulnerable to personal experience, making it difficult to ensure consistency and comprehensiveness. At the same time, problems such as information omission and redundancy in the classification process are common. Especially when faced with large and complex file data, manual processing will greatly increase the workload. Secondly, most of the existing retrieval methods are based on keyword search. The retrieval requests input by users often fail to effectively match the file content. The retrieval intentions, semantics, and context of users are often ignored, often resulting in the inability to retrieve relevant file information in a timely and effective manner. This method not only reduces the user's retrieval experience but also fails to fully explore the potential value of files. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for intelligent classification and retrieval of files to solve the deficiencies in the prior art, improve the relevance and accuracy of retrieval, and promote the modernization and intelligence of file management.
[0005] An embodiment of the present application provides a method for intelligent classification and retrieval of files, and the method includes: According to the original data format of the file, using a multi-modal data fusion algorithm, perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized file feature data set; Based on a deep learning-based adaptive classification model, perform classification processing on the standardized file feature data set to generate multi-level file classification labels; Through enhanced semantic network construction technology, analyze the semantic relationships between file classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph; According to the retrieval request input by the user, use the optimized semantic association graph for semantic matching and relevance ranking, and output the file information most relevant to the retrieval request.
[0006] Optionally, according to the original data format of the file, use a multimodal data fusion algorithm to perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized archive feature dataset, including: Perform semantic parsing on the text data through the BERT model to extract text feature vectors, perform feature extraction on the image data through the convolutional neural network CNN to generate image feature vectors, and perform spectral analysis on the audio data through the short-time Fourier transform STFT and adaptive filters to extract audio feature vectors; Normalize the obtained text feature vectors, image feature vectors, and audio feature vectors, and use a multimodal deep belief network DBN for feature fusion. Among them, input the normalized text feature vectors, image feature vectors, and audio feature vectors into a multi-layer DBN, and through layer-by-layer pre-training and backpropagation fine-tuning, learn the deep associations between features of each modality, and use an attention mechanism to weight the fused multimodal feature vectors to highlight key features and suppress redundant information; Perform unified format conversion on the fused multimodal feature vectors, use high-dimensional feature mapping technology to map them into a unified feature space, and use a kernel function to enhance the non-linear expression ability of the feature space to ensure the effective representation of different modality features in the unified space, and integrate the feature vectors in unified format into the standardized archive feature dataset.
[0007] Optionally, the adaptive classification model based on deep learning is used to classify the standardized archive feature dataset to generate multi-level archive classification labels, including: Construct a hybrid model based on the convolutional neural network CNN and the long short-term memory network LSTM. CNN is used to extract the spatial information of features, and LSTM is used to process the time series dynamic changes of features. Among them, when constructing the hybrid model, design multiple parallel CNN branches, each branch uses different convolutional kernel sizes and strides to extract features from different scales. At the same time, to prevent feature interference, use an attention mechanism to assign different weights to each branch; At the output layer of the model, use a hierarchical Softmax mechanism to construct multi-level classification labels. First, output the classification results of the first layer, and then generate subclass labels for the second layer classification according to the classification results of the first layer. Among them, for each category, by introducing label relationship modeling based on GCN, record the dependency relationships between each label, and use the graph structure information of the labels to enhance the inference ability of the model; Use the constructed hybrid model as an adaptive classification model to infer the standardized archive feature dataset and output multi-level archive classification labels.
[0008] Optionally, the method for analyzing the semantic relationship between archive classification tags through the enhanced semantic network construction technology, generating a semantic association graph, and optimizing the semantic association graph using a graph neural network includes: Perform co-occurrence analysis on the archive classification tags. By calculating the co-occurrence frequency of different classification tags in the same archive, construct a co-occurrence matrix; Based on the co-occurrence matrix, use the non-negative matrix factorization (NMF) technique to decompose the co-occurrence matrix into two matrices W and H, where W represents the latent semantic features of the tags, and H represents the expression of the tags on these features, so as to map the high-dimensional co-occurrence features to a low-dimensional semantic feature space and extract the latent semantic relationships of the tags; According to the latent semantic feature matrix W, use the broadcast mechanism to construct the semantic relationship between the tags to obtain a semantic association graph, where nodes are created for each tag, and the edge weights are defined by the similarity between the tags; Adopt graph embedding technology to generate the embedding vector of each node based on the context within the semantic association graph, and use the embedding vector to update the edge weights of the semantic association graph to optimize the semantic association graph, so that tags with similar semantic relationships have higher connectivity in the graph.
[0009] Optionally, the method for performing semantic matching and relevance ranking using the optimized semantic association graph according to the retrieval request input by the user and outputting the archive information most relevant to the retrieval request includes: For the retrieval request input by the user, first perform multimodal parsing on the retrieval content, and convert the text query, voice input, or image content in the retrieval content into embedding features of different modalities. Among them, if it is a text input query, use a pre-trained language model to parse the retrieval text and extract semantic embedding features to ensure that the extracted semantic embedding features can capture the user's intention; if it is a voice input query, convert the voice content into text through speech recognition and further use the language model to extract semantic embedding features; if it is an image input, use a convolutional neural network to extract image embedding features; Fuse the embedding features of different modalities into a unified retrieval request embedding feature through a multimodal fusion mechanism to map it into a feature space consistent with the archive feature dataset; Query the node closest to the retrieval request in the optimized semantic association graph, use graph embedding technology to generate the semantic embedding of each node, and calculate the similarity score between the retrieval request embedding feature and the node embedding feature in the graph; According to the similarity score, select several initial nodes and their direct adjacent nodes most relevant to the retrieval request as preliminary candidate nodes to obtain a corresponding preliminary candidate archive set; Use a graph neural network to perform message propagation on the preliminary candidate nodes, and propagate the semantic information of adjacent nodes into the retrieval matching nodes to enrich the context representation of the nodes; Optimize the semantic embedding of each node through multiple rounds of propagation iteration to better represent the semantic features in the current retrieval context; Recalculate the similarity score between the propagated node embedding and the retrieval request embedding, sort the preliminary candidate file set, and sequentially output the candidate file information most relevant to the retrieval request.
[0010] Another embodiment of the present application provides an intelligent file classification and retrieval system, and the system includes: A conversion module, configured to perform unified format conversion and feature extraction on text, image, and audio data by using a multimodal data fusion algorithm according to the original data format of the file to obtain a standardized file feature data set; A classification module, configured to perform classification processing on the standardized file feature data set based on an adaptive classification model of deep learning to generate multi-level file classification labels; A generation module, configured to analyze the semantic relationship between file classification labels through an enhanced semantic network construction technology, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph; An output module, configured to perform semantic matching and relevance ranking by using the optimized semantic association graph according to the retrieval request input by the user, and output the file information most relevant to the retrieval request.
[0011] Another embodiment of the present application provides a storage medium, in which a computer program is stored, and wherein the computer program is configured to execute the method described in any one of the above when running.
[0012] Another embodiment of the present application provides an electronic device, including a memory and a processor, a computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any one of the above.
[0013] Compared with the prior art, a method for intelligent classification and retrieval of archives provided by the present invention, according to the original data format of the archives, uses a multi-modal data fusion algorithm to perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized archive feature data set; based on an adaptive classification model of deep learning, classifies the standardized archive feature data set to generate multi-level archive classification labels; through an enhanced semantic network construction technology, analyzes the semantic relationships between the archive classification labels to generate a semantic association map, and optimizes the semantic association map to obtain an optimized semantic association map; according to the retrieval request input by the user, uses the optimized semantic association map to perform semantic matching and relevance ranking, and outputs the archive information most relevant to the retrieval request, thereby being able to improve the relevance and accuracy of the retrieval and promote the modernization and intelligence of archive management. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a hardware structure block diagram of a computer terminal for a method for intelligent classification and retrieval of archives provided by an embodiment of the present invention; Figure 2 It is a flowchart of a method for intelligent classification and retrieval of archives provided by an embodiment of the present invention; Figure 3 It is a structural diagram of a system for intelligent classification and retrieval of archives provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.
[0016] An intelligent electronic archive retrieval system constructed based on natural language processing (NLP) technology and machine learning algorithms is one of the solutions born to solve the above problems. Such systems use advanced NLP technology to parse the text content and extract the key information points therein; at the same time, use powerful machine learning models to classify and cluster these information, so that the computer can "understand" the meaning behind the files like humans. In this way, when the user inputs a query request, the system can quickly identify the most relevant results according to the pre-trained model and return them to the user, greatly improving the speed and quality of the search.
[0017] In addition, with the development of emerging technologies such as deep learning, future intelligent electronic archive retrieval will be even more intelligent. For example, it can continuously learn and optimize its knowledge base to adapt to the changes of specific terms in different fields; or combine functions such as image recognition to achieve cross-media information integration and provide users with a more comprehensive and rich service experience. In short, with the power of natural language processing and machine learning, we are gradually moving towards a more convenient and efficient digital information era.
[0018] An embodiment of the present invention first provides an intelligent file classification and retrieval method, which can be applied to electronic devices, such as computer terminals, specifically ordinary computers, etc.
[0019] The following takes running on a computer terminal as an example to illustrate it in detail. Figure 1 It is a hardware structure block diagram of a computer terminal for an intelligent file classification and retrieval method provided by an embodiment of the present invention. As Figure 1 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.
[0020] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any intelligent file classification and retrieval method.
[0021] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0022] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any intelligent file classification and retrieval method.
[0023] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 1 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0024] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0025] See Figure 2, embodiments of the present invention provide an intelligent file classification and retrieval method, which may include the following steps: S201, according to the original data format of the file, using a multimodal data fusion algorithm, perform unified format conversion and feature extraction on text, image and audio data to obtain a standardized file feature data set; This method uses a multimodal data fusion algorithm to perform unified format conversion and feature extraction on text, image and audio data according to the original data format of the file, forming a standardized file feature data set. This process first identifies the data features of different modalities, such as the semantic information of text content, the visual features of images, and the acoustic features of audio. By comprehensively using deep learning techniques and feature extraction methods, the system integrates multiple data types into a unified feature set, ensuring that different types of data can cooperate efficiently during processing and providing a more comprehensive information basis.
[0026] The implementation of this method significantly improves the intelligent level of file management and retrieval. By effectively fusing and standardizing different types of data, the system not only simplifies the subsequent analysis process of the data, but also greatly improves the processing ability of complex information. This file feature data set in a unified format lays a good foundation for subsequent classification and retrieval, enabling users to obtain relevant information faster and more accurately when searching for and managing files, improving work efficiency and the scientific nature of decision-making.
[0027] Specifically, the semantic analysis of text data can be performed through the BERT model to extract text feature vectors, the feature extraction of image data can be performed through the convolutional neural network CNN to generate image feature vectors, and the spectral analysis of audio data can be performed through the short-time Fourier transform STFT and adaptive filters to extract audio feature vectors; At this stage, the system uses specific technologies for feature extraction for different data modalities. The text data is semantically parsed through the BERT model to extract feature vectors with deep semantic information; the convolutional neural network (CNN) is used for the image data to extract spatial features, reflecting the visual information of the image; the audio data is spectroscopically analyzed through the short-time Fourier transform (STFT) and adaptive filters to obtain the frequency domain features of the audio signal. The extraction of these features not only ensures the accuracy of the information, but also helps with more complex subsequent data processing. This step enables the system to extract valuable information from different types of data, improving the flexibility and comprehensiveness of data processing. By using advanced feature extraction techniques, the features of each modality can be quantified and standardized, thus laying a solid foundation for subsequent feature fusion and intelligent classification, ensuring the accuracy and efficiency of subsequent analysis.
[0028] In this step, feature extraction is first performed on the text data. The preprocessing of the text data is a crucial step, including word segmentation and stop word removal, to ensure that important information is focused on during feature extraction. Next, the system adopts the BERT (Bidirectional Encoder Representations from Transformers) model, which is well-known for its bidirectional context understanding ability. At this stage, the input text is converted into corresponding word vectors, which can capture the deep semantic features of the text. For example, for the sentence "The cat is sitting on the mat", BERT will generate context-related embedding vectors for each word, enabling the relationship between "cat" and "mat" to be reflected semantically.
[0029] After completing the text feature extraction, the system turns to the processing of image data. Images usually contain rich visual information, and it is crucial to retain as much of this information as possible for subsequent classification and retrieval. To this end, a convolutional neural network (CNN) is used to process the image data. During this process, the image undergoes operations in multiple convolutional layers and pooling layers to extract local features and reduce the dimension. In specific operations, the system can adopt classic network architectures such as ResNet or VGG to extract the visual feature vectors of the image. These feature vectors can capture information such as the shape, color, and texture of the image. For example, after processing a photo containing a "cat", the resulting feature vector can present key visual features related to the cat, such as the shape of the ears and the coat color.
[0030] Finally, this step processes the audio data. Audio signals have time-domain and frequency-domain characteristics, so different processing methods are required. The system uses the short-time Fourier transform (STFT) to convert the audio signal into a spectrogram, which can simultaneously display the changes in time and frequency. After STFT processing, the system further applies an adaptive filter to extract the audio feature vectors, such as the pitch, volume, and frequency components of the speech. These feature vectors can reflect important information in the audio data, such as the mood or speech rate of the speaker. When the feature extraction of text, image, and audio data is all completed, the system collates and summarizes the feature vectors of each modality for entry into the subsequent fusion stage.
[0031] The obtained text feature vectors, image feature vectors, and audio feature vectors are normalized, and a multi-modal deep belief network DBN is used for feature fusion. Among them, the normalized text feature vectors, image feature vectors, and audio feature vectors are input into the multi-layer DBN. Through layer-by-layer pre-training and backpropagation fine-tuning, the deep associations between the features of each modality are learned. The attention mechanism is used to weight the fused multi-modal feature vectors to highlight the key features and suppress redundant information; After feature extraction is completed, the system normalizes the text feature vector, image feature vector, and audio feature vector, and then uses a multi-modal deep belief network (DBN) to fuse them. Normalization is to ensure the consistency of different features within the numerical range, facilitating the fusion operation. Through a layer-by-layer pre-training and back fine-tuning mechanism, DBN learns the deep associations between features of each modality, thus fusing out a comprehensive feature vector. In addition, an attention mechanism is adopted to weight the fused multi-modal features, significantly enhancing the importance of key features and reducing the impact of redundant information. The significance of this step is that through feature fusion, the system can form a more comprehensive feature representation, reflecting the multi-dimensional characteristics of the archival information. This fused feature vector can better represent the overall information of the archive, making subsequent classification and retrieval more accurate and efficient. At the same time, the attention mechanism helps the system focus on more important features, optimizing the performance of the model.
[0032] In this stage, the system first normalizes the feature vectors extracted from multi-modalities. The purpose of normalization is to eliminate the differences in the numerical ranges between different modalities, ensuring that each feature has the same weight in the subsequent fusion process. This can be achieved through methods such as max-min normalization or z-score standardization. After normalization, the text feature vector, image feature vector, and audio feature vector will be input into a multi-modal deep belief network (DBN) for fusion. At each layer of the DBN, the network uses unsupervised learning to extract high-level abstract information from the features of each modality, gradually learning the potential associations between features of different modalities.
[0033] During the fusion process, the training of DBN adopts a layer-by-layer pre-training and back fine-tuning mechanism. In the pre-training stage, DBN trains the weights of each layer according to the given features, effectively capturing the relationships between features. In the back fine-tuning stage, the network comprehensively considers the information of all modality features and optimizes the overall network performance through gradient descent. It is worth noting that an attention mechanism is adopted in this process, which dynamically adjusts the weights of different modality features, enabling more important features to occupy a larger proportion in the final fused feature vector. For example, when analyzing a video data, the color features of the image may be more important than the speech features, and the system will strengthen the performance of the image features in the fusion result through the attention mechanism.
[0034] After the fusion is completed, the system will obtain a feature vector in a unified format. This feature vector synthesizes multi-modal information and can comprehensively reflect the characteristics of the archives. This vector will be converted into a standardized format for subsequent processing steps. Finally, this fused feature vector will form a standardized archive feature dataset, providing stable basic data for subsequent classification and retrieval. This feature fusion not only improves the integrity of information but also lays a solid foundation for subsequent intelligent classification and efficient retrieval.
[0035] Convert the fused multi-modal feature vector into a unified format. Adopt high-dimensional feature mapping technology to map it into a unified feature space, and use kernel functions to enhance the non-linear expression ability of the feature space to ensure the effective representation of different modal features in the unified space. Integrate the feature vector in the unified format into the standardized archive feature dataset.
[0036] After the feature fusion is completed, the fused multi-modal feature vector will further undergo a unified format conversion. Adopt high-dimensional feature mapping technology to map it into a unified feature space. Use kernel function technology to enhance the non-linear expression ability of the feature space, enabling different modal features to be effectively represented in the same feature space. This process ensures that all features are compared and processed in the same dimension, thus simplifying subsequent classification and retrieval operations. Through non-linear mapping, the system can overcome the problem of dimensional inconsistency that may occur during the processing of multi-modal features, ensuring that all features can be analyzed in the same context. This unified feature representation form provides a stable input for subsequent classification models, helping to improve the accuracy and efficiency of classification and retrieval.
[0037] In this step, the fused multi-modal feature vector will be converted into a unified format to ensure that all data can be compared in the same feature space. First, the system calls high-dimensional feature mapping technology to map the fused feature vector into a high-dimensional space. Here, kernel functions such as the Gaussian kernel function are usually used to enhance the non-linear expression ability of the feature space. This mapping enables different modal features to be effectively compared and analyzed in the same dimension. This means that non-linear relationships that cannot be fully expressed by traditional linear features can also be distinguished in the new feature space.
[0038] When performing feature mapping, the system needs to consider the distribution of features to ensure that all feature vectors can perform well in the new feature space. Therefore, the system will calculate the distance for each feature vector to confirm the similarity between features. This step can use measurement methods such as Euclidean distance or cosine similarity to help the system evaluate the relationship between different features. For example, the similarity evaluation of text features and image features can reveal the aggregation of the two modalities of "cat" in the feature space, thus better understanding its semantics.
[0039] Finally, the feature vectors after kernel function mapping will be integrated into a standardized archive feature dataset, which is necessary for subsequent classification and retrieval steps. The feature set in a unified format provides clear inputs for the intelligent classification model and ensures that the relevance among all modal data is retained. When a user makes a retrieval request, the system can quickly and accurately provide relevant archive information based on this normalized dataset. This undoubtedly improves the efficiency of data processing and the accuracy of retrieval, making the archive management system more intelligent in operation.
[0040] S202, Use an adaptive classification model based on deep learning to classify the standardized archive feature dataset and generate multi-level archive classification labels; The adaptive classification model based on deep learning aims to classify the standardized archive feature dataset to generate multi-level archive classification labels. The model design combines a Convolutional Neural Network (CNN) and a Long Short-Term Memory Network (LSTM), giving full play to the advantages of these two algorithms. The CNN is responsible for extracting the spatial features of the input data, while the LSTM focuses on the changes in time series data. Such a combination enables the model to fully understand complex archive data. At the same time, to prevent interference between features, an attention mechanism is adopted in the model. This mechanism assigns different weights to different features, thereby emphasizing the influence of important features and ensuring that the generated classification labels are highly accurate and robust.
[0041] By using an adaptive classification model based on deep learning to classify the standardized archive features, not only the classification efficiency is improved, but also the burden of manual classification is reduced. The multi-level archive classification labels can provide a clearer and more accurate basis for subsequent retrieval and analysis, enabling users to find relevant information more quickly when querying archives. This intelligent processing method helps to improve the intelligence level of the archive management system, greatly shortening the time from data acquisition to information extraction and providing a better experience for users.
[0042] Specifically, a hybrid model based on the Convolutional Neural Network CNN and the Long Short-Term Memory Network LSTM can be constructed. The CNN is used to extract the spatial information of features, and the LSTM is used to process the time series dynamic changes of features. When constructing the hybrid model, design multiple parallel CNN branches, each branch using different convolutional kernel sizes and strides to extract features from different scales. At the same time, to prevent features from interfering with each other, an attention mechanism is adopted to assign different weights to each branch; At this stage, a hybrid model is designed by combining the advantages of CNN and LSTM to effectively extract the spatial and temporal information of archive features. CNN uses multiple parallel branches, each branch for different convolutional kernel sizes and strides, and the extracted features include static information such as edges and textures. LSTM is used to capture the dynamic changes in sequential data (such as time series). This combination allows the model to deeply understand the data from multiple dimensions and obtain a more comprehensive feature representation. The design of the hybrid model not only improves the effectiveness and accuracy of feature extraction, but also takes into account the diversity and complexity of the data, making the final generated classification labels more reliable when dealing with real-world data. Such a design provides a strong foundation for subsequent classification and information retrieval, and also enables the intelligent archive system to have higher flexibility and adaptability.
[0043] In the construction of the hybrid model, first select a suitable framework, such as TensorFlow or PyTorch, for model construction and training. The first part of the model uses a convolutional neural network (CNN) to extract spatial features. We design multiple parallel convolutional layers with different sizes of convolutional kernels (such as 3x3, 5x5, 7x7, etc.) to capture multi-scale feature information from the input data. After each convolutional layer, an activation function (such as ReLU) is connected, and then a pooling layer (such as max pooling or average pooling) is used for dimensionality reduction to reduce the size of the feature map while retaining key information. The final feature map is transformed into a one-dimensional vector through a Flatten layer to prepare for the subsequent fully connected layer. In addition, the outputs of multiple convolutional branches are concatenated in the network to form a comprehensive feature representation, laying the foundation for subsequent LSTM processing. To enhance the feature extraction ability, batch normalization technology is adopted to improve the training stability.
[0044] Next, a long short-term memory network (LSTM) layer is constructed to capture the time series information of the input data. LSTM can effectively preserve past information and dynamically adjust the memory, solving the problem of gradient disappearance that may occur in traditional RNNs when dealing with long sequences. In the model, the input to the LSTM layer is the feature vector output from the CNN layer, which is processed by several LSTM units to extract the temporal features in the data. At the same time, to improve the robustness and generalization ability of the model, a Dropout layer can be set to avoid overfitting. Specifically, a certain proportion of nodes can be randomly discarded after each LSTM unit. Through this randomization method, the performance of the model on new data is enhanced. Finally, the features output by the LSTM are concatenated with the features of the CNN part to form an integrated feature vector as the final representation of the model.
[0045] At the end of the model, an attention mechanism is added to improve the model's ability to focus on key information. The attention mechanism assigns different weights to different features, enabling the model to concentrate on important features during inference. Specifically, when implementing, the weighted average method can be used to weight the concatenated feature vectors to generate an optimized feature representation. The calculation of the weights can be achieved through a simple fully connected network that maps the feature vectors and then obtains the weight values through the softmax function. Finally, the optimized feature vectors are fed into the output layer to prepare for subsequent classification tasks. This hybrid model that combines CNN, LSTM, and the attention mechanism can efficiently extract the spatial and temporal features of the input data, laying a good foundation for the classification task.
[0046] At the output layer of the model, a hierarchical Softmax mechanism is adopted to construct multi-level classification labels. First, the classification results of the first layer are output, and then, according to the classification results of the first layer, the subclass labels for the second layer classification are generated correspondingly. Among them, for each category, by introducing label relationship modeling based on GCN, the dependencies between each label are recorded, and the inference ability of the model is enhanced by utilizing the graph structure information of the labels. The main purpose of this stage is to output classification labels layer by layer through the hierarchical Softmax mechanism. First, the model outputs the classification results of the first layer, and then generates the subclass labels for the second layer classification according to the classification results of the first layer. By introducing a graph convolutional network (GCN) to model the dependencies between labels, the inference ability of the model is enhanced. The hierarchical Softmax mechanism allows the model to flexibly manage labels at different levels, thereby achieving more refined classification. This method not only improves the classification accuracy but also enables the hierarchical relationship of the data to be more intuitively reflected during retrieval, ensuring that users can quickly understand and obtain the required information.
[0047] In the design of the model output layer, a hierarchical Softmax mechanism is adopted to construct multi-level classification labels. First, a multi-level label structure needs to be defined to clarify the hierarchical relationship of each layer of labels. For example, the top-level category can be "document type", and the lower-level categories can be subdivided into "report", "invoice", "letter", etc. When specifically implementing, first create a graph structure in the model to represent the relationship between different classification labels. This graph structure is generated by GCN (graph convolutional network) and can capture the dependencies between labels. For each category, record its similarity and relevance with other labels through the graph structure so that this information can be utilized to generate more accurate labels during model inference.
[0048] In actual operation, the implementation of the hierarchical Softmax mechanism involves performing independent Softmax calculations on the labels of each layer. First, the feature vector output by the model is linearly transformed to map to the dimension of the top-level labels, and the softmax function is applied to generate the probability distribution for each top-level label. Then, based on the top-level label with the highest probability, the corresponding subclass labels are further selected, and this process is repeated until the final classification label is generated. This way of deconstructing layer by layer not only effectively reduces the computational burden but also ensures that the hierarchical relationship between the labels is effectively utilized, improving the accuracy and precision of classification.
[0049] By introducing a loss function to guide the training process of the model, the hierarchical Softmax mechanism can optimize the classification effect. During the training process, appropriate loss functions (such as cross-entropy loss) and optimization algorithms (such as Adam or SGD) are selected to adjust the model parameters. In each training iteration, the loss is calculated based on the prediction results of the model and the actual labels, and the parameters of the network are updated according to the gradient descent method. This way of optimizing layer by layer ensures that the model can gradually improve the classification ability of each layer while retaining the overall hierarchical structure, thereby improving the overall classification performance and efficiency.
[0050] The constructed hybrid model is used as an adaptive classification model to infer the standardized archive feature dataset and output multi-level archive classification labels.
[0051] The last step is to apply the constructed hybrid model to the standardized archive feature dataset for actual classification inference and output multi-level archive classification labels. This model can automatically adapt to different data features and generate classification results that meet the user's needs. This inference process marks the transformation of the model from theory to practice, being able to process real data and produce effective classification labels. Through automated classification, the efficiency of archive management can be greatly improved, providing necessary support for subsequent information retrieval and analysis, and thus achieving the goal of intelligent archive management.
[0052] In this step, the previously constructed hybrid model is applied to the standardized archive feature dataset for classification inference. First, the input data is preprocessed, including different modalities such as text, images, and audio, to ensure that it conforms to the input format of the model. After these feature data are converted into a unified format, a standardized feature set is formed, which is convenient for the model to process effectively. Specifically, during the preprocessing process, the text is tokenized and denoised, the images are standardized and scaled, and the audio data needs to be denoised and feature extracted, such as MFCC (Mel Frequency Cepstral Coefficients), etc., to ensure that each data type is in a relatively unified feature space.
[0053] Next, input these preprocessed data into the hybrid model for inference. The model will gradually extract and integrate the input features through the previously trained CNN and LSTM layers. In this process, the CNN is responsible for extracting spatial features, while the LSTM analyzes the time-series features. After feature extraction, combined with the output of the attention mechanism, the model generates a complete feature vector, which will be used for subsequent classification. According to the model's architecture and training effect, the inference process will be automatic. Through the established hierarchical Softmax mechanism, the model can generate corresponding multi-level classification labels to ensure the accuracy and relevance of the output.
[0054] Finally, the inferred classification labels can be used for subsequent archival information processing and retrieval. In practical applications, the user's retrieval request can be directly matched with the model's classification labels to retrieve relevant archival information accordingly. This efficient classification and inference method can greatly improve the efficiency of archival retrieval. Users can quickly locate the required documents based on the multi-level labels. In an example application, in a large archival management system, when the user inputs keywords, the system can immediately present a list of relevant documents according to the inference results, and the user can then make a selection, meeting the immediacy and accuracy requirements for information acquisition. This ability of automatic classification and inference not only improves the intelligent level of archival management but also provides a more user-friendly operation experience.
[0055] S203, through enhanced semantic network construction technology, analyze the semantic relationships between archival classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph; In this stage, first, conduct co-occurrence analysis on the archival classification labels to calculate the simultaneous occurrence frequency of different classification labels in the same archive. By constructing a co-occurrence matrix, represent the relationships between the labels in numerical form, laying the foundation for subsequent semantic extraction. Subsequently, adopt non-negative matrix factorization (NMF) technology to decompose the co-occurrence matrix into two matrices, one representing the latent semantic features of the labels and the other representing the expression of the labels on these features. This process can map the high-dimensional co-occurrence features to a low-dimensional semantic feature space and extract the hidden semantic relationships between the labels. Then, based on the latent semantic feature matrix, use the broadcast mechanism to construct a semantic relationship graph between the labels. By defining nodes and edges, form a complete semantic association graph. Finally, apply graph neural network technology to optimize the graph to improve the connectivity and semantic relevance between the labels.
[0056] Through this process, the generated optimized semantic association graph can effectively reveal the deep semantic relationships between file classification tags, providing a more accurate semantic basis for subsequent classification and retrieval. The optimized graph not only makes the connection between tags closer, but also improves the semantic matching efficiency in the retrieval process, ensuring that users can obtain more relevant information when querying. This directly enhances the intelligence level of the system, improves the user experience, and facilitates quickly finding the required materials in complex file data.
[0057] Specifically, co-occurrence analysis of file classification tags can be performed. By calculating the co-occurrence frequency of different classification tags in the same file, a co-occurrence matrix is constructed. In this stage, first, collect the classification tag data of all files, and then analyze the occurrence of tags in each file. By traversing each file, record the number of times tags appear simultaneously, construct a co-occurrence matrix, with tags as rows and columns, and fill the elements of the matrix with frequency information. For example, if tag "A" and tag "B" co-occur in 10 files, then the value in the corresponding position in the co-occurrence matrix will be updated to 10. Continue this process until all combinations between tags are counted to form a complete co-occurrence matrix, thus laying a foundation for the subsequent non-negative matrix factorization.
[0058] This process provides basic data support for subsequent extraction of potential semantic associations. The co-occurrence matrix effectively shows the relationships between tags, helps identify which tags often co-occur, and further reveals their potential connections and similarities. Through this data-driven method, subsequent analysis and optimization work will be more accurate, laying a good foundation for constructing an effective semantic graph.
[0059] First, collect the classification tag data of all files and organize it into a structured database. The Pandas library in Python can be used to create a data frame, with each file as a row and each tag as a column. Then, by traversing each file, record the number of times each pair of tags coexist in the same file. For each file, read the list of tags it contains and use a double loop to traverse its tags, updating the corresponding position in the co-occurrence matrix. For example, if file A contains tags X and Y, then the connection between X and Y in the co-occurrence matrix will increase by 1.
[0060] After the co-occurrence analysis is completed, the co-occurrence matrix can be normalized. For example, each count value can be divided by the sum of that row to calculate the relative frequency and form a probability distribution, which can eliminate the influence of the number of different tags and make the analysis more fair and effective. At the same time, it is necessary to consider the occurrence frequency of tags to ensure that the co-occurrence matrix not only reflects simple occurrence situations but also can reveal frequently co-occurring tag combinations.
[0061] Finally, the constructed co-occurrence matrix will serve as the basis for subsequent analysis. This matrix can not only clearly display the relationships between tags but also lay a solid data foundation for further potential semantic analysis. By using visualization tools (such as Matplotlib or Seaborn), the co-occurrence matrix can be graphically presented to facilitate the team members' intuitive understanding of the tag relationships.
[0062] Based on the co-occurrence matrix, the non-negative matrix factorization (NMF) technique is used to decompose the co-occurrence matrix into two matrices W and H. Here, W represents the potential semantic features of the tags, and H represents the expression of the tags on these features, so as to map the high-dimensional co-occurrence features to a low-dimensional semantic feature space and extract the potential semantic relationships between the tags. In the process of decomposing the co-occurrence matrix using non-negative matrix factorization (NMF), the goal is first set to decompose the co-occurrence matrix into two non-negative matrices: W and H. The W matrix represents the relationship between the tags and the potential semantic features, while the H matrix represents the specific expression of the tags on these features. In the specific implementation, the NMF algorithm adjusts the values of W and H iteratively so that their product is as close as possible to the original co-occurrence matrix. This process not only retains the original information of the tags but also extracts the potential semantic features, enabling the relationships between the tags to be expressed in a more concise manner.
[0063] Through NMF decomposition, the data dimension can be effectively reduced and the potential semantic features between the tags can be extracted, thereby enhancing the learning ability of the model. This mapping is crucial because it allows the model to focus on the deep relationships between the tags rather than simply the surface co-occurrence situations, thus promoting subsequent semantic association analysis and graph construction. Through this method, it is possible to clearly identify which tags share similar features and establish their semantic positions in the entire dataset.
[0064] Before preparing for Non-Negative Matrix Factorization (NMF), it is first necessary to import relevant numerical calculation libraries such as NumPy, SciPy, or sklearn. Using the co-occurrence matrix as the input, set the parameters of NMF, which include the dimensions of W and H after decomposition, as well as the number of iterations, etc. Enter the execution stage of the NMF algorithm, and use the method of random initialization to generate the initial W and H matrices. In each iteration, the algorithm adjusts the elements of these matrices according to the multiplicative update rule to gradually approximate the original input co-occurrence matrix. This process will continue until the preset convergence criterion is reached. For example, stop the iteration when the reconstruction error is lower than a certain threshold. It is worth noting that NMF itself requires that the elements of the input matrix must be non-negative. To this end, it may be necessary to preprocess the co-occurrence matrix to ensure that all count values are positive. Through appropriate regularization techniques, NMF can be prevented from falling into local optima, ensuring that the final obtained W and H matrices have good generalization ability. At the same time, in order to better understand the potential semantic relationships between labels, each column vector of the W matrix will correspond to a potential semantic feature, while the H matrix can reflect the specific expressions of each label on these features. The final output of NMF will provide the necessary basic data for subsequent construction of the semantic association graph, ensuring the effective extraction of potential relationships between labels. Through subsequent analysis, the feature matrix W can be visualized, for example, using the clustering method, to help intuitively understand the relationship between label groups and semantic features.
[0065] According to the latent semantic feature matrix W, use the broadcasting mechanism to construct the semantic relationship between labels to obtain the semantic association graph. Among them, create nodes for each label and define the edge weights by the similarity between labels; At this stage, use the latent semantic feature matrix W to construct the semantic association graph. First, create a node for each classification label and embed the node into the graph structure. Then, define the weights of the edges by calculating the similarity between labels. Usually, cosine similarity or Euclidean distance can be used to quantify the degree of similarity. Using the broadcasting mechanism, the feature vectors in the matrix W can be extended so that each label node can be paired with other label nodes and calculate the similarity score between them, thus generating a complete semantic association graph to clearly show the relationship between labels.
[0066] This construction process can intuitively express the semantic relationship between labels, making each label node form corresponding connections in the graph. This not only enhances the visualization effect of the data but also lays a foundation for subsequent semantic matching and information retrieval. Through the optimized graph, the intelligent level of the model in processing complex queries can be improved, effectively enhancing the accuracy and efficiency of retrieval.
[0067] When constructing a semantic association graph, it is first necessary to create a node for each tag. In this process, graph theory-related toolkits (such as NetworkX) can be used to construct the graph structure. The creation of each node can be achieved by assigning a unique identifier to the tag while saving the tag information, including its name and relevant data on semantic features. Next, using the data in the latent semantic feature matrix W, the similarity between tags is calculated, usually using cosine similarity for evaluation. After calculating the similarity of all tags, the similarity values can be stored as edge weights. Using the broadcast mechanism, the row vectors of the W matrix can be extended to all combinations of each tag with other tags, thus simplifying the process of similarity calculation. By setting a threshold, the edges between tags with low similarity can be filtered out, and only relatively meaningful connections are retained, thereby ensuring the simplicity and informativeness of the semantic association graph. The semantic association graph will be able to visually display the relationships between various tags, helping to understand the similarities and interconnections between tags. Further, the graph structure can be used for visualization, enabling team members to quickly understand the relationships between tags and providing clues for later use to help optimize the retrieval system.
[0068] Adopt graph embedding technology to generate the embedding vector of each node based on the context within the semantic association graph, and use the embedding vector to update the edge weights of the semantic association graph to optimize the semantic association graph, so that tags with similar semantic relationships have higher connectivity in the graph.
[0069] At this stage, the semantic association graph is further optimized through graph embedding technology. The core idea of graph embedding technology is to map the structural information and attribute information of nodes in the graph to a low-dimensional dense space, making similar node features closer in the embedding space. Specifically, when operating, algorithms (such as DeepWalk or Node2Vec) are used to perform random walks on each node to capture the local relationships between nodes and generate the embedding vector of each node. These embedding vectors not only reflect the features of each tag but also consider its position and context information in the graph. Then, through these embedding vectors, the edge weights in the graph are adjusted and optimized, making tags with similar semantics more connected in the graph.
[0070] Through graph embedding technology, the expressive ability of the semantic association graph can be effectively improved, making tags with strong relevance more connected in the graph. The optimized graph further enhances the inference ability of the model, thus achieving more accurate and efficient semantic matching in the actual retrieval process. Users can more easily find tags related to their input during retrieval, thereby improving the sensitivity and practicality of the entire file retrieval system.
[0071] In this stage, first select a suitable graph embedding technique. Commonly used algorithms include DeepWalk, Node2Vec, or GraphSAGE, etc. These algorithms generate feature vectors of nodes by performing random walks or neighbor sampling on the graph. In actual implementation, define the parameter settings of the random walk, such as the number of walks, the length of the walk, and the step size of each walk. Flexible adjustment of these parameters can better capture the context relationships between nodes. After generating the node feature vectors, next, it is necessary to calculate the similarity between each pair of nodes and apply it to update the edge weights in the graph. For example, cosine similarity or L2 distance can be used to quantify the similarity between nodes. In this process, the update of the edge weights helps to reflect the changes in the current semantic information and enhance the connectivity between nodes with similar semantics. It should be noted that the edge weights may need to be updated iteratively multiple times to ensure that the final graph fully reflects the latest semantic relationships. After completing the graph embedding and edge weight update, the optimized semantic association graph will provide strong support for subsequent retrieval. When the user performs a retrieval, the system can more accurately understand and match the user's query intent and return the most relevant tag and profile information. Finally, the optimized graph not only improves the intelligence level of the system but also greatly enhances the user experience, making the information retrieval process more efficient and accurate.
[0072] S204, according to the retrieval request input by the user, use the optimized semantic association graph for semantic matching and relevance ranking, and output the profile information most relevant to the retrieval request.
[0073] According to the retrieval request input by the user, use the optimized semantic association graph for semantic matching and relevance ranking. First, perform multimodal parsing on the retrieval content, and convert the input from different formats such as text, voice, or image into corresponding embedding features. For text input, the system will use a pre-trained language model (such as BERT) to parse the text and extract semantic embedding features from it to ensure that the user's query intent can be accurately captured. When processing voice input, the system first converts the voice content into text through speech recognition technology and then uses the language model to extract semantic embedding features. For image input, convolutional neural networks are used to extract image embedding features. After processing all the embedding features of different modalities, the system merges these features into a unified retrieval request embedding feature through a multimodal fusion mechanism to facilitate querying in the feature space consistent with the profile feature dataset. Then, the system queries the nodes closest to the retrieval request in the optimized semantic association graph and uses graph embedding technology to generate the semantic embeddings of each node. The system calculates the similarity scores between the retrieval request embedding feature and the node embedding features in the graph, and selects several initial nodes and their direct adjacent nodes most relevant to the retrieval request as preliminary candidate nodes, and then obtains the corresponding preliminary candidate profile set.
[0074] The significance of this process lies in that through multimodal parsing and an optimized semantic association graph, the system can more accurately understand the user's retrieval intent. It is no longer limited to simple keyword matching but delves into the deep meaning of the user's query. This semantic matching ability enables the system to quickly discover the information most relevant to the user's needs in complex archival data, significantly improving the accuracy and efficiency of retrieval. Additionally, through the use of graph embedding technology, the system can effectively utilize the correlation between tags, associate the user's retrieval request with potentially similar archives, and provide richer and more accurate retrieval results, thereby enhancing the user experience.
[0075] Specifically, for the retrieval request input by the user, first, perform multimodal parsing on the retrieval content, converting the text query, voice input, or image content in the retrieval content into embedding features of different modalities. Among them, if it is a text input query, use a pre-trained language model to parse the retrieval text and extract semantic embedding features to ensure that the extracted semantic embedding features can capture the user's intent; if it is a voice input query, convert the voice content into text through speech recognition and further use the language model to extract semantic embedding features; if it is an image input, use a convolutional neural network to extract image embedding features. At this stage, the system needs to perform multimodal parsing on the retrieval content input by the user, identify the input format, and extract the corresponding features. For a text query, the system usually uses a pre-trained language model (such as BERT) to perform semantic parsing on the input text and extract the semantic embedding features of the text to accurately capture the user's intent. When processing voice input, first convert the voice information into text through speech recognition technology, and then use the text parsing model to extract its semantic embedding features. For image input, the system uses a convolutional neural network to extract the feature vectors of the image, and these feature vectors can reflect the semantic information of the image content. Through this process, the system can uniformly convert user inputs from different modalities into available feature representations, laying a foundation for subsequent semantic matching. This processing method greatly enhances the flexibility and adaptability of the system, can meet the scenarios where users express their needs through different media, and improves the accuracy and comprehensiveness of retrieval.
[0076] In this step, the system first needs to identify the type of the retrieval content input by the user and determine whether it is text, voice, or image. For text input, the system will call a pre-trained language model (such as BERT or GPT, etc.) and use natural language processing (NLP) technology to analyze the text. This includes word segmentation, stop word removal, word vector representation, etc., converting the text into computable semantic embedding features. Each word is mapped to a high-dimensional vector, and its context information is also taken into account, so as to capture the intent and nuances of the user's query to the greatest extent.
[0077] For voice input, the system will first use Automatic Speech Recognition (ASR) technology to convert the audio signal into text in real time. During this process, the system needs to ensure the recognition accuracy and use deep learning models (such as the sequence-to-sequence model with CTC loss function) to handle different accents and noise interferences. The converted text will be processed by the same NLP technology to generate semantic embedding features. In addition, the system also needs to consider the processing delay to ensure a quick response after the user input, thus enhancing the user experience.
[0078] When processing image input, the system will use Convolutional Neural Network (CNN) to extract features from the image. Specifically, after the image is uploaded, the system will first perform preprocessing, including image scaling, normalization, and enhancement, etc., to ensure that the model can better understand the image content. Then, the extracted feature vectors will be transformed into low-dimensional embeddings to reflect the key information of the image. After these processes are completed, the system will integrate the embedding features of the three modalities for subsequent use.
[0079] Fuse the embedding features of different modalities into a unified retrieval request embedding feature through a multimodal fusion mechanism, so as to map it into a feature space consistent with the archive feature dataset; Multimodal fusion is an important process of integrating features from different sources into a unified representation. In this step, the system will effectively combine the embedding features of text, voice, and image extracted in Step 1 through a multimodal fusion mechanism to generate a unified retrieval request embedding feature. This process requires designing appropriate fusion strategies, such as weighted average method, concatenation method, or other deep learning fusion models, to ensure that the features of each modality can complement each other and enhance the overall expression ability. By fusing the embedding features of different modalities, the system can comprehensively capture the user's query intention and reduce the information loss caused by insufficient single-modal information. The fused feature vector can better reflect the user's needs, improve the understanding of complex queries, and thus enhance the relevance and accuracy of the retrieval.
[0080] At this stage, in order to achieve effective fusion of multimodal features, the system designs a modular fusion mechanism. First, the system will introduce a data preprocessing step to normalize the features of different modalities to ensure that the features of each modality are fused on the same scale. Then, the system will assign different weights to each modality according to its characteristics. The weights may be set based on previous experimental results or optimized through a dynamic learning strategy. Through these steps, the system ensures that each modality occupies an appropriate proportion in the final feature representation.
[0081] Then, the system can adopt various fusion strategies, such as weighted average, concatenation, or more complex deep learning integration methods. Taking weighted average as an example, the system multiplies the features of each modality by their weights and then sums them up to generate a unified retrieval request embedding feature. If the concatenation method is used, the system concatenates the feature vectors of each modality by dimension to form a higher-dimensional feature representation. Through such strategies, the system can make full use of the advantages of different modalities, organically combine various types of information, and improve the accuracy of subsequent matching.
[0082] Finally, the unified retrieval request embedding feature after fusion will be mapped to the feature space consistent with the archive feature dataset. This mapping can be achieved through various methods such as linear transformation or non-linear transformation, ensuring that the fused features can be effectively compared with the features of the existing archive data. This process not only lays a solid foundation for subsequent similarity calculation but also ensures the sensitivity and adaptability of the system to multi-modal queries.
[0083] Query the node in the optimized semantic association graph that is closest to the retrieval request, use graph embedding technology to generate the semantic embedding of each node, and calculate the similarity score between the retrieval request embedding feature and the node embedding feature in the graph; Once the unified retrieval request embedding feature is obtained, the system will search in the optimized semantic association graph to find the node closest to this embedding feature. In this process, the system uses graph embedding technology to generate the semantic embedding of each node to capture the context relationship of the node in the graph. Then, by calculating the similarity score between the retrieval request embedding feature and the embedding features of each node in the graph, the nodes with higher similarity are found. In this way, the system can accurately identify the archive information most relevant to the user's retrieval request, and then provide higher-quality search results. This process uses the semantic association graph to strengthen the semantic understanding ability of the retrieval through similarity matching, ensuring that users can quickly obtain the relevant materials they need.
[0084] In this step, the system first needs to construct and optimize the semantic association graph to ensure that the nodes and edges in the graph can accurately express the relationships between each archive. The system will use a graph database (such as Neo4j) to design the attributes of the nodes and the weights of the edges, enabling the graph to be dynamically updated and reflect the latest semantic relationships. Then, the retrieval request embedding feature is used as a query vector to match the node features in the semantic association graph.
[0085] Next, the system will generate semantic embeddings for each node through graph embedding techniques. This process typically uses graph neural networks (GNNs) or node embedding algorithms (such as Node2Vec or GraphSAGE) to dynamically generate vector representations of nodes. The semantic embedding of each node will consider the feature information of its surrounding adjacent nodes to more comprehensively reflect the semantic context of the node in the graph. The system will continuously update the node features during this process to ensure that each node in the knowledge graph has the latest context semantic information.
[0086] Finally, the system will calculate the similarity scores between the retrieval request embedding features and the embedding features of each node. Similarity calculation can be performed through methods such as cosine similarity, Euclidean distance, or Manhattan distance. The system will filter out the nodes closest to the user's retrieval request according to the similarity scores to ensure that the subsequent candidate file set has the highest relevance. This process not only improves the accuracy of retrieval but also lays the foundation for the subsequent selection of candidate nodes.
[0087] According to the similarity scores, several initial nodes that are most relevant to the retrieval request and their direct adjacent nodes are selected as preliminary candidate nodes to obtain the corresponding preliminary candidate file set; After calculating the similarity scores between the retrieval request embedding features and the node embedding features in the graph, the system will select several initial nodes that are most relevant to the retrieval request based on these scores. The initial nodes include not only the nodes with the highest similarity but also their direct adjacent nodes in the graph to expand the scope of retrieval. This strategy enables the system to further enrich the number of candidate nodes based on the relevance ranking to ensure the comprehensiveness and diversity of the provided file information. Through this method, the system provides a set of preliminary candidate files, ensuring that the retrieval results cover both relevance and diversity, increasing the user's chance of obtaining information. At the same time, this method also reduces the risk of information bias caused by a single node, making the system's recommendation results more reliable.
[0088] In this step, the system first sorts all nodes according to the similarity scores calculated in the previous step and selects several initial nodes with the highest scores. These nodes represent the file information that is semantically closest to the user's query request. Then, the system will search for the adjacent nodes directly connected to these initial nodes according to the topological structure of the initial nodes. These adjacent nodes may also contain information related to the user's request, so their addition can enrich the diversity of candidate files.
[0089] Next, the system will integrate the selected initial nodes and their adjacent nodes to create a candidate profile set. To ensure the quality of the candidate set, the system will also set thresholds. For example, the similarity score must be higher than a certain set value to be included in the candidate set. This process will ensure that the generated candidate profile set is guaranteed in terms of relevance and avoid redundant information caused by the increase in the number of candidate nodes.
[0090] Finally, the system will conduct a preliminary search on the obtained candidate profile set and generate a list for subsequent information display. This candidate set not only includes the nodes with the highest scores but also integrates the information of adjacent nodes with strong relevance so that the system can comprehensively cover the user's search needs. This step provides a solid foundation for subsequent more in-depth searches and information presentation.
[0091] Use a graph neural network to perform message propagation on the preliminary candidate nodes, and spread the semantic information of adjacent nodes into the retrieval matching nodes to enrich the context representation of the nodes; After obtaining the preliminary candidate nodes, the system will use a graph neural network (GNN) for message propagation. Through the architecture of the GNN, nodes can communicate with each other, and the semantic features of adjacent nodes will be propagated to the target retrieval matching nodes. This process is very crucial because it can significantly enhance the context representation of the nodes, making the information of the nodes in the retrieval context more abundant. Through this message propagation mechanism, the system can better capture the deep semantic relationships between nodes, thereby improving the relevance and accuracy of the retrieval results. This dynamic information update also enables the system to adapt to different context changes and improve the flexibility and intelligence level of the retrieval.
[0092] At this stage, the system inputs the preliminary candidate nodes into a graph neural network (GNN) and uses its powerful message propagation ability to enhance the feature representation of each node. First, the system initializes the features of each candidate node and prepares for multiple rounds of message propagation in the GNN. In each round, the candidate nodes receive information from their adjacent nodes and then aggregate this information with their own features to enhance their context representation.
[0093] The message propagation mechanism of the GNN usually includes two main steps: neighborhood aggregation and feature update. In the neighborhood aggregation stage, the system summarizes the features of adjacent nodes according to the connection relationship of the nodes, possibly by means of averaging, summing, or weighted summing, etc. Then, the system combines the aggregated features of adjacent nodes with the features of the current node to update the feature representation of the node. This process is completed through multiple rounds of iteration, and the system continuously improves the context information of the nodes in each round, making the final node representation more abundant and accurate.
[0094] Finally, the system will normalize the node embeddings after multiple rounds of message propagation for subsequent similarity calculation. This process will ensure that the features of each node have a unified scale, facilitating effective comparison with retrieval requests. Through this message propagation mechanism, the system can better capture the complex relationships between nodes and extend the context of information to a broader semantic level, significantly enhancing the retrieval relevance of candidate archives.
[0095] Optimize the semantic embedding of each node through multiple rounds of propagation iteration to better represent its semantic features in the current retrieval context; After the initial message propagation is completed, the system will continue with multiple rounds of propagation iteration to further optimize the semantic embedding of each node. Each iteration will utilize the updated features of previous nodes and the feature information of adjacent nodes to continuously improve the context representation of the nodes. This process will ensure that the final node features not only reflect their own characteristics but also fully consider the relevant context information. Through multiple rounds of propagation, the system can more accurately capture the implicit semantics in the retrieval request, improving the quality of retrieval results. This iterative optimization method enables the nodes to adapt to different retrieval contexts, thus providing users with more relevant archive information.
[0096] In this step, the system will perform multiple rounds of iteration to further optimize the semantic embedding of each node. In each iteration, the feature vector of the node will be updated. The system will set an upper limit on the number of iterations (such as 3 to 5 times) and will also monitor the convergence of the node features. If the change in the node features in a certain iteration is less than the set threshold, the system can stop further iteration to avoid unnecessary computational overhead.
[0097] During the iteration process, the system will use the same message propagation mechanism to update the node features. In each iteration, the node will not only collect the features from adjacent nodes but also adjust the aggregation and update strategies to better adapt to the current retrieval context. For example, the system may introduce a context-sensitive weight adjustment mechanism to dynamically adjust the influence of adjacent nodes according to the features of the current retrieval request, thus ensuring that the most relevant information can be better incorporated into the context representation of the current node.
[0098] Finally, the system will ensure that the feature representation of each node is significantly improved after multiple rounds of iteration, forming an embedding that can fully reflect its semantic information in the current context. When the iteration ends, the system will store the optimized node embeddings for subsequent similarity calculation and candidate archive matching. This process not only enhances the context expression ability of the nodes but also lays a solid foundation for the final retrieval results.
[0099] Calculate the similarity score between the propagated node embeddings and the retrieval request embeddings again, sort the preliminary candidate file set, and sequentially output the candidate file information most relevant to the retrieval request.
[0100] After multiple rounds of propagation and feature optimization, the system compares the final node embeddings with the initially generated retrieval request embeddings and recalculates the similarity score. Using the same similarity measurement method (such as cosine similarity), it re-evaluates the relevance of each candidate node in the current retrieval context. Finally, the system sorts the preliminary candidate file set based on these similarity scores and sequentially outputs the candidate file information most relevant to the retrieval request. This process ensures that the search results finally returned to the user are based on the latest and most relevant context information, thereby improving the user's retrieval satisfaction and information acquisition efficiency. Through such platform-based content analysis and dynamic matching, the information retrieval system can meet the complex and changing user needs.
[0101] In the last step, the system compares the node vectors after multiple rounds of message propagation and semantic embedding optimization with the retrieval request embedding features. To calculate the similarity score, the system will use the same similarity measurement method as before, such as cosine similarity, Euclidean distance, etc. By comparing the retrieval request embedding with the embedding features of each candidate node, the system will generate a new list of similarity scores to determine which candidate file information is most relevant to the user's retrieval request.
[0102] Next, the system will sort the obtained similarity scores and rank the candidate nodes from high to low. To enhance the user experience, the system will set a threshold to ensure that only candidate files with scores exceeding a certain standard are returned, which can avoid returning results that are not relevant enough to the request. After sorting, the system will generate a clear list of candidate files, and these files will be ready to be presented to the user.
[0103] Finally, the system will sequentially output the candidate file information most relevant to the retrieval request according to the sorting result, including the key content, creation time, abstract, and other relevant metadata of the file. This process not only ensures that users can quickly find the information they need, but also provides rich background information to help users understand and select the most suitable file. At the same time, the system will record the user feedback of each retrieval for future further optimization and upgrade. This series of steps ensures the accuracy of information retrieval and user satisfaction.
[0104] It can be seen that according to the original data format of the archives, using the multimodal data fusion algorithm, the text, image, and audio data are subjected to unified format conversion and feature extraction to obtain a standardized archive feature dataset; based on the adaptive classification model of deep learning, the standardized archive feature dataset is classified to generate multi-level archive classification labels; through the enhanced semantic network construction technology, the semantic relationships between the archive classification labels are analyzed to generate a semantic association graph, and the semantic association graph is optimized to obtain an optimized semantic association graph; according to the retrieval request input by the user, the optimized semantic association graph is used for semantic matching and relevance ranking, and the archive information most relevant to the retrieval request is output, thereby being able to improve the relevance and accuracy of the retrieval and promote the modernization and intelligentization of archive management.
[0105] Another embodiment of the present invention provides an intelligent archive classification and retrieval system. Refer to Figure 3 , the system may include: A conversion module 301, configured to perform unified format conversion and feature extraction on text, image, and audio data according to the original data format of the archives by using a multimodal data fusion algorithm to obtain a standardized archive feature dataset; A classification module 302, configured to classify the standardized archive feature dataset based on an adaptive classification model of deep learning to generate multi-level archive classification labels; A generation module 303, configured to analyze the semantic relationships between the archive classification labels by using an enhanced semantic network construction technology to generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph; An output module 304, configured to perform semantic matching and relevance ranking according to the retrieval request input by the user by using the optimized semantic association graph, and output the archive information most relevant to the retrieval request.
[0106] It can be seen that according to the original data format of the archives, using the multimodal data fusion algorithm, the text, image, and audio data are subjected to unified format conversion and feature extraction to obtain a standardized archive feature dataset; based on the adaptive classification model of deep learning, the standardized archive feature dataset is classified to generate multi-level archive classification labels; through the enhanced semantic network construction technology, the semantic relationships between the archive classification labels are analyzed to generate a semantic association graph, and the semantic association graph is optimized to obtain an optimized semantic association graph; according to the retrieval request input by the user, the optimized semantic association graph is used for semantic matching and relevance ranking, and the archive information most relevant to the retrieval request is output, thereby being able to improve the relevance and accuracy of the retrieval and promote the modernization and intelligentization of archive management.
[0107] An embodiment of the present invention also provides a storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0108] Specifically, in this embodiment, the above storage medium can be configured to store a computer program for executing the following steps: S201, according to the original data format of the file, using a multi-modal data fusion algorithm, perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized file feature data set; S202, based on an adaptive classification model of deep learning, perform classification processing on the standardized file feature data set to generate multi-level file classification labels; S203, through an enhanced semantic network construction technology, analyze the semantic relationships between file classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph; S204, according to the retrieval request input by the user, use the optimized semantic association graph to perform semantic matching and relevance ranking, and output the file information most relevant to the retrieval request.
[0109] It can be seen that according to the original data format of the file, using a multi-modal data fusion algorithm, perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized file feature data set; based on an adaptive classification model of deep learning, perform classification processing on the standardized file feature data set to generate multi-level file classification labels; through an enhanced semantic network construction technology, analyze the semantic relationships between file classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph; according to the retrieval request input by the user, use the optimized semantic association graph to perform semantic matching and relevance ranking, and output the file information most relevant to the retrieval request, thereby being able to improve the relevance and accuracy of retrieval and promote the modernization and intelligence of file management.
[0110] An embodiment of the present invention also provides an electronic device, including a memory and a processor, a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0111] Specifically, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0112] Specifically, in this embodiment, the above processor can be configured to execute the following steps through a computer program: S201. According to the original data format of the archives, use the multi-modal data fusion algorithm to perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized archive feature dataset; S202. Based on the deep learning-based adaptive classification model, classify the standardized archive feature dataset to generate multi-level archive classification labels; S203. Through the enhanced semantic network construction technology, analyze the semantic relationships between the archive classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph; S204. According to the retrieval request input by the user, use the optimized semantic association graph for semantic matching and relevance ranking, and output the archive information most relevant to the retrieval request.
[0113] It can be seen that according to the original data format of the archives, use the multi-modal data fusion algorithm to perform unified format conversion and feature extraction on text, image, and audio data to obtain a standardized archive feature dataset; based on the deep learning-based adaptive classification model, classify the standardized archive feature dataset to generate multi-level archive classification labels; through the enhanced semantic network construction technology, analyze the semantic relationships between the archive classification labels, generate a semantic association graph, and optimize the semantic association graph to obtain an optimized semantic association graph; according to the retrieval request input by the user, use the optimized semantic association graph for semantic matching and relevance ranking, and output the archive information most relevant to the retrieval request, so as to improve the relevance and accuracy of retrieval and promote the modernization and intelligentization of archive management.
[0114] The above has detailed the structure, features, and function effects of the present invention according to the illustrated embodiments. The above is only a preferred embodiment of the present invention, but the present invention is not limited to the scope shown in the drawings. Any changes made according to the concept of the present invention, or modified into equivalent embodiments with equivalent changes, still within the spirit covered by the specification and the drawings, shall be within the protection scope of the present invention.
Claims
1. A method for intelligent classification and retrieval of archives, characterized in that: The method comprises: According to the original data format of the archives, a multimodal data fusion algorithm is used to convert the text, image and audio data into a unified format and extract features to obtain a standardized archive feature data set. Based on the deep learning adaptive classification model, the standardized archival feature data set is classified and processed to generate multi-level archival classification labels; Through enhanced semantic network construction technology, the semantic relationship between archive classification labels is analyzed to generate a semantic association map, and the semantic association map is optimized to obtain an optimized semantic association map; According to the search request input by the user, the optimized semantic association graph is used to perform semantic matching and relevance sorting, and the archive information most relevant to the search request is output.
2. The method according to claim 1, characterized in that According to the original data format of the archive, a multimodal data fusion algorithm is used to perform unified format conversion and feature extraction on text, image and audio data to obtain a standardized archive feature data set, including: The BERT model is used to perform semantic analysis on text data and extract text feature vectors. The convolutional neural network (CNN) is used to extract features from image data and generate image feature vectors. The short-time Fourier transform (STFT) and adaptive filter are used to perform spectrum analysis on audio data and extract audio feature vectors. The obtained text feature vector, image feature vector and audio feature vector are normalized, and feature fusion is performed using a multimodal deep belief network (DBN), wherein the normalized text feature vector, image feature vector and audio feature vector are input into a multi-layer DBN, and the deep associations between the features of each modality are learned through layer-by-layer pre-training and reverse fine-tuning. The fused multimodal feature vector is weighted using an attention mechanism to highlight key features and suppress redundant information. The fused multimodal feature vectors are converted into a unified format and mapped into a unified feature space using high-dimensional feature mapping technology. The kernel function is used to enhance the nonlinear expression ability of the feature space to ensure the effective representation of different modal features in the unified space. The feature vectors in a unified format are integrated into a standardized archival feature dataset.
3. The method according to claim 2, characterized in that The deep learning-based adaptive classification model performs classification processing on the standardized archive feature data set to generate multi-level archive classification labels, including: A hybrid model based on convolutional neural network (CNN) and long short-term memory (LSTM) is constructed. CNN is used to extract spatial information of features, and LSTM is used to process the dynamic changes of time series of features. When constructing the hybrid model, multiple parallel CNN branches are designed, and each branch uses a different convolution kernel size and stride to extract features from different scales. At the same time, in order to prevent features from interfering with each other, an attention mechanism is used to assign different weights to each branch. In the output layer of the model, a hierarchical Softmax mechanism is used to construct multi-level classification labels. First, the classification results of the first layer are output, and then the sub-class labels of the second layer are generated according to the classification results of the first layer. For each category, the dependency relationship between each label is recorded by introducing GCN-based label relationship modeling, and the graph structure information of the label is used to enhance the reasoning ability of the model. The constructed hybrid model is used as an adaptive classification model to infer the standardized archival feature dataset and output multi-level archival classification labels.
4. The method according to claim 3, characterized in that The enhanced semantic network construction technology is used to analyze the semantic relationship between the archive classification labels, generate a semantic association map, and optimize the semantic association map using a graph neural network to obtain an optimized semantic association map, including: Conduct label co-occurrence analysis on archive classification labels, and construct a co-occurrence matrix by calculating the co-occurrence frequency of different classification labels in the same archive; Based on the co-occurrence matrix, the non-negative matrix factorization (NMF) technology is used to decompose the co-occurrence matrix into two matrices W and H, where W represents the potential semantic features of the label and H represents the expression of the label on these features, so as to map the high-dimensional co-occurrence features into a low-dimensional semantic feature space and extract the potential semantic relationship of the label; According to the latent semantic feature matrix W, the semantic relationship between tags is constructed using the broadcast mechanism to obtain a semantic association graph, in which a node is created for each tag and the edge weight is defined by the similarity between tags; Using graph embedding technology, an embedding vector of each node is generated based on the context in the semantic association graph, and the embedding vector is used to update the edge weight of the semantic association graph to optimize the semantic association graph so that labels with similar semantic relationships have higher connectivity in the graph.
5. The method according to claim 4, characterized in that According to the search request input by the user, the optimized semantic association graph is used to perform semantic matching and relevance sorting, and the archive information most relevant to the search request is output, including: For the search request input by the user, the search content is first parsed in a multimodal manner, and the text query, voice input or image content in the search content is converted into embedding features of different modalities. If it is a text input query, the search text is parsed using a pre-trained language model to extract semantic embedding features to ensure that the extracted semantic embedding features can capture the user's intention; if it is a voice input query, the voice content is converted into text through voice recognition, and the language model is further used to extract semantic embedding features; if it is an image input, the image embedding features are extracted using a convolutional neural network; The embedded features of different modalities are fused into a unified retrieval request embedding feature through a multimodal fusion mechanism to map it into a feature space consistent with the archive feature dataset; In the optimized semantic association graph, search for the node closest to the search request, generate the semantic embedding of each node using graph embedding technology, and calculate the similarity score between the embedding features of the search request and the embedding features of the nodes in the graph; According to the similarity scores, several initial nodes and their directly adjacent nodes that are most relevant to the search request are selected as preliminary candidate nodes to obtain a corresponding preliminary candidate archive set; Use graph neural networks to propagate messages on preliminary candidate nodes and propagate the semantic information of adjacent nodes to the retrieved matching nodes to enrich the contextual representation of the nodes. The semantic embedding of each node is iteratively optimized through multiple rounds of propagation to better represent the semantic features in the current retrieval context; The similarity score of the propagated node embedding and the retrieval request embedding is calculated again, and the preliminary candidate profile set is sorted, and the candidate profile information most relevant to the retrieval request is output in sequence.
6. An intelligent archive classification and retrieval system, characterized in that: The system comprises: The conversion module is used to convert the text, image and audio data into a unified format and extract features based on the original data format of the archives using a multimodal data fusion algorithm to obtain a standardized archive feature data set; The classification module is used to classify the standardized archival feature data set based on the adaptive classification model of deep learning and generate multi-level archival classification labels; A generation module is used to analyze the semantic relationship between archive classification labels through enhanced semantic network construction technology, generate a semantic association map, and optimize the semantic association map to obtain an optimized semantic association map; The output module is used to perform semantic matching and relevance sorting based on the search request input by the user using the optimized semantic association graph, and output the archive information most relevant to the search request.
7. The system according to claim 6, characterized in that The conversion module is specifically used for: The BERT model is used to perform semantic analysis on text data and extract text feature vectors. The convolutional neural network (CNN) is used to extract features from image data and generate image feature vectors. The short-time Fourier transform (STFT) and adaptive filter are used to perform spectrum analysis on audio data and extract audio feature vectors. The obtained text feature vector, image feature vector and audio feature vector are normalized, and feature fusion is performed using a multimodal deep belief network (DBN), wherein the normalized text feature vector, image feature vector and audio feature vector are input into a multi-layer DBN, and the deep associations between the features of each modality are learned through layer-by-layer pre-training and reverse fine-tuning. The fused multimodal feature vector is weighted using an attention mechanism to highlight key features and suppress redundant information. The fused multimodal feature vectors are converted into a unified format and mapped into a unified feature space using high-dimensional feature mapping technology. The kernel function is used to enhance the nonlinear expression ability of the feature space to ensure the effective representation of different modal features in the unified space. The feature vectors in a unified format are integrated into a standardized archival feature dataset.
8. The system according to claim 7, characterized in that The classification module is specifically used for: A hybrid model based on convolutional neural network (CNN) and long short-term memory (LSTM) is constructed. CNN is used to extract spatial information of features, and LSTM is used to process the dynamic changes of time series of features. When constructing the hybrid model, multiple parallel CNN branches are designed, and each branch uses a different convolution kernel size and stride to extract features from different scales. At the same time, in order to prevent features from interfering with each other, an attention mechanism is used to assign different weights to each branch. In the output layer of the model, a hierarchical Softmax mechanism is used to construct multi-level classification labels. First, the classification results of the first layer are output, and then the sub-class labels of the second layer are generated according to the classification results of the first layer. For each category, the dependency relationship between each label is recorded by introducing GCN-based label relationship modeling, and the graph structure information of the label is used to enhance the reasoning ability of the model. The constructed hybrid model is used as an adaptive classification model to infer the standardized archival feature dataset and output multi-level archival classification labels.
9. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 5 when executed.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-scale structure and feature fusion gearbox intelligent fault diagnosis method
CN112116029A
Regional air quality prediction method based on deep belief network
CN117973583A
Tourism resource hierarchical multi-label classification method and system
CN118312833A
Enterprise portrait label intelligent generation method
CN119168504A
Archive data retrieval method, system and device
CN119271630A
Cited By
Enterprise information generation and retrieval method based on AI and knowledge graph
CN120578742A
Digital archive automatic quality inspection method and system based on image recognition
CN120635922A
Industrial document intelligent retrieval method and system based on large model
CN120744081A
Long text classification method based on multi-modal knowledge enhancement
CN120873995A
A Long Text Classification Method Based on Multimodal Knowledge Enhancement
CN120873995B