Archive management method based on artificial intelligence
Through the archive management method based on artificial intelligence, the problem that traditional systems are difficult to deal with multimodal data is solved, and the unified feature representation of multimodal data and the rapid retrieval of deep learning models is realized, which improves the efficiency and accuracy of archive management.
Patent Information
- Application Number
- CN202510270966.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
Traditional archive management systems are difficult to effectively process multimodal data, including audio, image and text data, resulting in data heterogeneity, difficulty in extracting features and low retrieval efficiency.
Using an artificial intelligence-based archive management method, the archival data of different modalities is converted into a unified multimodal feature representation through preprocessing, feature extraction and standardization, and trained using deep learning models to achieve fast indexing and accurate retrieval.
The heterogeneity elimination and feature extraction of multimodal data are realized, which improves the overall expression ability of the data. Users can achieve accurate retrieval through voice query, which improves the efficiency and accuracy of the retrieval.
Smart Images

Figure CN120196801A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of file management, and specifically provides a file management method based on artificial intelligence. Background Art
[0002] With the rapid development of information technology, the types of file data have become increasingly rich, no longer limited to traditional text data, but covering various modalities such as audio and images. The emergence of such multimodal data has brought unprecedented challenges to the storage, management, and retrieval of files. Most traditional file management systems are based on single text data and conduct retrieval through keyword matching, which can no longer meet the current complex and changing requirements for file data processing.
[0003] Audio data can record sound information such as speech and music, has time-domain and frequency-domain characteristics, and can convey rich semantic and emotional information. Image data intuitively shows the appearance and form of things through features such as color, texture, and shape. Text data contains explicit semantic information, which is convenient for understanding and analysis. However, the following problems exist in the storage and retrieval of these different types of file data: data heterogeneity: different modalities of data have differences in format, structure, and features, making it difficult to uniformly process; feature extraction difficulty: how to extract effective and stable features from complex audio, image, and text data is a technical problem; low retrieval efficiency: traditional keyword-based retrieval methods have poor effects when dealing with multimodal data and are difficult to quickly and accurately find the target file. Therefore, it has important practical significance and application value to propose a technical solution for accurate retrieval, and thus a file management method based on artificial intelligence is proposed. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a file management method based on artificial intelligence, which has the advantages of fast indexing and accurate detection, and solves the problem that the detection of files in the prior art is relatively cumbersome.
[0005] To achieve the above object, the present invention provides the following technical solution: A file management method based on artificial intelligence, including the following steps: S1: Collect various types of file data, including audio data, image data, and text data; S2: Preprocess the collected file data, extract features from the preprocessed file data, standardize the extracted features, and perform multimodal fusion on the standardized features to obtain multimodal feature F fused ; S3: The multimodal feature F containing the time-domain feature F time and the frequency-domain feature F frep in the audio data fusedLabeled as R, the image data contains the color feature F color , the texture feature F texture and the shape feature F shape of the multimodal feature F fused Labeled as S, the text data contains the text feature F text of the multimodal feature F fused Labeled as T and stored in the archive database; S4: Construct a deep learning model, use the loss function to train and optimize the deep learning model until the performance of the deep learning model no longer improves; S5: According to the user's voice query, convert the voice query into text, calculate the semantic information of the text, use the semantic information as the input of the deep learning model, and retrieve the target file in the archive database.
[0006] Preferably, in step S2, the audio data preprocessing includes first removing the background noise in the audio data, unifying the format of the audio data, and then using the Fourier transform method to extract the time-domain feature F time and the frequency-domain feature F frep from the audio data; The image data preprocessing includes first grayscaling the image data, then scaling, cropping, and rotating it to make the size and format of the image data the same, denoising the image with the unified size and format, and finally using the RGB-to-grayscale method to extract the color feature F color , the texture feature F texture and the shape feature F shape from the denoised image; The text data preprocessing includes first removing the stop words, punctuation marks, and useless characters in the text data, then vectorizing the text data, and using TF-IDF to extract the text feature F text from the text data; Normalize the features extracted above, and the expression is: where F αj is the j-th feature vector of the α-th type, is the j-th normalized feature vector of the α-th type, μ αj is the mean vector of the j-th feature vector of the α-th type, σ αj is the standard deviation vector of the j-th feature vector of the α-th type, and α is the time-domain feature F time , the color feature F color , the texture feature F texture , the shape feature F shape and the text feature F textAny one of them; Use the sigmoid algorithm to normalize the time-domain feature F time , color feature F color , texture feature F texture , shape feature F shape and text feature F text to perform fusion to obtain the multi-modal feature F fused , and the expression is: where α is the time-domain feature F time , color feature F color , texture feature F texture , shape feature F shape and text feature F text , W i is the weight corresponding to different modal features, σ is the activation function, b is the bias term, and τ is a combination of any one or more of R, S, and T.
[0007] Preferably, in the step S4, specifically: Construct a deep learning model, including an input layer, a hidden layer, a fusion layer, and an output layer. The expression of the forward propagation process of each layer of the deep learning model is: z (l) = W (l) a (l-1) + b (l) a (l) = e(z (l) ) where Z (l) is the weighted input of the l-th layer, W (l) is the weight matrix of the l-th layer, b (l) is the bias vector of the l-th layer, and e(·) is a non-linear activation function; The expression of the output layer outputting P is: P = β(W out a last + b out ) where β(·) is the sigmoid activation function, W out is the weight matrix of the output layer, b out is the bias of the output layer, and a last is the activation output of the last layer of the hidden layer; Train the deep learning model through the loss function and use the SGD algorithm to optimize and minimize the loss function, thereby updating the parameters of the deep learning model.
[0008] Preferably, the hidden layer includes a fully connected layer, a convolutional layer, a recurrent layer, and an attention mechanism layer.
[0009] Preferably, the semantics of the user can be divided into seven categories according to tags R, S, and T, including R, S, T, RS, RT, ST, and RST.
[0010] Preferably, the specific steps in step S5 are as follows: S5.1: Use ASR technology to process the audio signal O of the user's voice query into text O'; S5.2: Use NLP technology to extract the semantic information vector γ from text O'; S5.3: Perform the same semantic analysis on each record δ in the database i and extract the semantic information vector γ δi ; S5.4: Calculate the similarity between the semantic information vector γ of the query and the semantic information vector γ i of each record δ in the database δi The expression is: S5.5: Select the document record with the highest score as the retrieval result according to the similarity score. The expression is: where δ i∈ is the data archive, and δ is the document with the highest similarity.
[0011] Preferably, the weight W in step S2 i The specific steps are as follows: Step 1: Take the time-domain feature F time , color feature F color , texture feature F texture , shape feature F shape and text feature F text as the input of the model. Let G be the target variable, then the expression of the model is: G = W 0+ W1F time + W2F time + W3F time + W4F time + W5F time + ε where W0 is the intercept term, W1, W2, W3, W4, W5 ∈ W i , and ε is the error term; Step 2: Use the normal distribution as the prior. The expression is: where is the prior variance, and i = 0, 1, 2, 3, 4, 5; Step 3: Update the distribution of weights using the observed data and combine it with the prior distribution to obtain the posterior distribution.
[0012] Compared with the prior art, the present invention provides an artificial intelligence-based file management method, which has the following beneficial effects: This artificial intelligence-based file management method converts file data of different modalities into a unified multi-modal feature representation through steps such as preprocessing, feature extraction, and standardization, achieving the elimination of data heterogeneity and the effective extraction of features. This multi-modal fusion method can make full use of the complementarity between different modality data and improve the overall expression ability of the data. Users can convert speech to text through voice query and calculate the semantic information of the text. This semantic-level retrieval method is more accurate and flexible than traditional keyword matching. Based on the input of the deep learning model, the target file in the file database can be quickly retrieved, improving the efficiency and accuracy of retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic flowchart of an artificial intelligence-based file management method proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0015] Please refer to Figure 1 , an artificial intelligence-based file management method, including the following steps: S1: Collect various types of file data, including audio data, image data, and text data; S2: Preprocess the collected file data, extract features from the preprocessed file data, standardize the extracted features, and perform multi-modal fusion on the standardized features to obtain multi-modal feature F fused ; S3: Mark the multi-modal feature F time containing time-domain feature F frep and frequency-domain feature F fused in the audio data as R, mark the multi-modal feature F color containing color feature F texture , texture feature F shape , and shape feature F fused in the image data as S, and mark the multi-modal feature F containing text feature F in the text datatext Multimodal feature F fused It is labeled as T and stored in the archive database; according to the tags R, S, and T, the semantics of the user can be divided into seven categories, including R, S, T, RS, RT, ST, and RST. The model can reduce the dimension of the retrieval process according to the semantics of the user. In the first layer, it can first index the category belonging to one of R, S, T, RS, RT, ST, and RST, and then further index within one of R, S, T, RS, RT, ST, and RST to reduce the dimension of the processing process.
[0016] S4: Build a deep learning model, use the loss function to train and optimize the deep learning model until the performance of the deep learning model no longer improves; For each training batch, randomly select a batch of samples from the archive database.
[0017] Input this batch of samples into the model, perform forward propagation, and calculate the loss function.
[0018] Perform backpropagation and calculate the gradient of the loss function with respect to the model parameters.
[0019] Use an optimization algorithm to update the model parameters.
[0020] Repeat the above steps until a predetermined number of training epochs is reached or the performance on the validation set no longer improves S5: According to the user's voice query, convert the voice query into text, calculate the semantic information of the text, use the semantic information as the input of the deep learning model, and retrieve the target archive in the archive database.
[0021] Before that, it can also be: if the model output indicates that the input data is unauthorized access, trigger the alarm system. The alarm system can include sound alarms, sending email or text message notifications, recording event logs, etc.
[0022] Authorized access and archive location detection: If the model output indicates that the input data is authorized access, use the input features to quickly detect the location of the archive in the archive database.
[0023] Display the retrieved target archive to the user, and the user can select the interested archive for further viewing or processing according to the displayed results.
[0024] Furthermore, it can also be: regularly evaluate the performance of the model in actual applications, and adjust the model parameters or retrain the model as needed. Data update: As new data is collected, regularly update the training set to ensure that the model can adapt to new access patterns and abnormal behaviors. Security maintenance: Regularly check the security of the system to prevent potential attacks or data leaks.
[0025] In step S2, the preprocessing of audio data includes first removing background noise from the audio data, unifying the format of the audio data, and then using the Fourier transform method to extract the time-domain feature F in the audio data time and the frequency-domain feature F frep from it; The preprocessing of image data includes first grayscaling the image data, then scaling, cropping, and rotating it to make the size and format of the image data the same, denoising the image with the unified size and format, and finally using the RGB-to-grayscale method to extract the color feature F color , texture feature F texture and shape feature F shape from the denoised image; The preprocessing of text data includes first removing stop words, punctuation marks, and useless characters from the text data, then vectorizing the text data, and using TF-IDF to extract the text feature F text from the text data; Normalize the features extracted above. The expression is: where F αj is the j-th feature vector of the α-th type, is the j-th normalized feature vector of the α-th type, μ αj is the mean vector of the j-th feature vector of the α-th type, σ αj is the standard deviation vector of the j-th feature vector of the α-th type, and α is any one of the time-domain feature F time , color feature F color , texture feature F texture , shape feature F shape and text feature F text among them; Use the sigmoid algorithm to fuse the normalized time-domain feature F time , color feature F color , texture feature F texture , shape feature F shape and text feature F text to obtain the multi-modal feature F fused . The expression is: where α is the time-domain feature F time , color feature F color , texture feature F texture , shape feature F shape and text feature F text , and W iis the weight corresponding to different modal features, σ is the activation function, b is the bias term, and τ is a combination of any one or more of R, S, and T.
[0026] Specifically in step S4: Construct a deep learning model, including an input layer, a hidden layer, a fusion layer, and an output layer. The expression for the forward propagation process of each layer of the deep learning model is: z (l) = W (l) a (l-1) + b (l) a (l) = e(z (l) ) where Z (l) is the weighted input of the l-th layer, W (l) is the weight matrix of the l-th layer, b (l) is the bias vector of the l-th layer, and e(·) is a non-linear activation function; The expression for the output P of the output layer is: P = β(W out a last + b out ) where β(·) is the sigmoid activation function, W out is the weight matrix of the output layer, b out is the bias of the output layer, and a last is the activation output of the last layer of the hidden layer; Train the deep learning model through the loss function and use the SGD algorithm to optimize and minimize the loss function, thereby updating the parameters of the deep learning model.
[0027] The hidden layer includes a fully connected layer, a convolutional layer, a recurrent layer, and an attention mechanism layer.
[0028] Specific steps in step S5 are: S5.1: Use ASR technology to process the audio signal O of the user's voice query into text O'; S5.2: Use NLP technology to extract the semantic information vector γ from the text O'; S5.3: Perform the same semantic analysis on each record δ in the database i and extract the semantic information vector γ δi ; S5.4: Calculate the similarity between the semantic information vector γ of the query and the semantic information vector γ i of each record δ in the database δi , and the expression is: S5.5: Select the document record with the highest score as the retrieval result according to the similarity score. The expression is: where δ i∈ is the data archive, and δ is the document with the highest similarity.
[0029] The weight W in step S2 i The specific steps are as follows: Step 1: Take the time-domain feature F time , color feature F color , texture feature F texture , shape feature F shape and text feature F text as the input of the model. Let G be the target variable, then the expression of the model is: G = W 0+ W1F time + W2F time + W3F time + W4F time + W5F time + ε. Where W0 is the intercept term, W1, W2, W3, W4, W5 ∈ W i , and ε is the error term; Step 2: Use the normal distribution as the prior. The expression is: where is the prior variance, and i = 0, 1, 2, 3, 4, 5; Step 3: Update the distribution of the weights using the observed data and combine it with the prior distribution to obtain the posterior distribution.
[0030] In the Bayesian framework, methods such as Bayesian networks or Bayesian optimization can be used to estimate the weights of features. These methods usually need to consider the prior distribution of features and their dependencies.
[0031] In summary, this AI-based archive management method converts archive data of different modalities into a unified multi-modal feature representation through steps such as preprocessing, feature extraction, and standardization, achieving the elimination of data heterogeneity and the effective extraction of features. This multi-modal fusion method can make full use of the complementarity between different modality data, improve the overall expression ability of data. Users can convert speech to text through voice queries and calculate the semantic information of the text. This semantic-level retrieval method is more accurate and flexible than traditional keyword matching. Based on the input of the deep learning model, it can quickly retrieve the target archives in the archive database, improving the efficiency and accuracy of retrieval.
[0032] It should be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising said element.
[0033] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An archive management method based on artificial intelligence, characterized in that: The following steps are involved: S1: Collect various types of archival data, including audio data, image data and text data; S2: Preprocess the collected archival data, extract features from the processed archival data, standardize the extracted features, and perform multimodal fusion on the standardized features to obtain multimodal features F fused ; S3: The audio data contains time domain features F time and the frequency domain feature F frep The multimodal features F fused Marked as R, the image data contains color features F color , texture feature F texture With shape feature F shape The multimodal features F fused Marked as S, the text data contains text features F text The multimodal features F fused Mark it as T and store it in the archive database; S4: Build a deep learning model and use the loss function to train and optimize the deep learning model until the performance of the deep learning model no longer improves; S5: According to the user's voice query, the voice query is converted into text, and the semantic information of the text is calculated. The semantic information is used as input based on the deep learning model to retrieve the target archive in the archive database.
2. The artificial intelligence-based archive management method according to claim 1, characterized in that: The audio data preprocessing in step S2 includes first removing the background noise in the audio data, unifying the format of the audio data, and then using the Fourier transform method to convert the time domain features F in the audio data into time and the frequency domain feature F frep Perform extraction; Image data preprocessing includes graying the image data, scaling, cropping, and rotating it to make the image data size and format consistent, and denoising the image after the size and format are unified. Finally, the denoised image is converted to color feature F using the RGB to grayscale method. color , texture feature F texture With shape feature F shape Extraction of Text data preprocessing includes removing stop words, punctuation marks and useless characters from the text data, vectorizing the text data, and using TF-IDF to vectorize the text features of the text data. text Perform extraction; The above extracted features are standardized and the expression is: Among them, F αj is the j-th eigenvector of the α-th type, is the jth normalized eigenvector of the αth type, μ αj is the mean vector of the j-th eigenvector of the α-th type, σ αj is the standard deviation vector of the jth feature vector of the αth type, α is the time domain feature F time 、Color feature F color , texture feature F texture , shape feature F shape With the text feature F text Any of these; Use the sigmoid algorithm to normalize the time domain features F time 、Color feature F color , texture feature F texture , shape feature F shape With the text feature F text Fusion is performed to obtain the multimodal feature F fused , the expression is: Among them, α is the time domain feature F time 、Color feature F color , texture feature F texture , shape feature F shape With the text feature F text , W i is the weight corresponding to different modal features, σ is the activation function, b is the bias term, and τ is any one or more combinations of R, S, and T.
3. The artificial intelligence-based archive management method according to claim 1, characterized in that: The step S4 specifically includes: The deep learning model is constructed, including the input layer, hidden layer, fusion layer and output layer. The forward propagation process of each layer based on the deep learning model is expressed as: z (l) =W (l) a (l-1) +b (l) a (l) =e(with (l) ) Among them, Z (l) is the weighted input of the lth layer, W (l) is the weight matrix of the lth layer, b (l) is the bias vector of the lth layer, e(·) is the nonlinear activation function; The expression of the output layer output P is: P=β(W out a last +b out ) Among them, β(·) is the sigmoid activation function, W out is the weight matrix of the output layer, b out is the bias of the output layer, a last is the activation output of the last hidden layer; The deep learning model is trained through the loss function, and the SGD algorithm is used to optimize and minimize the loss function, thereby updating the parameters of the deep learning model.
4. The artificial intelligence-based archive management method according to claim 3, characterized in that: The hidden layer includes a fully connected layer, a convolutional layer, a recurrent layer and an attention mechanism layer.
5. The artificial intelligence-based archive management method according to claim 1, characterized in that ,According to the labels R, S, and T, the user's semantics can be divided into seven categories, ,including R, S, T, RS, RT, ST, and RST.
6. The artificial intelligence-based archive management method according to claim 1, characterized in that: The specific steps in step S5 are: S5.1: Using ASR technology to process the audio signal O of the user's voice query into text O'; S5.2: Use NLP technology to extract the semantic information vector γ from the text O'; S5.3: For each record in the database δ i Perform the same semantic analysis and extract the semantic information vector γ δi ; S5.4: Compute the semantic information vector γ of the query and each record δ in the database i The semantic information vector γ δi The similarity between them is expressed as: S5.5: According to the similarity score, select the document record with the highest score as the search result. The expression is: Among them, δ i∈ Data archive, δ is the document with the highest similarity.
7. The artificial intelligence-based archive management method according to claim 1, characterized in that: The weight W in step S2 i The specific steps are: Step 1: Transform the time domain feature F time 、Color feature F color , texture feature F texture , shape feature F shape With the text feature F text As the input of the model, G is the target variable, and the expression of the model is: G=W 0+ W1F time +W2F time +W3F time +W4F time +W5F time +ε Where W0 is the intercept term, W1, W2, W3, W4, W5∈W i ,ε is the error term; Step 2: Use normal distribution as a priori, the expression is: in, is the prior variance, i = 0, 1, 2, 3, 4, 5; Step 3: Use the observed data to update the distribution of weights and combine it with the prior distribution to get the posterior distribution.
Citation Information
Cited By
Intelligent file classification and retrieval method and system based on artificial intelligence
CN121166924A