Archive management method and system based on multi-modal data analysis technology

Through multimodal data analysis technology, the problem of inefficiency in traditional archive management is solved, automatic data classification and high-quality summary generation are realized, and the quality and efficiency of archive management are improved.

CN120067325AActive Publication Date: 2025-05-30FUJIAN YIRONG INFORMATION TECH

Patent Information

Application Number
CN202510544115.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

Traditional archive management methods are inefficient and error-prone, making it difficult to effectively manage and analyze large amounts of unstructured data.

Method used

The archive management method based on multimodal data analysis technology is adopted, and the automatic classification, abstract generation and personalized archiving strategies of data are realized through adaptive denoising technology, text classification model, sequence-to-sequence generation model and reinforcement learning model.

Benefits of technology

It improves the quality and efficiency of archive management, ensures the accessibility and availability of data, and realizes accurate classification of text data and high-quality summary generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067325A_ABST
    Figure CN120067325A_ABST
Patent Text Reader

Abstract

The invention relates to an archive management method and system based on a multi-modal data analysis technology, and the method comprises the following steps: S1, collecting archive data related to electric power, and employing a self-adaptive denoising technology to improve the processing quality of various types of files; s2, based on the de-noised archive data related to the electric power, converting a text into an integer index sequence by using tokenizer, standardizing the sequence into the same length by using padding, and performing automatic data classification and label generation by using a text classification model; s3, automatically generating abstracts for technical files and reports by adopting a sequence-to-sequence generation model on the basis of the classified text data; and S4, making a personalized archive filing strategy according to user behavior analysis so as to meet specific requirements of different departments or needs. According to the method, not only are the accessibility and availability of data improved, but also the file management quality and efficiency are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of file management, and particularly to a file management method and system based on multi-modal data analysis technology. Background Art

[0002] In today's information age, enterprises and organizations generate and accumulate a large amount of data during their operations. This data not only includes structured data but also a large amount of unstructured data such as text, images, and audio. With the in-depth digital transformation, especially in the power industry, the amount of data and data complexity are both increasing continuously. These data contain rich information value, but at the same time, they also pose higher requirements for data storage, management, and analysis.

[0003] Traditional file management methods usually rely on manual classification and storage, which are inefficient and error-prone. With the development of artificial intelligence and data science technologies, intelligent file management technologies have gradually emerged. Summary of the Invention

[0004] To solve the above problems, the purpose of the present invention is to provide a file management method and system based on multi-modal data analysis technology, which not only improves the accessibility and usability of data but also effectively improves the quality and efficiency of file management.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A file management method based on multi-modal data analysis technology, comprising the following steps: S1: Collect file data related to electricity and adopt an adaptive denoising technology to improve the processing quality of various types of files; S2: Based on the electricity-related file data after denoising processing, use a tokenizer to convert the text into an integer index sequence, use padding to standardize the sequences to the same length, and adopt a text classification model for automatic data classification and label generation; S3: Based on the classified text data, for technical documents and reports, adopt a sequence-to-sequence generation model to automatically generate abstracts; S4: Develop a personalized file archiving strategy according to user behavior analysis to meet the specific requirements of different departments or needs; The specific content of S4 is as follows: S41: Collect user interaction data and user meta-information. The user interaction data includes usage logs, clickstreams, search queries, access frequencies, and file modification records; the user meta-information includes user roles, departments, and permission levels; S42: Extract features, including the correlation between access frequency, stay time, access order, and the most frequently searched keywords, and use cluster analysis for behavior patterns to divide users into different groups; S43: Construct a graph network model for user interaction, where nodes represent files and users, and edge weights represent interaction intensity. Apply graph analysis techniques to identify highly connected groups or nodes; S44: Apply association rule mining to identify co-occurrence behaviors and formulate targeted strategies; S45: Use a reinforcement learning model to generate dynamic archiving strategies, adaptively adjust the organizational structure of the archive according to user behavior, and define specific archiving rules based on the user request speed and data storage efficiency. The archiving rules include priorities and automatic archiving trigger conditions.

[0006] Further, S1 specifically includes: Collect power-related archive data, which includes equipment maintenance records, operation logs, sensor data, and contract texts; Perform text data denoising on equipment maintenance records, operation logs, and contract texts, including using TextBlob for spelling and grammar correction, unifying text formats, and removing redundant information; use regular expressions to remove unstructured noise, define common noise patterns, and automatically filter the identified noise; Perform missing value imputation, smoothing, and anomaly detection and correction on sensor data.

[0007] Further, use a tokenizer to convert the text into an integer index sequence, and use padding to standardize the sequences to the same length, as follows: Use the tokenizer to create a vocabulary and index mapping for the entire corpus. Traverse the corpus, count the frequency of each word, create a unique vocabulary, and assign a unique integer index to each word in the vocabulary; Convert each text sentence into an index sequence. According to the vocabulary mapping, replace each word in the text with its index in the vocabulary; Preset the maximum sequence length according to the task requirements and computing power. Truncate sequences that are too long and pad sequences that are too short; for sequences with a length less than the maximum length max_len, use padding tokens to pad; for sequences that exceed the maximum length, truncate them; For the sequence [x1,x2,…,xm]: m is the sequence length; If the sequence length m < max_len, then pad the sequence to: [x1,x2,…,xm,0,0,…,0] Where the padded sequence length is max_len.

[0008] If m > max_len, truncate the sequence to: [x1, x2, …, xmax_len].

[0009] Furthermore, the construction of the text classification model is as follows: The input layer uses a pre-trained word embedding layer to convert words into vectors; The input size is: (batch_size, sequence_length, embedding_dim); where batch_size is the batch size, sequence_length is the maximum number of words in each line of text, and embedding_dim is the dimension of the word vector; Add multiple convolutional layers, each using a different kernel size, and add corresponding pooling layers after each convolutional layer; ; where, h i:j represents a segment of word embedding sequence starting at position i with a length of , representing any segment of the text; w k is the weight parameter of the convolutional kernel, b is the bias, f is the activation function ReLU; E(x i+k ): represents the word embedding vector at the i + k-th position in the input sequence, and K represents the width of the convolutional kernel; The pooling layer uses max pooling to extract the important features of the output of each convolutional layer: p = max(h); where, h is the result after convolution, and p is the feature after pooling; The pooled features are flattened and used as the input to the fully connected layer; Finally, the output of the fully connected layer passes through Softmax to produce the final class prediction; The cross-entropy loss is used as the loss function during the training process, and the adaptive optimization algorithm AMS is combined to optimize the weights; Deploy the trained text classification model to the production environment to achieve automatic classification of text data.

[0010] Furthermore, the weight optimization is combined with AMS as follows: Initialize the first-order momentum m 0 and the second-order momentum v 0 for each parameter θ: m 0 ← 0, v 0 ← 0; Set the learning rate α, the decay rates β1 and β2 of the first-order and second-order momenta; Set the small value Prevent division-by-zero errors, where subscript t represents the corresponding time step; Iteratively update at each time step t: Calculate the gradient g t : ; Update the first-order momentum m t : ; Update the second-order momentum v t : ; Maintain the max second-order momentum estimate : ; Bias-correct the first-order momentum: ; Parameter update: ; By keeping the maximum historical value of the second-order momentum, the problem of over-relaxation in momentum updates is avoided.

[0011] Furthermore, the sequence-to-sequence generation model is specifically: In the sequence-to-sequence generation model, through the encoder-decoder architecture, the generated language model is used to predict the sequence, sharing the information of the context representation; Encoder: Convert the input text into a context-aware hidden state, and use the Transformer model to generate the embedded representation of the input sequence; Decoder: Generate the target sequence from the hidden state of the encoder, using the architecture design of the Transformer model; The decoder looks at the output of the encoder at each generation step according to importance to form a dynamic context vector, using the multi-head self-attention mechanism to adjust the attention weights and focus on the key information of the input text; For the attention mechanism, given the input matrix X, after projecting it into query Q, key K, and value V matrices: Q = XW Q , K = XW K , V = XW V ; where W Q , W K , W V are the weight matrices for query, key, and value respectively; Attention calculation: ; where d k is the dimension of the key vector; During decoding, the beam search strategy is used to improve stability by exploring multiple sequences simultaneously; the Copy Mechanism is adopted to increase the ability to directly copy words from the input, which is especially helpful for generating keywords in technical texts; The standard cross-entropy loss of the sequence-to-sequence model is used to drive the model to generate the target sentence, and the coverage loss is introduced. The relative intensity weights are adjusted by combining them with the standard cross-entropy loss in a weighted manner, making the generated text have a certain creativity while still maintaining consistency. The BLEU and ROUGE metrics are used to quantitatively analyze the n-gram or subsequence overlap between the generated summary and the reference summary, evaluating the accuracy and consistency of the generated text.

[0012] Furthermore, during decoding, the beam search strategy is used to improve stability by exploring multiple sequences simultaneously, as follows: Set e, that is, the number of candidate sequences retained at each time step; Initialize the candidate sequence set to the starting state; For each time step t, calculate the probability of each current sequence expanding the next word. For a given candidate sequence S, the probability of its expansion is: ; where y t is the generated word at time step t, X is the input sequence; T is the time step length; Select the e×r words with the highest probability from the next words of each candidate sequence, where r is the vocabulary size; Retain the e extended sequences with the highest probability as the candidate sequences for the next time step; Continue to expand until the maximum sequence length is reached or all sequences output the end marker <eos>。

[0013] Furthermore, the Copy Mechanism is adopted to increase the ability to directly copy words from the input as follows: Calculate the probabilities for generating words and copying words from the input respectively; Use a learnable parameter P gen To determine whether to generate or copy from the words: ; where σ is the sigmoid function; h t and s t are the encoder hidden state and the decoder hidden state at time step t respectively; is the input at the previous time step; w gen is the weight matrix; Calculate the word distribution P of the generation model vocab : ; s t is the hidden state of the decoder at time step t; W out is the weight matrix of the output layer; b out is the bias of the output layer; Calculate the input copy probability P based on each input position copy ; ; where represents the input position where the input word is equal to the current output word y t ; where is the weight that the attention assigns to the input position ; Combine the probabilities of generation and copying as the final output word probability P ( y t ): 。

[0014] An archive management system based on multimodal data analysis technology, including a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned archive management method based on multimodal data analysis technology.

[0015] The present invention has the following beneficial effects: 1. The present invention not only improves the accessibility and usability of data, but also effectively enhances the quality and efficiency of file management; 2. The present invention ensures that all text data is converted into integer index sequences of the same length, making it suitable for input into a deep learning model for batch processing, efficiently processing text inputs of different lengths, enabling the text classification model to achieve accurate classification and label generation, and effectively improving the classification quality; and using the cross-entropy loss function and the AMSGrad optimization algorithm, the model can obtain a more stable learning process during training, while enhancing the modeling ability for complex classification tasks, effectively improving the stability and accuracy of the text classification model; 3. By combining Beam Search and Copy Mechanism, the present invention enables the generation model to improve the quality of generated outputs when processing natural language, combining the exploratory nature of Beam Search with the precision of Copy Mechanism to generate more fluent and appropriate text paragraphs, capable of obtaining more accurate abstract content, and effectively improving the quality of file management. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following further elaborates on the present invention in detail with reference to the accompanying drawings and specific embodiments: Reference Figure 1 , in this embodiment, a file management method based on multi-modal data analysis technology is provided, including the following steps: S1: Collect power-related file data and adopt adaptive denoising technology to improve the processing quality of various types of files; S2: Based on the power-related file data after denoising processing, use a tokenizer to convert the text into an integer index sequence, use padding to standardize the sequence to the same length, and adopt a text classification model for automatic data classification and label generation; S3: Based on the classified text data, for technical documents and reports, adopt a sequence-to-sequence generation model to automatically generate abstracts; S4: Develop a personalized file archiving strategy according to user behavior analysis to meet the specific requirements of different departments or needs; Specifically, S4 is as follows: S41: Collect user interaction data and user meta-information, where the user interaction data includes usage logs, clickstreams, search queries, access frequencies, and file modification records; the user meta-information includes user roles, departments, and permission levels; S42: Extract features, including the correlation between access frequency, stay time, access order, and most frequently searched keywords, and adopt cluster analysis for behavior patterns to divide users into different groups; S43: Construct a graph network model for user interaction, where nodes represent files and users, and edge weights represent interaction intensity. Apply graph analysis techniques to identify highly connected groups or nodes; S44: Apply association rule mining (such as the Apriori algorithm) to identify co-occurrence behaviors and formulate targeted strategies; for example, users who frequently access a certain type of file often have organizational requirements for similar archives.

[0018] S45: Use a reinforcement learning model to generate dynamic archiving strategies, adaptively adjust the organizational structure of archives according to user behavior, and define specific archiving rules based on the user request speed and data storage efficiency. The archiving rules include priorities and automatic archiving trigger conditions.

[0019] In this implementation, the specific content of S1 is as follows: Collect power-related archive data, where the archive data includes equipment maintenance records, operation logs, sensor data, and contract texts; Equipment maintenance records: mainly structured text data, which may contain inconsistent formats and spelling mistakes.

[0020] Operation logs: semi-structured text data, which may contain redundant information and error records. Sensor data: numerical time series data, which may contain noise, missing values, and outliers. Contract texts: unstructured text data, which may have format problems and semantic errors due to OCR recognition; For equipment maintenance records, operation logs, and contract texts, perform text data denoising processing, including using TextBlob for spelling and grammar correction, unifying text formats (such as date formats, unit formats), and removing redundant information; use regular expressions to remove unstructured noise, define common noise patterns, and automatically filter out the identified noise; For sensor data, perform missing value filling, smoothing processing, and anomaly detection and correction processing.

[0021] In this implementation, use a tokenizer to convert text into an integer index sequence, and use padding to standardize the sequences to the same length, as follows: Use a tokenizer to create a vocabulary and index mapping for the entire corpus. Traverse the corpus, count the frequency of each word, create a unique vocabulary, and assign a unique integer index to each word in the vocabulary; Convert each text sentence into an index sequence. According to the vocabulary mapping, replace each word in the text with its index in the vocabulary; Preset the maximum sequence length according to the task requirements and computing power, truncate sequences that are too long, and pad sequences that are too short; for sequences with a length less than the maximum length max_len, use padding tokens to pad them; for sequences that exceed the maximum length, truncate them. For a sequence [x1, x2, …, xm]: m is the sequence length. If the sequence length m < max_len, then pad the sequence to: [x1, x2, …, xm, 0, 0, …, 0] where the padded sequence length is max_len.

[0022] If m > max_len, then truncate the sequence to: [x1, x2, …, xmax_len].

[0023] In this implementation, the construction of the text classification model is as follows: The input layer uses a pre-trained word embedding layer to convert words into vectors. The input size is: (batch_size, sequence_length, embedding_dim) where batch_size is the batch size, sequence_length is the maximum number of words per line of text, and embedding_dim is the dimension of the word vector. Add multiple convolutional layers, each using a different kernel size, and add corresponding pooling layers after each convolutional layer. ; where, h i:j represents a segment of word embedding sequence starting at position i with a length of , representing any segment of the text; w k is the weight parameter of the convolutional kernel, b is the bias, f is the activation function ReLU; E(x i+k ): represents the word embedding vector at the i + k-th position in the input sequence, and K represents the width of the convolutional kernel. The pooling layer uses max pooling to extract important features from the output of each convolutional layer: p = max(h) where h is the result after convolution, and p is the feature after pooling. The pooled features are processed by flattening and used as the input to the fully connected layer. Finally, the output of the fully connected layer passes through Softmax to produce the final class prediction. The cross-entropy loss is used as the loss function during the training process, and AMS is combined for weight optimization. Deploy the trained text classification model to the production environment to achieve automatic classification of text data.

[0024] In this implementation, weight optimization is combined with AMS as follows: Initialize the first-order momentum m for each parameter θ 0 and the second-order momentum v 0 : m 0 ←0, v 0 ←0; Set the learning rate α, the decay rates β1 and β2 of the first-order and second-order momenta; Set a small value to prevent division-by-zero errors, and the subscript t represents the time step t; Iteratively update at each time step t: Calculate the gradient g t : ; Update the first-order momentum m t : ; Update the second-order momentum v t : ; Maintain the max second-order momentum estimate : ; Bias-correct the first-order momentum: ; Parameter update: ; By maintaining the maximum historical value of the second-order momentum, the problem of over-relaxation in momentum updates is avoided.

[0025] In this implementation, the sequence-to-sequence generation model is specifically: In the sequence-to-sequence generation model, through the encoder-decoder architecture, use the generative language model to predict the sequence and share the information of the context representation; Encoder: Convert the input text into a context-aware hidden state and use the transformer model to generate the embedded representation of the input sequence; Decoder: Generates the target sequence from the hidden state of the encoder, using the architecture design of Transformer, including an input layer, multiple Transformer decoder layers, and an output layer; the input layer inputs to generate the target sequence from the hidden state of the encoder, and the multiple Transformer decoder layers include masked multi-head self-attention, encoder-decoder attention, feed-forward network, residual, and normalization layers; the output layer outputs the vocabulary probability distribution through softmax; decoding generates each word of the target sequence in turn, and the output of the current step is used as the input of the next step until the end symbol is encountered.

[0026] The decoder looks at the output of the encoder at each generation step according to importance to form a dynamic context vector, uses the multi-head self-attention mechanism to adjust the attention weights, and focuses on the key information of the input text. For the attention mechanism, given the input matrix X, after projecting it into query Q, key K, and value V matrices: Q = XW Q , K = XW K , V = XW V ; where, W Q , W K , W V are the weight matrices for query, key, and value respectively. Attention calculation: ; where, d k is the dimension of the key vector. During decoding, the beam search strategy is used to improve stability by exploring multiple sequences simultaneously; the Copy Mechanism is adopted to increase the ability to directly copy words from the input, which is especially helpful for generating keywords in technical texts. The standard cross-entropy loss of the sequence-to-sequence model is used to drive the model to generate the target sentence, and a diversification loss (such as encouraging the coverage of new n-grams, non-repetition, and punishing patterned outputs) is introduced. The relative strength weights are adjusted by weighted combination with the standard cross-entropy loss to make the generated text have a certain degree of creativity while still maintaining consistency. The BLEU and ROUGE metrics are used to quantitatively analyze the n-gram or subsequence overlap degree between the generated summary and the reference summary to evaluate the accuracy and consistency of the generated text.

[0027] In this implementation, during decoding, the beam search strategy is used to improve stability by exploring multiple sequences simultaneously, as follows: Set e, that is, the number of candidate sequences retained at each time step; Initialize the candidate sequence set to the starting state; For each time step t, calculate the probability of the next word for each current sequence extension. For a given candidate sequence S, the probability of its extension is: ; where y t is the generated word at time step t, X is the input sequence; T is the time step length; Select the top e×r words with the highest probabilities from the next words of each candidate sequence, where r is the vocabulary size; Retain the top e extended sequences with the highest probabilities as the candidate sequences for the next time step; Continue to extend until the maximum sequence length is reached or all sequences output the end marker <eos>。

[0028] In this implementation, the Copy Mechanism is adopted to increase the ability to directly copy words from the input as follows: Calculate the probabilities for generating words and copying words from the input respectively; Use a learnable parameter P gen (generation probability) to determine whether to generate or copy from the word: ; where σ is the sigmoid function; h t and s t are the encoder hidden state and decoder hidden state at time step t respectively; is the input at the previous time step; w gen is the weight matrix; Calculate the word distribution P of the generation model vocab : ; s t is the hidden state of the decoder at time step t; W out is the weight matrix of the output layer; b out is the bias of the output layer; Calculate the input copy probability P based on each input position copy ; ; where, represents the input position where the input word is equal to the current output word y t ; where, is the weight that the attention assigns to the input position ; Combine the probabilities of generation and copying as the final output word probability P ( y t ): 。

[0029] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0030] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0031] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0032] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0033] As mentioned above, it is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.< / eos> < / eos>

Claims

1. An archive management method based on multimodal data analysis technology, characterized in that: The following steps are involved: S1: Collect archival data related to electricity and use adaptive denoising technology to improve the processing quality of various types of files; S2: Based on the denoised power-related archival data, a tokenizer is used to convert the text into an integer index sequence, padding is used to standardize the sequences to the same length, and a text classification model is used to automatically classify the data and generate labels; S3: Based on the classified text data, a sequence-to-sequence generation model is used to automatically generate summaries for technical documents and reports; S4: Develop personalized archiving strategies based on user behavior analysis to meet the specific requirements of different departments or needs; The S4 is specifically: S41: Collecting user interaction data and user meta information, wherein the user interaction data includes usage logs, click streams, search queries, access frequencies, and file modification records; and the user meta information includes user roles, departments, and authority levels; S42: Extract features, including visit frequency, dwell time, visit sequence, and correlation between the most frequently searched keywords, and use cluster analysis behavior patterns to divide users into different groups; S43: Build a graph network model of user interaction, where nodes represent files and users, edge weights represent interaction strength, and apply graph analysis techniques to identify highly connected groups or nodes; S44: Apply association rule mining to identify co-occurrence behaviors and formulate targeted strategies; S45: Use a reinforcement learning model to generate a dynamic archiving strategy, adaptively adjust the organizational structure of the archive according to user behavior, and define specific archiving rules based on user request speed and data storage efficiency. The archiving rules include priorities and automatic archiving trigger conditions.

2. The archive management method based on multimodal data analysis technology according to claim 1, characterized in that: The S1 is specifically: Collecting archival data related to electricity, including equipment maintenance records, operation logs, sensor data, and contract texts; Perform text data denoising on equipment maintenance records, operation logs and contract texts, including using TextBlob for spelling and grammar correction, unifying text formats, and removing redundant information; using regular expressions to remove unstructured noise, define common noise patterns, and automatically filter identified noise; Sensor data is processed by filling missing values, smoothing, and anomaly detection and correction.

3. The archive management method based on multimodal data analysis technology according to claim 1, characterized in that: The tokenizer is used to convert the text into a sequence of integer indexes, and padding is used to standardize the sequences to the same length, as follows: Use tokenizer to create vocabulary and index mapping for the entire corpus, traverse the corpus, count the frequency of each word, create a unique vocabulary, and assign a unique integer index to each word in the vocabulary; Convert each text sentence into a sequence of indices, replacing each word in the text with its index in the vocabulary according to the vocabulary mapping; According to task requirements and computing power, the maximum length of the sequence is preset, and sequences that are too long are truncated and sequences that are too short are padded. For sequences that are less than the maximum length max_len, padding is used to fill them. For sequences that exceed the maximum length, truncations are performed. For the sequence [x1, x2, …, xm]: m is the length of the sequence; If the sequence length m < max_len, then pad the sequence to: [x1, x2, …, xm, 0, 0, …, 0] where the padded sequence length is max_len; If m > max_len, then truncate the sequence to: [x1, x2, …, xmax_len].

4. The archive management method based on multimodal data analysis technology according to claim 3 is characterized in that: The construction of the text classification model is as follows: The input layer uses a pre-trained word embedding layer to convert words into vectors; The input size is: (batch_size, sequence_length, embedding_dim); where batch_size is the batch size, sequence_length is the maximum number of words per line of text, and embedding_dim is the dimension of the word vector; Add multiple convolutional layers, each using a different kernel size, and add a corresponding pooling layer after each convolutional layer; ; Among them, h i:j It means starting at position i and having a length of A word embedding sequence of w represents any section of text; k is the weight parameter of the convolution kernel, b is the bias, and f is the activation function ReLU; E(x i+k ): represents the word embedding vector at the i+kth position of the input sequence, and K represents the width of the convolution kernel; The pooling layer uses max pooling to extract important features from the output of each convolutional layer: p = max(h); where h is the result after convolution and p is the pooled feature; The pooled features are flattened and used as the input to the fully connected layer; Finally, the output of the fully connected layer passes through Softmax to produce the final class prediction; The cross-entropy loss is used as the loss function during the training process, and AMS is combined for weight optimization; Deploy the trained text classification model to the production environment to achieve automatic classification of text data.

5. The archive management method based on multimodal data analysis technology according to claim 4 is characterized in that: The weight optimization in combination with AMS is as follows: Initialize the first-order momentum m0 and second-order momentum v0 for each parameter θ: m0←0, v0←0; set the learning rate α, the decay rates of the first-order momentum and the second-order momentum β1, β2; set the small value To prevent division by zero errors, the subscript t indicates the corresponding time step; Iteratively update, at each time step t: Calculate the gradient g t : ; Update the first-order momentum m t : ; Update the second-order momentum v t : ; Maintaining the max second-order momentum estimate : ; Bias-corrected first-order momentum: ; Parameter update: ; By maintaining the maximum historical value of the second-order momentum, the problem of over-relaxation in momentum updates is avoided.

6. The archive management method based on multimodal data analysis technology according to claim 1, characterized in that: The sequence-to-sequence generation model is specifically: In the sequence-to-sequence generation model, through the encoder-decoder architecture, a generative language model is used to predict sequences, sharing the information of the context representation; Encoder: Convert the input text into a context-aware hidden state, and use the Transformer model to generate the embedded representation of the input sequence; Decoder: Generate the target sequence from the hidden state of the encoder, using the architecture design of Transformer; The decoder views the output of the encoder at each generation step according to importance to form a dynamic context vector, uses the multi-head self-attention mechanism, adjusts the attention weights, and focuses on the key information of the input text; For the attention mechanism, assuming the input matrix is X, after projecting it into query Q, key K, and value V matrices: Q=XW Q ,K=XW K ,V=XW V ; Among them, W Q , W K , W V Weight matrices for query, key, and value, respectively; Attention calculation: ; Among them, d k is the dimension of the key vector; During decoding, the beam search strategy is used to improve stability by exploring multiple sequences simultaneously; the CopyMechanism is adopted to increase the ability to directly copy words from the input, which is especially helpful for the generation of keywords in technical texts; The standard cross entropy loss of the sequence-to-sequence model is used to drive the model to generate target sentences, and coverage loss is introduced. The relative strength weights are adjusted in combination with the standard cross entropy loss weighting to make the generated text have a certain degree of creativity while still maintaining consistency. The BLEU and ROUGE indicators are used to quantitatively analyze the n-gram or subsequence overlap between the generated summary and the reference summary, and to evaluate the accuracy and consistency of the generated text.

7. The archive management method based on multimodal data analysis technology according to claim 6, characterized in that: The beam search strategy used during decoding improves stability by exploring multiple sequences simultaneously, as follows: Set e, the number of candidate sequences retained at each time step; Initialize the candidate sequence set to the starting state; For each time step t, the probability of expanding the next word in each current sequence is calculated. For a given candidate sequence S, the probability of its expansion is: ; Among them, y t is the generated word at time step t, X is the input sequence; T is the time step length; Select the e×r words with the highest probability from the next word of each candidate sequence, where r is the vocabulary size; Keep the e extended sequences with the highest probability as candidate sequences for the next time step; Continue to expand until the maximum sequence length is reached or all sequences output end markers <eos> 。< / eos> 8. The archive management method based on multimodal data analysis technology according to claim 7, characterized in that: The Copy Mechanism is used to add the ability to copy words directly from the input, as follows: Compute probabilities for generating words and copying words from input separately; Using a learnable parameter P gen Determines whether to generate or copy from vocabulary: ; Among them, σ is the sigmoid function; h t and t are the encoder hidden state and decoder hidden state at time step t respectively; is the input of the previous time step; w gen is the weight matrix; Calculate the vocabulary distribution P of the generative model vocab : ; s t is the hidden state of the decoder at time step t; W out is the weight matrix of the output layer; b out is the bias of the output layer; Calculate the input duplication probability P based on each input position copy ; ; in, Indicates the input location The input word Equal to the current output word y t ; in, is the attention allocated to the input position The weight of The combined generation and replication probabilities are used as the final output word probability P ( y t ): 。 9. An archive management system based on multimodal data analysis technology, characterized in that: It includes a processor, a memory and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the archive management method based on multimodal data analysis technology as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Archive management model construction method and system based on knowledge graph

    CN111737471A

  • Document recommendation method and system, terminal and storage medium

    CN115630170A

  • Intelligent archive construction method fusing artificial intelligence and knowledge graph technology

    CN115994230A

  • Convolutional neural network-based health record integration method and device, and medium

    CN117423423A

  • Automatic filing method and system for personnel archives

    CN117827750A

Cited By

  • Prescription information compression method and device and paper prescription

    CN120874879A