An artificial intelligence-based archive system and method
By introducing intelligent search modules, information sharing optimization modules, deep learning engines and adaptive classification modules into the archive management system, the problems of search difficulties and poor information sharing performance in archive management are solved, and efficient and accurate archive management and information sharing are achieved.
Patent Information
- Application Number
- CN202411229278.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-09-03
AI Technical Summary
The prior art has problems in the archive management and poor information sharing performance, resulting in inefficient file management and insufficient use of information resources.
Adopt an archive system based on artificial intelligence, including intelligent retrieval module, information sharing optimization module, deep learning engine and adaptive classification module. The intelligent search module uses natural language processing and deep neural networks for accurate search. The information sharing optimization module uses advanced encryption standards generated by quantum random number to ensure secure information sharing. The deep learning engine performs feature extraction and pattern recognition. The adaptive classification module dynamically adjusts the classification strategy.
It realizes efficient and accurate file retrieval and information sharing, improves the efficiency and service quality of archive management, and ensures the security and integrity of information.
Smart Images

Figure CN119066201B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and specifically provides an archive system and method based on artificial intelligence. Background Technique
[0002] With the rapid development of information technology, the importance of archive management has become increasingly prominent. In various fields, a large amount of information is recorded and stored as archive materials, covering various forms such as text, images, audio, and video. Traditional archive management mainly relies on manual operations and simple computer assistance. However, with the continuous increase in the number of archives and the growing variety of types, this method gradually becomes difficult to meet the requirements of modern society for efficient and accurate archive management.
[0003] There are certain defects in the existing archive system and method based on artificial intelligence. On the one hand, archive retrieval is difficult. Facing a large number and variety of archive materials, manual retrieval is difficult and error-prone, consuming a lot of time. Especially when the scale of the archive room is large, this problem is more prominent. On the other hand, in terms of archive information sharing, some existing computer management systems only regard it as a tool for storing materials, and the information sharing performance is poor. The lack of an effective sharing mechanism makes archive information unable to flow efficiently between different departments and different users, and cannot truly play the service function of the information. This not only hinders the improvement of work efficiency but also restricts the full utilization of information resources. For this reason, we propose an archive system and method based on artificial intelligence. Summary of the Invention
[0004] The purpose of the present invention is to provide an archive system and method based on artificial intelligence.
[0005] To solve the problems raised in the above background technique, the present invention provides the following technical solution: An archive system based on artificial intelligence, the archive system includes an intelligent retrieval module, an information sharing optimization module, a deep learning engine, and an adaptive classification module.
[0006] The intelligent retrieval module utilizes natural language processing and deep neural networks. Combining with the word vector model, it converts the retrieval request and archive text into vector form, and realizes accurate and timely retrieval through the cosine similarity and the correlation calculation formula introducing a time decay factor. The retrieved relevant archive data will be passed as input to the deep learning engine and the adaptive classification module. The information sharing optimization module adopts an improved algorithm of the Advanced Encryption Standard integrated with quantum random number generation. It uses a quantum random number generator to generate an initial key, encrypts the archive information in groups, and dynamically adjusts the number of encryption rounds according to the key length and security level to ensure the secure sharing of information. The deep learning engine uses the stochastic gradient descent algorithm to train a multi-layer convolutional neural network, and realizes accurate feature extraction and pattern recognition of archive data by calculating the gradient of the loss function and updating the network parameters. After obtaining the archive data from the intelligent retrieval module, the deep learning engine performs feature extraction. The extracted features can be fed back to the intelligent retrieval module to optimize the retrieval results, and can also provide richer feature information for the adaptive classification module for classification. The adaptive classification module dynamically adjusts the classification strategy by calculating the information gain value according to the archive attributes. The adaptive classification module receives the feature information extracted by the deep learning engine and the archive data provided by the intelligent retrieval module, and comprehensively considers these factors for classification adjustment, making the archive classification more reasonable and efficient. At the same time, the classification results can also provide a more accurate retrieval range and direction for the intelligent retrieval module;
[0007] The archive system further includes a data cleaning module, which uses the Z-score standardization method for numerical data and performs stemming and stop word filtering on text data.
[0008] As a further solution of the present invention: The intelligent retrieval module is provided with a word vector model, which uses the word vector model to convert the retrieval request and archive text into vector form, calculates the correlation between the retrieval request and the archive text through the cosine similarity. At the same time, a time decay factor is introduced to reflect the timeliness of the archive, and higher weights are given to newer archives. The correlation score between the retrieval request and the archive text is calculated through the correlation calculation formula. The specific correlation score calculation formula is as follows:
[0009]
[0010] Where: represents the correlation score, represents the retrieval request, archive document, represents the retrieval request and the archive document semantic similarity, represents the word frequency of the retrieval term in the document, represents the update frequency of the document, and λ is the time decay coefficient. represents the current time, represents the file creation time, α , β and γ are adjustable weight parameters;
[0011] The retrieval results are sorted in reverse order through the relevance score calculation formula, and files with high relevance and strong timeliness are preferentially displayed.
[0012] As a further solution of the present invention: The information sharing optimization module uses an improved algorithm of the Advanced Encryption Standard (AES) integrated with quantum random number generation to encrypt the shared file information. In the initial stage of encryption, a quantum random number generator is used to generate a highly random and unpredictable initial key. During the encryption process, first, the file information is grouped, with each group being 128 bits. Then, multiple round keys are generated by combining the quantum key and the traditional key expansion algorithm. The number of encryption rounds is dynamically adjusted according to the key length and the security level of the file. For a 128-bit key, 15 rounds of encryption are performed; for a 192-bit key, 18 rounds of encryption are performed; for a 256-bit key, 20 rounds of encryption are performed. In each round, byte substitution, row shift, column confusion, and round key addition operations are performed in sequence to ensure the confidentiality and integrity of the file information during sharing, and it is determined that only users with the correct key can restore the original file information through the corresponding decryption algorithm, thereby achieving secure and reliable information sharing.
[0013] As a further solution of the present invention: The deep learning engine uses the Stochastic Gradient Descent (SGD) algorithm to train the multi-layer convolutional neural network. First, the network parameters are initialized, including the weights and biases of the convolutional kernels. In each iteration, a small batch of file data is randomly selected as the training sample, and the gradient of the loss function with respect to the network parameters is calculated. Here, the entropy loss function is used as the loss function, and the specific loss function formula is as follows:
[0014]
[0015] Where: represents the loss value, represents the true label, represents the predicted output;
[0016] Next, the parameters are updated according to the learning rate, and the specific formula is as follows:
[0017]
[0018] Where: the updated network parameters, the current network parameters, the gradient of the loss function;
[0019] Through continuous iteration, the network gradually converges, enabling accurate feature extraction and pattern recognition of archival data.
[0020] As a further solution of the present invention: The adaptive classification module dynamically adjusts the classification strategy by calculating the information gain value of the archives. First, for each attribute, calculate its information entropy under the current classification , where represents the probability of a certain attribute value appearing. Then, calculate the information gain of this attribute for classification. The specific calculation formula is as follows:
[0021]
[0022] Information gain, D represents the data set, represents the attribute A takes the value of v when the data subset, represents the data set of the information entropy, represents the data subset of the information entropy;
[0023] For the calculated information gain, the greater it is, the greater the contribution of this attribute to classification. Based on the size of the information gain value, dynamically adjust the classification level and category division of the archives.
[0024] As a further solution of the present invention: The data cleaning module uses data standardization technology to clean and normalize the original archival data. For numerical data, the Z-score standardization method is adopted to convert the data into a distribution with a mean of 0 and a standard deviation of 1. The calculation formula of Z-score standardization is as follows:
[0025]
[0026] Where: represents the standardized value, represents the original data value, represents the data mean, represents the data standard deviation;
[0027] For text data, through stemming and stop word filtering operations, redundant and irrelevant information is removed.
[0028] As a further solution of the present invention: the intelligent retrieval module also has an automatic error correction function. When the search request input by the user contains spelling errors and semantic ambiguity, the similarity with the correct search term is calculated through the edit distance algorithm. First, the character string input by the user is converted into a character array, and then the edit distance between the two arrays is calculated by the dynamic programming method. The editing operations include inserting, deleting and replacing characters. According to the size of the edit distance, the similarity with the possible correct search term is judged, and automatic error correction and optimization are performed. At the same time, the context information and historical search records are combined to further improve the accuracy of error correction.
[0029] As a further solution of the present invention: the information sharing optimization module supports attribute-based access control (ABAC) strategy. The information sharing optimization module dynamically allocates access rights according to user attributes, environment attributes, operation attributes and object attributes. First, the access control policy rules are defined. Then, when the user initiates an access request, the relevant attribute information is extracted, and these attributes are evaluated and matched using a rule engine. If the match is successful, the corresponding access rights are granted, otherwise, access is denied.
[0030] As a further solution of the present invention: When performing feature extraction, the deep learning engine uses a local binary pattern (LBP) algorithm to extract features from archival image data. First, the image is divided into several small areas. The central pixel of each area is compared with its neighborhood pixels. If the neighborhood pixel value is greater than the central pixel value, it is marked as 1, otherwise, it is marked as 0. These marked values are combined into a binary number in a certain order, and converted into a decimal number as the LBP value of the area. Then, the frequency of occurrence of different LBP values in the image is counted to form a feature vector. In this way, the detailed features of the archival image data can be effectively captured.
[0031] In addition, the present invention also provides an artificial intelligence-based archive method, which comprises the following steps:
[0032] Step 1: Use the word vector model in the intelligent retrieval module. This model is optimized and trained by a deep neural network to convert the search request and archive text entered by the user into a vector form, helping to accurately capture the semantic features of the text and provide a basis for subsequent relevance calculations;
[0033] Step 2: Calculate the semantic similarity between the search request and the archive document through cosine similarity. In the calculation process, further analyze and optimize the vectors in combination with natural language processing technology to improve the accuracy of semantic similarity calculation;
[0034] Step 3: Introduce a time decay factor to reflect the timeliness of the archives, assign higher weights to newer archives, consider the time characteristics of the archives, and make the retrieval results more in line with the user's needs for the latest information. By dynamically adjusting the time decay coefficient, the influence degree of timeliness on the relevance score can be flexibly controlled according to different archive types and application scenarios;
[0035] Step 4: Calculate the relevance score between the retrieval request and the archive text according to the relevance calculation formula Sort the retrieval results in reverse order through the relevance score calculation formula, and preferentially display the archives with high relevance and strong timeliness to provide the most suitable retrieval results for users.
[0036] Adopting the above technical solution, compared with the prior art, the beneficial effects of the present invention are as follows:
[0037] 1. Through the intelligent retrieval module using natural language processing technology, the present invention can deeply understand the semantics of the retrieval request input by the user. The deep neural network further strengthens the analysis ability of complex semantic relationships. Combining with the word vector model, the retrieval request and the archive text are transformed into high-dimensional vectors. Through accurate cosine similarity calculation and the relevance calculation formula introducing the time decay factor, the most suitable archives can be quickly and accurately screened out from the massive archive database. Even in a large-scale archive room, the deep learning engine can continuously optimize the feature extraction and pattern recognition capabilities through the training of a large amount of archive data, greatly improving the accuracy and efficiency of retrieval, saving a large amount of time and reducing the possibility of errors;
[0038] 2. Through the information sharing optimization module, the present invention adopts an improved algorithm of the advanced encryption standard integrating quantum random number generation, uses a quantum random number generator to generate highly random and unpredictable initial keys to ensure the security of archive information during the sharing process. During the encryption process, the archive information is grouped, and multiple round keys are generated by combining the quantum key and the traditional key expansion algorithm, and the encryption round number is dynamically adjusted according to the key length and the security level of the archive. At the same time, the intelligent retrieval module and the information sharing optimization module work together to enable the archive information to flow efficiently between different users and systems. The data cleaning module provides a high-quality data foundation for information sharing, truly realizing the service function of the information, promoting the collaboration and communication between departments, and improving the overall work efficiency;
[0039] 3. The present invention can classify archives more accurately by calculating the information gain value according to the archive attributes through the adaptive classification module and dynamically adjusting the classification strategy. Combining the rich feature information extracted by the deep learning engine, when faced with a large number of archives with various types, this dynamic classification method ensures that the archives are always in a reasonable classification system, facilitating users to quickly locate the required archives. At the same time, the deep learning engine uses the local binary pattern algorithm to process the archive image data during feature extraction, effectively capturing the detailed features of the image, further enriching the feature information of the archives, providing strong support for the accurate classification and efficient retrieval of archives, and greatly improving the overall efficiency and service quality of archive management. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of the system flow in the embodiment of the present invention;
[0041] Figure 2 It is a schematic diagram of the method steps in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following further describes the specific embodiments of the present invention with reference to the drawings. It should be noted here that the description of these embodiments is for helping to understand the present invention, but does not constitute a limitation to the present invention.
[0043] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0044] Embodiment 1:
[0045] With the rapid development of information technology, the importance of archive management has become increasingly prominent. In various fields, a large amount of information is recorded and saved as archive materials, covering various forms such as text, images, audio, and video. Traditional archive management mainly relies on manual operations and simple computer assistance. However, with the continuous increase in the number of archives and the increasing variety, this method gradually fails to meet the requirements of modern society for efficient and accurate archive management;
[0046] There are certain defects in an existing file system and method based on artificial intelligence. On the one hand, file retrieval is difficult. Facing a large number and various types of file materials, manual retrieval is difficult, error-prone, and time-consuming. Especially when the scale of the file room is large, this problem is more prominent. On the other hand, in terms of file information sharing, some existing computer management systems only regard it as a tool for storing materials, and the information sharing performance is poor. The lack of an effective sharing mechanism makes file information unable to flow efficiently between different departments and users, and cannot truly play the service function of the information. This not only hinders the improvement of work efficiency but also limits the full utilization of information resources. Therefore, we propose a file system and method based on artificial intelligence;
[0047] Therefore, in order to effectively solve the above problems, this application proposes a file system and method based on artificial intelligence, as shown in the attached drawings of the specification Figure 1 - Figure 2 which includes an intelligent retrieval module, an information sharing optimization module, a deep learning engine, and an adaptive classification module;
[0048] The intelligent retrieval module uses natural language processing and deep neural networks, combines the word vector model to convert the retrieval request and file text into vector form, and realizes accurate and timely retrieval through the cosine similarity and the correlation calculation formula introducing a time decay factor. The retrieved relevant file data will be passed as input to the deep learning engine and the adaptive classification module. The information sharing optimization module adopts an improved algorithm of the advanced encryption standard integrating quantum random number generation, uses a quantum random number generator to generate an initial key, encrypts the file information in groups and dynamically adjusts the encryption rounds according to the key length and security level to ensure the secure sharing of information. The deep learning engine uses the stochastic gradient descent algorithm to train the multi-layer convolutional neural network, and realizes the accurate feature extraction and pattern recognition of file data by calculating the gradient of the loss function and updating the network parameters. After obtaining the file data from the intelligent retrieval module, the deep learning engine performs feature extraction. The extracted features can be fed back to the intelligent retrieval module to optimize the retrieval results, and at the same time can also provide richer feature information for the adaptive classification module for classification. The adaptive classification module dynamically adjusts the classification strategy by calculating the information gain value according to the file attributes. The adaptive classification module receives the feature information extracted by the deep learning engine and the file data provided by the intelligent retrieval module, and comprehensively considers these factors for classification adjustment, making the file classification more reasonable and efficient. At the same time, the classification result can also provide a more accurate retrieval range and direction for the intelligent retrieval module;
[0049] The file system also includes a data cleaning module, which uses the Z-score standardization method for numerical data and performs stemming and stop word filtering on text data;
[0050] Specific workflow: Utilize the word vector model in the intelligent retrieval module. The word vector model is optimized and trained through a deep neural network, converting the retrieval request input by the user and the archive text into vector form, helping to accurately capture the semantic features of the text and providing a basis for subsequent relevance calculation. Calculate the semantic similarity between the retrieval request and the archive document through cosine similarity. During the calculation process, combine natural language processing technology to further analyze and optimize the vector, improving the accuracy of semantic similarity calculation;
[0051] Furthermore, through the intelligent retrieval module using natural language processing technology, it can deeply understand the semantics of the retrieval request input by the user. The deep neural network further enhances the analysis ability of complex semantic relationships. Combining with the word vector model, the retrieval request and the archive text are converted into high-dimensional vectors. Through precise cosine similarity calculation and a relevance calculation formula introducing a time decay factor, the most suitable archives can be quickly and accurately screened out from a large number of archives. Even in a large-scale archive room, the deep learning engine can continuously optimize the feature extraction and pattern recognition capabilities through training on a large amount of archive data, greatly improving the accuracy and efficiency of retrieval, saving a large amount of time and reducing the possibility of errors.
[0052] Embodiment 2:
[0053] Based on Embodiment 1, as shown in the accompanying drawings of the specification Figure 1 The intelligent retrieval module is provided with a word vector model. Use the word vector model to convert the retrieval request and the archive text into vector form, calculate the relevance between the retrieval request and the archive text through cosine similarity. At the same time, introduce a time decay factor to reflect the timeliness of the archive, giving higher weights to newer archives, and calculate the relevance score between the retrieval request and the archive text through the relevance calculation formula. The specific relevance score calculation formula is as follows:
[0054]
[0055] Where: Represents the relevance score, Represents the retrieval request, Archive document, Represents the retrieval request And the archive document Semantic similarity, Represents the word frequency of the retrieval term in the document, Represents the update frequency of the document, λ is the time decay coefficient, Represents the current time, Represents the archive creation time, α 、 β And γ Are adjustable weight parameters;
[0056] The retrieval results are sorted in reverse order through the relevance score calculation formula, and the files with high relevance and strong timeliness are preferentially displayed. The information sharing optimization module uses an improved algorithm of the Advanced Encryption Standard (AES) integrated with quantum random number generation to encrypt the shared file information. In the initial stage of encryption, a quantum random number generator is used to generate a highly random and unpredictable initial key. During the encryption process, first, the file information is grouped, with each group being 128 bits. Then, multiple round keys are generated by combining the quantum key and the traditional key expansion algorithm. The number of encryption rounds is dynamically adjusted according to the key length and the security level of the file. For a 128-bit key, 15 rounds of encryption are performed; for a 192-bit key, 18 rounds of encryption are performed; and for a 256-bit key, 20 rounds of encryption are performed. In each round, byte substitution, row shift, column confusion, and round key addition operations are sequentially performed to ensure the confidentiality and integrity of the file information during sharing, and it is determined that only users with the correct key can restore the original file information through the corresponding decryption algorithm, thus achieving secure and reliable information sharing. The deep learning engine uses the Stochastic Gradient Descent (SGD) algorithm to train a multi-layer convolutional neural network. First, the network parameters are initialized, including the weights and biases of the convolutional kernels. In each iteration, a small batch of file data is randomly selected as the training sample, and the gradient of the loss function with respect to the network parameters is calculated. Here, the entropy loss function is used as the loss function, and the specific loss function formula is as follows:
[0057]
[0058] where: represents the loss value, represents the true label, represents the predicted output;
[0059] Then, the parameters are updated according to the learning rate, and the specific formula is as follows:
[0060]
[0061] where: are the updated network parameters, are the current network parameters, is the gradient of the loss function;
[0062] Through continuous iteration, the network gradually converges and can accurately extract features and recognize patterns from file data. The adaptive classification module dynamically adjusts the classification strategy by calculating the information gain value of the file. First, for each attribute, its information entropy under the current classification is calculated , where represents the probability of a certain attribute value occurring. Then, the information gain of this attribute for classification is calculated, and the specific calculation formula is as follows:
[0063]
[0064] Information gain D represents a data set represents an attribute A takes a value of v when the data subset represents a data set the information entropy of represents a data subset the information entropy of;
[0065] For the calculated information gain, the greater it is, the greater the contribution of the attribute to classification. Based on the magnitude of the information gain value, dynamically adjust the classification hierarchy and category division of the archives;
[0066] Specific workflow: For the information sharing optimization module, use the Advanced Encryption Standard (AES) algorithm to encrypt the shared archive information. During the encryption process, first group the archive information, with each group being 128 bits. Assume the original archive information is a binary data segment "10101100 01110011 11010100 10001011", divide it into 128-bit groups. Then, generate multiple round keys through the key expansion algorithm. The number of encryption rounds depends on the key length. For a 128-bit key, perform 10 rounds of encryption; for a 192-bit key, perform 12 rounds of encryption; for a 256-bit key, perform 14 rounds of encryption. Taking a 128-bit key as an example, assume the key is "01010101 10101010 11110000 00001111". In each round of encryption, first perform the byte substitution operation, that is, replace the input byte with the corresponding output byte through a predefined lookup table. For example, replace "10101100" with "01100101" to change the data distribution. Then perform the row shift operation, cyclically shift the bytes in each row of the matrix by a specific offset. Assume the first row is not shifted, the second row is shifted left by one bit, the third row is shifted left by two bits, and the fourth row is shifted left by three bits to increase the data confusion. Then perform the column mixing operation, process the bytes in each column through a linear transformation so that the change of a single byte can affect the data of the entire column. Finally, perform the round key addition operation, perform a bitwise exclusive OR operation on the current round key and the data processed previously. For example, in the first round, perform an exclusive OR operation on the processed data "0110010100111001 10010110 01001001" and the round key "01010101 10101010 11110000 00001111" to further enhance the security of encryption. Only users with the correct key can restore the original archive information through the corresponding decryption algorithm, thus achieving secure and reliable information sharing;
[0067] Furthermore, the information sharing optimization module adopts an improved algorithm of the Advanced Encryption Standard (AES) integrated with quantum random number generation. The quantum random number generator is used to generate highly random and unpredictable initial keys to ensure the security of archival information during the sharing process. During the encryption process, the archival information is grouped, and multiple round keys are generated by combining quantum keys and traditional key expansion algorithms. The number of encryption rounds is dynamically adjusted according to the key length and the security level of the archives. At the same time, the intelligent retrieval module and the information sharing optimization module work together to enable the efficient transfer of archival information between different users and systems. The data cleaning module provides a high-quality data foundation for information sharing, truly realizing the service function of information, promoting collaboration and communication among departments, and improving the overall work efficiency.
[0068] Embodiment 3:
[0069] Based on Embodiment 2, as shown in the accompanying drawings of the specification Figure 1 - Figure 2 The data cleaning module uses data standardization technology to clean and standardize the original archival data. For numerical data, the Z-score standardization method is adopted to convert the data into a distribution with a mean of 0 and a standard deviation of 1. The calculation formula for Z-score standardization is as follows:
[0070]
[0071] Where: represents the standardized value, represents the original data value, represents the data mean, represents the data standard deviation;
[0072] For text data, redundant and irrelevant information is removed through stemming and stop-word filtering operations. The intelligent retrieval module also has an automatic error correction function. When there are spelling mistakes and semantic ambiguities in the retrieval request entered by the user, the similarity with the correct retrieval term is calculated through the edit distance algorithm. First, the string entered by the user is converted into a character array, and then the edit distance between the two arrays is calculated through dynamic programming. The edit operations include inserting, deleting, and replacing characters. According to the size of the edit distance, the similarity with the possible correct retrieval term is judged, and automatic error correction and optimization are performed. At the same time, combined with context information and historical retrieval records, the accuracy of error correction is further improved. The information sharing optimization module supports the Attribute-Based Access Control (ABAC) policy. The information sharing optimization module dynamically assigns access permissions according to the user's attributes, environmental attributes, operation attributes, and object attributes. First, access control policy rules are defined. Then, when the user initiates an access request, relevant attribute information is extracted, and a rule engine is used to evaluate and match these attributes. If the match is successful, the corresponding access permission is granted; otherwise, access is denied. When the deep learning engine performs feature extraction, the Local Binary Pattern (LBP) algorithm is used to extract features from the archival image data. First, the image is divided into several small regions. For the central pixel of each region, it is compared with its neighboring pixels. If the neighboring pixel value is greater than the central pixel value, it is marked as 1; otherwise, it is marked as 0. These marked values are combined into a binary number in a certain order and converted into a decimal number as the LBP value of the region. Then, the frequencies of different LBP values appearing in the image are counted to form a feature vector. In this way, the detailed features of the archival image data can be effectively captured;
[0073] Specific workflow: For the intelligent retrieval module, when there are spelling mistakes or semantic ambiguities in the retrieval request entered by the user, the similarity with the correct retrieval term is calculated through the edit distance algorithm. First, the string entered by the user is converted into a character array. For example, if the user enters "cmoputer", it is converted into the character array ['c','','o','p','u','t','e','r']. Then, the edit distance between the two arrays is calculated through the method of dynamic programming. The edit operations include inserting, deleting, and replacing characters. Suppose the correct retrieval term is "computer", which is converted into the character array ['c','o','','p','u','t','e','r']. A two-dimensional array is created to calculate the edit distance. The initial state is that the first row and the first column respectively represent the number of operations from an empty string to the target string. Then, the two character arrays are compared bit by bit. According to whether the current characters are the same and the previous calculation results, the value of the edit distance is updated. For example, for the first character 'c' which is the same, no operation is required. For the second character '' and 'o' which are different, calculate the minimum cost of replacement, insertion, or deletion. According to the size of the edit distance, judge the similarity with the possible correct retrieval term. The smaller the edit distance, the higher the similarity, and automatic error correction and optimization are performed. At the same time, combined with context information and historical retrieval records, the accuracy and intelligence of error correction are further improved. For example, if the user often actually wants to retrieve files related to "computer" when entering "cmoputer", the system will correct it to "computer" according to the historical record first for retrieval to improve the retrieval success rate and user experience. For example, in an academic literature archive system, when the user enters "neorlogy", the system finds that it has a high similarity with "neurology" through calculating the edit distance, automatically corrects it to "neurology" and displays the relevant literature archives to help the user quickly obtain the required information;
[0074] Furthermore, through the adaptive classification module, the information gain value is calculated based on the archive attributes to dynamically adjust the classification strategy. Combining the rich feature information extracted by the deep learning engine, the archives can be classified more accurately. When facing a large number and diverse types of archive materials, this dynamic classification method ensures that the archives are always in a reasonable classification system, facilitating users to quickly locate the required archives. At the same time, when extracting features, the deep learning engine uses the local binary pattern algorithm to process the archive image data, effectively capturing the detailed features of the image, further enriching the feature information of the archives, providing strong support for the accurate classification and efficient retrieval of archives, and greatly improving the overall efficiency and service quality of archive management.
[0075] Working principle:
[0076] First, the word vector model in the intelligent retrieval module is used. This model has been optimized and trained by deep neural networks to convert the retrieval requests and archive texts entered by users into vector form, which helps to accurately capture the semantic features of the text and provide a basis for subsequent relevance calculations. Then, the semantic similarity between the retrieval request and the archive document is calculated by cosine similarity. During the calculation process, the vector is further analyzed and optimized in combination with natural language processing technology to improve the accuracy of semantic similarity calculations. A time decay factor is introduced to reflect the timeliness of the archives, giving higher weights to newer archives. Considering the time characteristics of the archives, the retrieval results are more in line with the user's demand for the latest information. By dynamically adjusting the time decay coefficient, the degree of influence of timeliness on the relevance score can be flexibly controlled according to different archive types and application scenarios. According to the relevance calculation formula Calculate the relevance scores between the search request and the archive text, sort the search results in reverse order using the relevance score calculation formula, and give priority to displaying archives with high relevance and timeliness, providing users with the search results that best meet their needs. At this point, the entire workflow ends.
[0077] At the same time, the present application uses specific words to describe the embodiments of the present application. For example, "one embodiment", "an embodiment", and / or "some embodiments" refer to a certain feature, structure or characteristic related to at least one embodiment of the present application. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more in different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures or characteristics in one or more embodiments of the present application can be appropriately combined.
[0078] Some aspects of the present application may be performed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software may be referred to as "data blocks", "modules", "engines", "units", "components" or "systems". The processor may be one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, or combinations thereof. In addition, various aspects of the present application may be represented as computer products located in one or more computer-readable media, which include computer-readable program codes. For example, computer-readable media may include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, tapes ...), optical disks (e.g., compact disks CDs, digital versatile disks DVDs ...), smart cards, and flash memory devices (e.g., cards, sticks, key drives ...).
[0079] A computer-readable medium may include a propagated data signal that contains computer program code, for example, on a baseband or as part of a carrier wave. The propagated signal may have various manifestations, including electromagnetic form, optical form, etc., or a suitable combination form. A computer-readable medium can be any computer-readable medium other than a computer-readable storage medium, and this medium can be connected to an instruction execution system, apparatus, or device to enable communication, propagation, or transmission of a program for use. The program code located on the computer-readable medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, radio frequency signal, or similar media, or any combination of the above media.
[0080] Similarly, it should be noted that, in order to simplify the description of the disclosure of this application and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of this application, sometimes multiple features are merged into one embodiment, drawing, or description thereof. However, this disclosure method does not mean that the features required by the subject matter of this application are more than those mentioned in the claims. In fact, the features of the embodiments are fewer than all the features of the single embodiment disclosed above. In some embodiments, numbers are used to describe components and the quantity of attributes. It should be understood that such numbers used for the description of the embodiments are, in some examples, modified by the modifiers "about", "approximate", or "substantially". Unless otherwise stated, "about", "approximate", or "substantially" indicate that the number allows a ±20% variation. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, and this approximate value can change according to the characteristics required by individual embodiments. In some embodiments, the numerical parameters should consider the specified significant digits and adopt the method of retaining the general number of digits. Although the numerical ranges and parameters used to confirm the breadth of the scope in some embodiments of this application are approximate values, in specific embodiments, such numerical settings are as precise as possible within the feasible range.
[0081] Although the present invention is disclosed above in preferred embodiments, it is not used to limit the present invention. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, all modifications, equivalent changes, and decorations made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention fall within the protection scope defined by the claims of the present invention.
Claims
1. An artificial intelligence-based archive system, characterized by: The archive system includes an intelligent retrieval module, an information sharing optimization module, a deep learning engine and an adaptive classification module; The intelligent retrieval module uses natural language processing and deep neural networks, combined with word vector models, to convert retrieval requests and archive texts into vector forms, and achieves accurate and timely retrieval through cosine similarity and the correlation calculation formula that introduces the time decay factor. The retrieved relevant archive data will be passed as input to the deep learning engine and the adaptive classification module. The information sharing optimization module adopts the advanced encryption standard improved algorithm that integrates quantum random number generation, uses the quantum random number generator to generate the initial key, encrypts the archive information in groups, and dynamically adjusts the number of encryption rounds according to the key length and security level to ensure the safe sharing of information. The deep learning engine uses the stochastic gradient descent algorithm to train the multi-layer convolutional neural network, and calculates the loss. The loss function gradient and updated network parameters realize accurate feature extraction and pattern recognition of archival data. The deep learning engine extracts features after acquiring archival data from the intelligent retrieval module. The extracted features can be fed back to the intelligent retrieval module to optimize the retrieval results. At the same time, it can also provide more abundant feature information for classification for the adaptive classification module. The adaptive classification module dynamically adjusts the classification strategy by calculating the information gain value according to the archival attributes. The adaptive classification module receives the feature information extracted by the deep learning engine and the archival data provided by the intelligent retrieval module, and comprehensively considers these factors to make classification adjustments, so that the archival classification is more reasonable and efficient. At the same time, the classification results can also provide a more accurate retrieval scope and direction for the intelligent retrieval module. The archive system also includes a data cleaning module, which uses a Z-score standardization method for numerical data and performs stem extraction and stop word filtering on text data; The intelligent retrieval module is provided with a word vector model, which is used to convert the retrieval request and the archive text into a vector form, and the correlation between the retrieval request and the archive text is calculated by cosine similarity. At the same time, a time decay factor is introduced to reflect the timeliness of the archive, and a higher weight is given to newer archives. The correlation score between the retrieval request and the archive text is calculated by a correlation calculation formula. The specific correlation score calculation formula is as follows: in: represents the relevance score, Represents a retrieval request. Archival documents, Represents a search request With archive documents The semantic similarity of Indicates the frequency of the search term in the document. represents the update frequency of the document, λ is the time decay coefficient, Indicates the current time. Indicates the time when the file was created. α , β and γ is an adjustable weight parameter; The search results are sorted in reverse order using the relevance score calculation formula, giving priority to displaying archives with high relevance and strong timeliness The information sharing optimization module uses an improved Advanced Encryption Standard (AES) algorithm that incorporates quantum random number generation to encrypt the shared archive information. In the initial stage of encryption, a quantum random number generator is used to generate a highly random and unpredictable initial key. During the encryption process, the archive information is first grouped into 128 bits per group. Then, multiple round keys are generated by combining quantum keys and traditional key expansion algorithms. The number of encryption rounds is dynamically adjusted according to the key length and the security level of the archive. A 128-bit key performs 15 rounds of encryption, a 192-bit key performs 18 rounds of encryption, and a 256-bit key performs 20 rounds of encryption. In each round, byte replacement, row shift, column confusion, and round key addition operations are performed in sequence to ensure the confidentiality and integrity of the archive information during the sharing process, and to ensure that only users with the correct key can restore the original archive information through the corresponding decryption algorithm, thereby achieving safe and reliable information sharing. The adaptive classification module dynamically adjusts the classification strategy by calculating the information gain value of the archive. First, for each attribute, its information entropy under the current classification is calculated. ,in Represents the probability of a certain attribute value appearing, and then calculates the information gain of the attribute for classification. The specific calculation formula is as follows: Information gain, D represents a data set, Representation attributes A The value is v The data subset at time Representation dataset The information entropy of Representing a subset of data Information entropy of The larger the calculated information gain is, the greater the contribution of the attribute to the classification. The classification level and category division of the archives are dynamically adjusted according to the size of the information gain value.
2. The artificial intelligence-based archive system according to claim 1, characterized in that: The deep learning engine uses the stochastic gradient descent (SGD) algorithm to train the multi-layer convolutional neural network. First, the network parameters, including the weights and biases of the convolution kernels, are initialized. In each iteration, a small batch of archive data is randomly selected as training samples, and the gradient of the loss function to the network parameters is calculated. The loss function here uses the entropy loss function. The specific loss function formula is as follows: in: represents the loss value, represents the true label, Represents the predicted output; Next, the parameters are updated according to the learning rate. The specific formula is as follows: in: Updated network parameters, Current network parameters, The gradient of the loss function; Through continuous iteration, the network gradually converges and can accurately extract features and recognize patterns of archival data.
3. The artificial intelligence-based archive system according to claim 1, characterized in that: The data cleaning module uses data standardization technology to clean and normalize the original archive data. For numerical data, the Z-score standardization method is used to convert the data into a distribution with a mean of 0 and a standard deviation of 1. The calculation formula for Z-score standardization is as follows: in: represents the normalized value, Represents the original data value, represents the data mean, represents the standard deviation of the data; For text data, redundant and irrelevant information is removed through stemming and stop word filtering operations.
4. The artificial intelligence-based archive system according to claim 1, characterized in that: The intelligent retrieval module also has an automatic error correction function. When the search request input by the user contains spelling errors and semantic ambiguity, the similarity with the correct search term is calculated through the edit distance algorithm. First, the string input by the user is converted into a character array, and then the edit distance between the two arrays is calculated through the dynamic programming method. The editing operations include inserting, deleting and replacing characters. According to the size of the edit distance, the similarity with the possible correct search term is judged, and automatic error correction and optimization are performed. At the same time, the context information and historical search records are combined to further improve the accuracy of error correction.
5. The artificial intelligence-based archive system according to claim 1, characterized in that: The information sharing optimization module supports attribute-based access control (ABAC) strategy. The information sharing optimization module dynamically allocates access rights based on user attributes, environmental attributes, operation attributes and object attributes. First, the access control policy rules are defined. Then, when the user initiates an access request, the relevant attribute information is extracted and the rule engine is used to evaluate and match these attributes. If the match is successful, the corresponding access rights are granted, otherwise, access is denied.
6. The artificial intelligence-based archive system according to claim 1, characterized in that: When performing feature extraction, the deep learning engine uses the local binary pattern (LBP) algorithm to extract features from the archival image data. First, the image is divided into several small areas. The central pixel of each area is compared with its neighborhood pixels. If the neighborhood pixel value is greater than the central pixel value, it is marked as 1, otherwise, it is marked as 0. These marked values are combined into a binary number in a certain order and converted into a decimal number as the LBP value of the area. Then, the frequency of occurrence of different LBP values in the image is counted to form a feature vector. In this way, the detailed features of the archival image data can be effectively captured.
7. An artificial intelligence-based archival method applicable to the artificial intelligence-based archival system of any one of claims 1 to 6, characterized in that: The archival method comprises the following steps: Step 1: Use the word vector model in the intelligent retrieval module. This model is optimized and trained by a deep neural network to convert the search request and archive text entered by the user into a vector form, helping to accurately capture the semantic features of the text and provide a basis for subsequent relevance calculations; Step 2: Calculate the semantic similarity between the search request and the archive document through cosine similarity. In the calculation process, further analyze and optimize the vectors in combination with natural language processing technology to improve the accuracy of semantic similarity calculation; Step 3: Introduce a time decay factor to reflect the timeliness of archives, give higher weights to newer archives, and consider the time characteristics of archives, so that the search results are more in line with the user's needs for the latest information. By dynamically adjusting the time decay coefficient, the influence of timeliness on the relevance score can be flexibly controlled according to different archive types and application scenarios; Step 4: Calculate the correlation formula Calculate the relevance scores between the search request and the archive text, sort the search results in reverse order using the relevance score calculation formula, give priority to displaying archives with high relevance and timeliness, and provide users with the search results that best meet their needs.
Citation Information
Patent Citations
Power file question and answer type intelligent retrieval method and system
CN117171333A
Archive management system based on artificial intelligence
CN118245652A