Accurate archive information checking method and system based on cloud architecture

By adopting distributed storage and semantic analysis technologies in the cloud architecture, the problems of accuracy and efficiency in archival information verification are solved, and efficient and accurate archival information management and inspection are achieved.

CN120104697APending Publication Date: 2025-06-06SHANDONG QIANHE ARCHIVES MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510179629.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing technology has defects in the accuracy and efficiency of archival information verification, and it is impossible to fully utilize the advantages of cloud architecture in archival information management, especially in processing massive data and semantic understanding.

Method used

A cloud-based architecture method is adopted to obtain, preprocess and distributed storage of archive information data, combine semantic analysis and machine learning algorithms to understand the verification keywords entered by users, and provide verification results with different accuracy according to user needs.

Benefits of technology

It improves the accuracy and efficiency of archival information verification, can flexibly match accurate and fuzzy queries, adapt to different needs, shortens the inspection time, and ensures the security and confidentiality of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104697A_ABST
    Figure CN120104697A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of archive information management, in particular to an archive information accurate checking method and system based on a cloud architecture. The method comprises the following steps: acquiring archive information data; preprocessing the acquired archive information data; storing the preprocessed archive information data on a plurality of cloud nodes in a distributed manner based on a distributed cloud storage architecture; performing semantic understanding on the examination keywords input by the user based on semantic analysis and a machine learning algorithm; and inspection results with different precisions are provided according to user requirements. Through distributed storage and elastic computing resources of the cloud architecture and an efficient checking algorithm, the checking time of the archive information is greatly shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of archive information management, and in particular to a method and system for accurately checking archive information based on a cloud architecture. Background Art

[0002] In the digital age, the importance of archival information management has become increasingly prominent. Traditional archival information verification methods mostly rely on local storage and single-machine processing modes, which gradually expose many problems when facing the growing amount of archival data.

[0003] From a storage perspective, local storage is limited by the capacity of hardware devices and cannot meet the long-term storage needs of large-scale archival data. In addition, data security depends on the stability of local hardware. Once the hardware is damaged, data loss is very likely to occur. In terms of inspection efficiency, the single-machine processing mode cannot meet the needs of quickly obtaining archival information due to limited computing resources and slow response when processing complex query requests. For example, in some large enterprises or government agencies, when a large number of historical archives need to be queried, traditional methods may take hours or even days, seriously affecting work efficiency.

[0004] With the development of cloud computing technology, the application of cloud architecture in the field of archival information management has gradually emerged. Cloud architecture has powerful storage capacity and elastic computing resources, which can easily cope with the storage and processing needs of massive data, realize remote storage and access of archival information, and break the geographical restrictions. However, the archival information verification technology based on cloud architecture still has many shortcomings. On the one hand, the existing verification algorithms are often based on simple keyword matching and lack in-depth understanding of semantics, resulting in inaccurate verification results and unable to meet the diverse query needs of users. For example, when a user queries "important recent meeting minutes", traditional algorithms may return a large number of irrelevant results because they cannot accurately understand the semantics of "recent" and "important". On the other hand, data storage and management under the cloud architecture lack effective data organization and indexing mechanisms, making it difficult to quickly locate target archives in massive data, further reducing the verification efficiency.

[0005] In summary, the existing technology has obvious defects in the accuracy and efficiency of archival information verification, and cannot give full play to the advantages of cloud architecture in archival information management. Therefore, developing an archival information accurate verification method based on cloud architecture is of great practical significance for improving the intelligent level of archival management and meeting the needs of various industries for fast and accurate verification of archival information. Summary of the invention

[0006] In order to solve the above-mentioned problems, the present invention provides a method and system for accurately checking archive information based on cloud architecture.

[0007] In the first aspect, the present invention provides a method for accurately checking archive information based on a cloud architecture, which adopts the following technical solutions:

[0008] A method for accurately checking archive information based on cloud architecture, comprising:

[0009] Obtain archival information data;

[0010] Preprocess the acquired archival information data;

[0011] Based on the distributed cloud storage architecture, the pre-processed archive information data is distributed and stored on multiple cloud nodes;

[0012] Semantic understanding of the query keywords entered by users based on semantic analysis and machine learning algorithms;

[0013] Provide inspection results with different accuracy according to user needs.

[0014] Furthermore, the acquisition of archival information data includes converting paper data into electronic format based on data scanning and OCR technology; acquiring electronic archival information in the business system according to predetermined rules and frequencies based on docking with various business system databases; and collecting information from archival information source websites based on web crawler methods.

[0015] Furthermore, the preprocessing of the acquired archival information data includes unifying the file format and data encoding of the acquired archival information data, and extracting key metadata information based on natural language processing metadata.

[0016] Furthermore, the preprocessed archival information data is distributedly stored on multiple cloud nodes based on the distributed cloud storage architecture, including dividing the preprocessed archival information data into data blocks according to logical rules, calculating the hash value of each data block through a hash function; determining the cloud node of each data block using a load balancing-based node selection algorithm; and finally transmitting the data block to the selected cloud node through the network.

[0017] Further, the transmitting of the data block to the selected cloud node through the network includes copying each data block to multiple different cloud nodes based on a three-copy redundancy strategy, wherein, in another

[0018] Furthermore, the semantic analysis and machine learning algorithm are used to perform semantic understanding on the verification keywords input by the user, including segmenting the verification keywords input by the user to obtain independent word blocks, and converting the independent word blocks into feature vectors, using the trained machine learning model to perform semantic understanding on the feature vectors, and outputting the semantic representation and semantic label of the keywords.

[0019] Furthermore, the provision of verification results of different precisions according to user needs includes reasoning on semantic extension based on the output results of the machine learning model, combining the fuzzy query algorithm and the precise query algorithm, and providing verification results of different precisions according to user needs, wherein a decision formula is constructed by calculating the cosine similarity of the keyword vector and the data element vector and setting a dynamic adjustment factor, and a judgment rule is used to match the user's keyword query method, and the decision formula is expressed as:

[0020]

[0021] , where the confidence of the exact semantics is P precise , the confidence of fuzzy semantics is P vague , keyword vector and the archive metadata vector in the archive Cosine similarity of α is the dynamic adjustment factor.

[0022] Second, a cloud-based archival information accurate inspection system, including:

[0023] The data acquisition module is configured to acquire archive information data;

[0024] A preprocessing module is configured to preprocess the acquired archive information data;

[0025] The storage module is configured to distribute and store the pre-processed archive information data on multiple cloud nodes based on a distributed cloud storage architecture;

[0026] A semantic understanding module is configured to perform semantic understanding of the query keywords input by the user based on semantic analysis and machine learning algorithms;

[0027] The inspection module is configured to provide inspection results of different precisions according to user needs.

[0028] In a third aspect, the present invention provides a computer-readable storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor of a terminal device and executing the method for accurately checking archival information based on a cloud architecture.

[0029] In a fourth aspect, the present invention provides a terminal device comprising a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store a plurality of instructions, wherein the instructions are suitable for being loaded by the processor and executing the method for accurately verifying archival information based on a cloud architecture.

[0030] In summary, the present invention has the following beneficial technical effects:

[0031] 1. Through the above technical solution, the present invention can flexibly match precise queries and fuzzy queries according to the search terms input by the customer, and adaptively provide the search results required by the customer, rather than providing precise searches for all search questions, which is conducive to saving resources and improving efficiency, while meeting the psychological expectations of customers.

[0032] The distributed storage and elastic computing resources of the cloud architecture, as well as efficient verification algorithms, have greatly shortened the time required to verify archival information.

[0033] 2. Through the application of semantic analysis and machine learning algorithms, the inspection results are more in line with user needs and the accuracy is improved.

[0034] 3. Users can access the cloud platform to check archival information anytime and anywhere through the Internet, without geographical and time restrictions; distributed storage and permission management mechanisms ensure the security and confidentiality of archival information. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic diagram of a method for accurately checking archive information based on cloud architecture according to Example 1 of the present invention. DETAILED DESCRIPTION

[0036] The present invention is further described in detail below in conjunction with the accompanying drawings.

[0037] Example 1

[0038] Reference Figure 1 , a method for accurately checking archive information based on cloud architecture in this embodiment includes:

[0039] Obtain archival information data;

[0040] Preprocess the acquired archival information data;

[0041] Based on the distributed cloud storage architecture, the pre-processed archive information data is distributed and stored on multiple cloud nodes;

[0042] Semantic understanding of the query keywords entered by users based on semantic analysis and machine learning algorithms;

[0043] Provide inspection results with different accuracy according to user needs.

[0044] Specifically, the following steps are included:

[0045] S1. Data acquisition

[0046] Digital collection: For paper archives, professional scanning equipment is used to convert them into electronic image formats, such as PDF, JPEG, etc. At the same time, equipped with OCR (optical character recognition) technology, the text information in the scanned image is converted into editable and searchable text data for subsequent metadata extraction and index establishment.

[0047] Database connection: connect with the database of each business system, and obtain the electronic file information generated in the business system through the data interface according to certain rules and frequencies, such as the company's financial statements, personnel files, etc. During the connection process, ensure the consistency and integrity of the data to avoid data loss or duplicate collection.

[0048] Web crawler collection: For some public and valuable archival information source websites, use web crawler technology to collect information according to the set rules and scope. In the collection process, strictly abide by laws and regulations and relevant regulations of the website to avoid infringing on the rights and interests of others. After cleaning and screening, the collected data is included in the archival information management system.

[0049] User upload: Provide users with a convenient upload interface, allowing users to upload personal or unit-owned archive information to the cloud platform in accordance with the specified format and specifications. During the upload process, the file format, size, etc. are verified, and users are prompted to add necessary metadata information.

[0050] S2. Data Preprocessing

[0051] (1) Data cleaning

[0052] Remove duplicate data: By comparing the key identification information of the archival data (such as file number, title, etc.), identify and delete duplicate archival records to avoid interference from redundant information during subsequent inspections.

[0053] Processing missing values: For missing fields in archival data, we process them according to the characteristics of the data and business rules. For some important metadata fields (such as date, author, etc.), if they are missing, we can complete them by analyzing the association with other relevant archival data or supplementing them manually; for some missing non-key fields, we can use the default value filling method.

[0054] Correct erroneous data: Check for format errors, logical errors, etc. in the archive data. For example, convert the date format that does not meet the specifications; correct or delete the values ​​of numerical data that are obviously unreasonable (such as negative age).

[0055] (2) Unified format

[0056] File format unification: Unify the file formats of archives from different sources into one or several formats that are easy to store and process. Convert various image formats into JPEG or PNG; unify document formats into PDF or DOCX.

[0057] Data encoding is unified: ensure that the character encoding of archive data is consistent to avoid garbled characters caused by encoding problems. Common encoding formats can be unified as UTF-8.

[0058] (3) Metadata extraction

[0059] Automatic extraction: Use natural language processing technology and regular expressions to automatically extract key metadata information from the archive text content, such as title, date, author, subject terms, etc. Among them, the regular expression is used to match the string in the text that conforms to the date format as the archive date.

[0060] Manual supplement: For some metadata information that cannot be accurately extracted automatically, a manual interface is provided for administrators or relevant personnel to manually supplement and correct it to ensure the integrity and accuracy of the metadata.

[0061] (4) Establish an index library

[0062] Determine index fields: Based on the key information and inspection requirements of the archive, determine the fields used to create the index, such as title, date, keywords, author, etc.

[0063] Index construction: The inverted index algorithm is used to construct the index for the determined index fields. For the keyword field, the inverted index is used to record in which files each keyword appears, so as to quickly locate the relevant files.

[0064] S3. Distributed storage steps

[0065] (1) Data segmentation: The pre-processed archive information data is segmented according to certain size or logical rules. The size of each data block is set to a fixed 10MB. For large archive files, they are segmented into multiple 10MB data blocks; for structured archive data records, they are segmented according to a certain number of records.

[0066] (2) Calculate hash value: For each data block, calculate its unique hash value through a hash function (such as MD5, SHA-256, etc.). The hash value is used as the identifier of the data block for subsequent storage and data integrity verification. Taking the MD5 algorithm as an example, the steps for calculating the hash value are as follows:

[0067] Filling Data

[0068] First, pad the data to be hashed so that its length modulo 512 is 448. The specific padding method is to add a 1 after the data, and then add several 0s until the length requirement is met. Let the length of the original data be L, and the length of the padded data be L', then: L'=L+1+k, where k is the smallest non-negative integer that satisfies (L'\bmod 512)=448.

[0069] Add data length

[0070] Add 64 bits after the padded data to represent the length of the original data (in bits). The total length of the data is now an integer multiple of 512. If the length of the original data is L (in bits), the value represented by the added 64 bits is L.

[0071] Initialize the buffer

[0072] The MD5 algorithm uses four 32-bit registers (A, B, C, D) as buffers, and the initial values ​​are:

[0073] A = 0x67452301

[0074] B = 0xefcdab89

[0075] C = 0x98badcfe

[0076] D = 0x10325476

[0077] Group processing

[0078] Divide the data after padding and length addition into 512-bit (64-byte) groups, and process each group in turn. The processing of each group is as follows:

[0079] 1. Save initial values: Save the values ​​of A, B, C, and D in the current buffer as a=A, b=B, c=C, and d=D respectively.

[0080] 2. Perform four rounds of loop operations: each loop contains 16 operations, and each operation will update the buffer register according to different logical functions.

[0081] First round: The logic function is F(X,Y,Z)=(X\land Y)\lor(\neg X\land Z), and the operation order is a=b+((a+F(b,c,d)+X[i]+T[i])\ll s), where X[i] is the i-th 32-bit word in the current group (i ranges from 0 to 15), T[i] is the i-th value in the constant table, \ll s means circular left shift of s bits, and s is the displacement set according to different operations.

[0082] Second round: The logic function is G(X,Y,Z)=(X\land Z)\lor(Y\land\neg Z), and the order of operations is a=b+((a+G(b,c,d)+X[i]+T[i])\ll s), where i ranges from 0 to 15, but the order is different.

[0083] Round 3: The logical function is H(X,Y,Z)=X\oplus Y\oplus Z, and the order of operations is a=b+((a+H(b,c,d)+X[i]+T[i])\ll s), and the value and order of i have changed again.

[0084] Round 4: The logical function is I(X,Y,Z)=Y\oplus(X\lor\neg Z), and the order of operations is a=b+((a+I(b,c,d)+X[i]+T[i])\ll s), and the value and order of i change again.

[0085] 3. Update the buffer: After the four rounds of operation are completed, the result obtained by the current group processing is added to the initial value of the buffer, that is, A=A+a, B=B+b, C=C+c, D=D+d.

[0086] Output hash value

[0087] After all the groups have been processed, the values ​​of the four registers A, B, C, and D in the buffer are connected in sequence to obtain the final MD5 hash value, which is 128 bits (16 bytes) in length and is usually expressed in the form of a 32-bit hexadecimal number.

[0088] (3) Node selection algorithm: Based on the characteristics of the distributed cloud storage architecture, a consistent hashing algorithm or a node selection algorithm based on load balancing is used to determine the cloud node where each data block is stored. The consistent hashing algorithm maps the hash values ​​of cloud nodes and data blocks to a ring space, and selects the nearest cloud node for storage based on the position of the data block hash value on the ring; the load balancing algorithm monitors the storage load of each cloud node in real time and stores the data block on the node with the lightest load.

[0089] Among them, (1) defines the hash space. The consistent hashing algorithm usually uses a fixed-size hash space, which is generally expressed as a ring space ranging from 0 to 2^{n}-1, where n is the number of bits output by the hash function. For example, if the hash function outputs 32 bits, the hash space ranges from 0 to 2^{32}-1. This ring space is like a ring connected from beginning to end, and all hash values ​​will fall on this ring.

[0090] (2) Select a hash function. Select a suitable hash function, such as MD5, SHA-1, or SHA-256. The hash function should be able to map any input (including cloud node identifiers and data block identifiers) into the hash space defined in step 1. For example, if the SHA-256 hash function is used, it will output a 256-bit hash value, which needs to be mapped into the above ring space.

[0091] (3) Calculate the hash value of the cloud node. For each cloud node, use the hash function selected in step 2 to calculate its hash value. The identity of the cloud node can be the node's IP address, domain name, or other unique identifier. Assume that there are three cloud nodes, namely Node1, Node2, and Node3. Calculate their hash values ​​H_{Node1}, H_{Node2}, and H_{Node3} using the following formula:

[0092] H_{Node1}=Hash(Node1)

[0093] H_{Node2}=Hash(Node2)

[0094] H_{Node3}=Hash(Node3)

[0095] Hash represents the selected hash function. The calculated hash value will be mapped to the circular hash space defined in step 1 to form a corresponding point.

[0096] (4) Calculate the hash value of the data block, and also use the selected hash function to perform hash calculation on each data block. The identifier of the data block can be the data block number, file name, or other unique identification information. Assume that there are data blocks Data1 and Data2, calculate their hash values ​​H_{Data1} and H_{Data2}, the formula is:

[0097] H_{Data1}=Hash(Data1)

[0098] H_{Data2}=Hash(Data2)

[0099] These hash values ​​are also mapped to the ring hash space.

[0100] (5) Determine the storage node in the annular space, and search the cloud node hash value closest to the data block hash value in the annular hash space in a clockwise direction. The specific steps are as follows:

[0101] ① For data block Data1, its hash value is H_{Data1}. Starting from the point where H_{Data1} is located, search clockwise along the annular space to find the first cloud node hash value encountered. Assuming it is H_{Node2}, the data block Data1 is stored on the cloud node Node2.

[0102] ② For data block Data2, its hash value is H_{Data2}. Similarly, start searching clockwise from H_{Data2}. If the first cloud node hash value encountered is H_{Node3}, then store data block Data2 on cloud node Node3.

[0103] (6) Handling node addition and deletion

[0104] Node addition: When a new cloud node is added, the hash value of the new node is calculated and mapped to the ring space. At this time, some data blocks may need to be reallocated to storage nodes. Specifically, starting from the hash value position of the new node, the data blocks that originally belonged to other nodes in the clockwise direction are migrated to the new node if they are closer to the new node.

[0105] Node deletion: When a cloud node is removed, the data blocks stored on the node need to be reallocated. Starting from the hash value position of the node, search for the next cloud node in a clockwise direction and migrate the data blocks on the original node to the newly found cloud node.

[0106] Data transmission and storage: The divided data is transmitted to the selected cloud node through the network. During the transmission process, the encrypted transmission protocol (such as SSL / TLS) is used to ensure the security of the data. After receiving the data block, the cloud node stores it in the local storage medium and records the relevant metadata information of the data block, such as the storage path, hash value, and the archive identifier.

[0107] Redundant storage and backup: To improve data reliability, each data block is copied to multiple different cloud nodes according to a pre-set redundancy strategy. For example, a three-copy redundancy strategy is adopted to back up and store each data block on two other cloud nodes in different geographical locations. At the same time, the backup data is regularly checked for consistency to ensure the consistency of the backup data with the original data.

[0108] S4. Steps to implement semantic understanding

[0109] (1) Data preparation

[0110] Collect corpus: Collect a large amount of text data related to the archives field, including archive titles, content summaries, keywords, etc., to build a corpus dedicated to the archives field. These corpora can come from historical inspection records, existing archives, and related industry literature.

[0111] Annotated data: Manually annotate the data in the corpus and mark out key words, topics, semantic categories, etc. For example, for a text about an employee's entry file, key information such as "employee name", "entry date", "position" and their categories are annotated.

[0112] (2) Model selection and training

[0113] Model selection: Choose a machine learning model suitable for semantic analysis. The BERT model with Transformer architecture can effectively capture the semantic features in the text.

[0114] Model training: Use the annotated corpus to train the selected model. During the training process, adjust the model parameters so that the model can accurately learn the semantic information in the text. The weights of the BERT model are updated through the back-propagation algorithm so that it can accurately predict keywords and semantic categories based on the input text.

[0115] By collecting large-scale text data, including news articles, books, web page texts, etc., the WordPiece word segmentation algorithm is used to segment the text into sub-words. For example, "unhappiness" is segmented into "un", "happy" and "ness".

[0116] Add special tags: add [CLS] tags at the beginning of each sentence for classification tasks; add [SEP] tags between and at the end of sentence pairs.

[0117] Generate input features: Convert the subwords after segmentation into corresponding word embeddings, and generate segment embeddings to distinguish different sentences, and position embeddings to represent the position information of words. The final input is the sum of these three embeddings.

[0118] Define the BERT model architecture. BERT is based on the Transformer encoder architecture and consists of multiple stacked Transformer encoder layers. Each encoder layer contains a multi-head self-attention mechanism and a feed-forward neural network.

[0119] Defining pre-training tasks, BERT has two main pre-training tasks: Masked Language Model (MLM) and Next Sentence Prediction (NSP).

[0120] Masked Language Model (MLM), masking operation: randomly select some words in the input sequence and replace them with [MASK] tags (about 15% of the words). Among them, 80% are replaced with [MASK], 10% are randomly replaced with other words, and 10% remain unchanged.

[0121] Calculate loss: The model predicts the masked word and uses the cross entropy loss function to calculate the loss between the predicted result and the real word. Suppose the predicted word probability distribution is The label of the real word is y, then the MLM loss is L MLM for:

[0122]

[0123] Where N is the size of the vocabulary.

[0124] (3) Next Sentence Prediction (NSP)

[0125] Construct sentence pairs: Randomly select some sentence pairs, 50% of which are truly adjacent sentence pairs and 50% are randomly combined sentence pairs.

[0126] Calculate loss: The model determines whether the input sentence pair is adjacent, and also uses the cross entropy loss function to calculate the loss between the predicted result and the true label. Suppose the probability of whether the predicted sentence pair is adjacent is The true label is p, then the NSP loss is L NSP for:

[0127]

[0128] (4) Forward propagation calculation

[0129] Input features are passed into the model: The preprocessed input features (the sum of word embedding, segment embedding, and position embedding) are passed into the BERT model.

[0130] After the encoder layer: The input features pass through multiple stacked Transformer encoder layers in sequence. The calculation steps of each encoder layer are as follows:

[0131] Multi-head self-attention mechanism: Calculate the attention score of each head, calculate the attention weight through the query, key and value matrix, and then sum the weighted sum to get the output of each head. Finally, concatenate the outputs of all heads and use linear transformation to get the final output of the multi-head self-attention mechanism.

[0132] Feedforward neural network: The output of the multi-head self-attention mechanism is linearly transformed, activated, and transformed again to obtain the final output of the encoder layer.

[0133] Output prediction results: The model outputs predictions for masked words and whether sentence pairs are adjacent.

[0134] (5) Calculate total loss

[0135] Adding the MLM loss and the NSP loss gives the total loss L:

[0136] L=L MLM +L NSP

[0137] Among them, the weights of the BERT model are updated through the back-propagation algorithm. At the beginning of each training iteration, the gradients of all trainable parameters in the model are initialized to zero to avoid gradient accumulation.

[0138] To calculate the gradient, we use the chain rule to start from the loss function and work backwards to calculate the gradient of each trainable parameter (such as the weight matrix and bias vector) with respect to the loss, based on the total loss L. For example, for a simple linear layer y = Wx + b, where W is the weight matrix, b is the bias vector, x is the input, and y is the output. If the loss function is L, then we calculate it using the chain rule and

[0139] Back propagation process: Starting from the output layer of the model, the parameter gradients of the multi-head self-attention mechanism and the feedforward neural network in each encoder layer are calculated in reverse order. In the multi-head self-attention mechanism, the gradients of the query, key, and value matrices need to be calculated; in the feedforward neural network, the gradients of the weights and biases of the two linear transformation layers need to be calculated.

[0140] Use the optimizer to update the weights,

[0141] Update formula: Let θ be the trainable parameter of the model, is the gradient of parameter θ with respect to loss L, α is the learning rate, m t and v t They are first-order moment estimation and second-order moment estimation respectively. The update formula of Adam optimizer is as follows:

[0142] Compute the first moment estimate:

[0143]

[0144] Compute the second moment estimate:

[0145]

[0146] Correct the first moment estimate bias:

[0147]

[0148] Correct the second-order moment estimate bias:

[0149]

[0150] Update parameters:

[0151]

[0152] where β 1 and β 2 is the decay rate, usually set to 0.9 and 0.999 respectively, and ∈ is a small constant used to avoid the denominator being zero.

[0153] Repeat the training process

[0154] Repeat the above process of forward propagation, loss calculation, backpropagation, and weight update until the model converges or reaches the preset number of training rounds.

[0155] (3) Keyword processing

[0156] Word segmentation: Segment the keywords entered by the user into words, and divide the continuous text sequence into independent words or word blocks. Use the Jieba word segmentation tool to segment "2023 financial statements" into "2023" and "financial statements".

[0157] Feature extraction: Convert the segmented keywords into feature vectors that the model can process. For the BERT model, use pre-trained word vectors to map keywords into vector representations;

[0158] Among them, when the word segmentation tool WordPiece is used to segment the query keyword input by the user, a series of subwords are obtained. For the keyword "XXX company", it may be segmented into "XXX" and "company".

[0159] For each subword after segmentation, the corresponding vector representation is searched in the pre-trained word vector table. During the pre-training process, the BERT model has generated corresponding word vectors for each possible subword (from its vocabulary). By querying the weight matrix of the pre-trained word vector, each subword is mapped to a vector of fixed dimension. Assuming that the word vector dimension of the BERT model is 768 (BERT-Base model), then "someone" and "company" will be mapped to 768-dimensional vectors.

[0160] Subword vector merging: When a keyword is divided into multiple subwords, the vectors of these subwords need to be merged into a vector representing the entire keyword. The average pooling method is used to average the corresponding dimensions of multiple subword vectors. So and so Company, merged keyword vector The i-th dimension of keywords is:

[0161]

[0162] If a keyword corresponds to only one subword, the vector of the subword is directly used as the vector representation of the keyword.

[0163] Add position encoding. In order to allow the model to better capture the order information of words in keywords, position encoding can be added. When processing a text sequence, the BERT model generates a position encoding vector for each position. For a single keyword (which can be regarded as a short sequence), position encoding can be similarly added to each of its subword positions. The dimension of the position encoding vector is the same as the word vector dimension. The position encoding vector is added to the subword vector to obtain the final feature vector used to input the model. For example, for the subword "apple", its position encoding vector is Apple, the word vector is Apple, final eigenvector In this way, the model can better understand the position information of words in keywords, thereby performing more accurate semantic understanding.

[0164] (4) Semantic understanding and reasoning

[0165] Model prediction: The extracted keyword feature vector is input into the trained model. The model performs semantic understanding of the keyword based on the learned semantic knowledge and outputs the semantic representation of the keyword and related semantic tags. The model determines that "financial statements" belong to the "financial archives" category.

[0166] Semantic expansion and reasoning: Based on the output of the model, the keywords are semantically expanded and reasoned. When the user enters "sales contract", the model infers related semantic concepts such as "purchase contract" and "contract management" based on semantic associations to expand the scope of verification and improve the completeness of the search.

[0167] (5) Accurate verification algorithm

[0168] After semantic analysis and machine learning algorithms are used to understand the search keywords entered by users, fuzzy query and precise query technologies are combined to provide search results of different precisions according to user needs. For example, for precise matching keywords, relevant files are directly located; for fuzzy keywords, possible related files are screened out through semantic analysis and similarity calculation.

[0169] The confidence that the keyword output by the semantic understanding model is accurate semantics is P precise , the value range is [0,1], and the confidence level of fuzzy semantics is P vague , and P precise +P vague =1.

[0170] Calculating keyword vectors and the archive metadata vector in the archive Cosine similarity of The formula is

[0171] Set the expected similarity threshold T for precise query precise , fuzzy query expected similarity threshold T vague ,and

[0172] T precise >T vague , T precise , T vague ∈[0,1].

[0173] A dynamic adjustment factor α is introduced with a value range of [0,1]. It is dynamically adjusted according to factors such as the real-time system load and data volume. For example, when the system load is high, it tends to be more concise and efficient, and the α value increases; when the data volume is large, in order to obtain more comprehensive results, the α value decreases.

[0174] Constructing the decision formula:

[0175]

[0176] Judgment rules:

[0177]

[0178] The formula combines the confidence of semantic understanding, the similarity between keywords and archive metadata, and dynamic adjustment factors to decide whether to perform precise or fuzzy queries.

[0179] For precise query, firstly, for each pre-processed keyword, the corresponding record is searched in the index table according to the pre-built inverted index. The inverted index is a data structure that maps keywords in a document to a list of documents containing the keyword.

[0180] For the keyword "employee training", a list of all corresponding document IDs is found in the inverted index. These document IDs uniquely identify the files containing the keyword "employee training".

[0181] Result intersection operation: When querying for multiple keywords, the document ID list found in the inverted index for each keyword is intersected, because accurate query requires that all keywords must appear in the target archive at the same time.

[0182] Suppose the document ID list corresponding to keyword A is D1, D3, D5, and the document ID list corresponding to keyword B is D3, D4, D5. Then the resulting document ID list after the intersection operation is D3, D5. These documents are the exact match files that contain both keyword A and keyword B.

[0183] Archive acquisition and return: Based on the document ID obtained by the intersection operation, the corresponding archive details are obtained from the archive storage system, including the archive's title, content, creation time and other metadata.

[0184] The acquired archive information is presented to the user in a certain format and order (such as sorted by archive creation time) to complete the precise query operation.

[0185] The expanded keywords and all archive metadata in the archive (such as title, abstract, keyword field, etc.) are converted into vector representations through word vector models. Commonly used word vector models include Word2Vec and GloVe, which can map words in the text to low-dimensional vector space, so that words with similar semantics are close in the vector space.

[0186] For each expanded keyword vector, calculate its similarity with all archive metadata vectors in the archive. Use the cosine similarity method:

[0187] Computes the dot product of two vectors, that is, multiplying the elements of corresponding dimensions and then summing them.

[0188] Calculate the modulus of the two vectors respectively, that is, the square root of the sum of the squares of the elements in each dimension.

[0189] Divide the dot product by the product of the two vector moduli to get the cosine similarity value. The closer the value is to 1, the more similar the text semantics represented by the two vectors are.

[0190] Fuzzy matching rule application:

[0191] Partial matching: In addition to calculating semantic similarity, partial matching of keywords is also performed. For the keywords entered by the user or the expanded keywords, the content containing the keyword substring is searched in the archive metadata. For example, for the keyword "finance", the archive title and content are searched for archives containing the word "finance". Even if the archive contains a longer string such as "financial management" or "financial report", it is considered a match result.

[0192] Synonym matching: With the help of pre-built synonym dictionaries, keywords are replaced with their synonyms for query. For example, the synonyms of "contract" are "contract" and "agreement". When the user enters "contract", the system will search for files containing "contract" and "agreement" at the same time.

[0193] Filter and sort results:

[0194] Threshold screening: Set a similarity threshold of 0.6. Only files with a similarity greater than the threshold and files obtained through partial matching and synonym matching will be included in the preliminary screening results.

[0195] Comprehensive sorting: sort the results after preliminary screening by comprehensively considering the semantic similarity score, the degree of partial matching (such as the length of the matching substring, the number of occurrences, etc.) and other relevant factors (such as the update time of the archive, the access frequency, etc.). The sorted results are presented to the user to complete the fuzzy query operation.

[0196] Example 2

[0197] This embodiment provides a file information accurate inspection system based on cloud architecture, including:

[0198] The data acquisition module is configured as

[0199] A computer-readable storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded and executed by a processor of a terminal device, for a method for accurately checking archival information based on a cloud architecture.

[0200] A terminal device includes a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor to implement a cloud-based archival information accurate verification method.

[0201] The above are all preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for accurately checking archive information based on cloud architecture, characterized in that: include: Obtain archival information data; Preprocess the acquired archival information data; Based on the distributed cloud storage architecture, the pre-processed archive information data is distributed and stored on multiple cloud nodes; Semantic understanding of the query keywords entered by users based on semantic analysis and machine learning algorithms; Provide inspection results with different accuracy according to user needs.

2. According to the cloud architecture-based archival information accurate inspection method of claim 1, it is characterized in that: The acquisition of archival information data includes converting paper data into electronic format based on data scanning and OCR technology; obtaining electronic archival information in the business system according to predetermined rules and frequencies based on docking with various business system databases; and collecting information from archival information source websites based on web crawler methods.

3. According to the cloud architecture-based archival information accurate inspection method of claim 2, it is characterized in that: The preprocessing of the acquired archival information data includes unifying the file format and data encoding of the acquired archival information data, and extracting key metadata information based on natural language processing metadata.

4. According to the cloud architecture-based archival information accurate inspection method of claim 3, it is characterized in that: The method comprises distributing and storing the pre-processed archive information data on multiple cloud nodes based on a distributed cloud storage architecture, including dividing the pre-processed archive information data into data blocks according to logical rules, and calculating a hash value of each data block by a hash function; The cloud node for each data block is determined using a node selection algorithm based on load balancing; finally, the data block is transmitted to the selected cloud node through the network.

5. According to the cloud architecture-based archival information accurate inspection method of claim 4, it is characterized in that: The data block is transmitted to the selected cloud node through the network, including copying each data block to multiple different cloud nodes based on a three-copy redundancy strategy, including backup storage on two other cloud nodes in different geographical locations.

6. According to the cloud architecture-based archival information accurate inspection method of claim 5, it is characterized in that: The method performs semantic understanding on the verification keywords input by the user based on semantic analysis and machine learning algorithms, including segmenting the verification keywords input by the user to obtain independent word blocks, and converting the independent word blocks into feature vectors, using the trained machine learning model to perform semantic understanding on the feature vectors, and outputting the semantic representation and semantic label of the keywords.

7. According to claim 6, a method for accurately checking archive information based on cloud architecture is characterized in that: The method of providing the inspection results of different precisions according to the user's needs includes reasoning about semantic extension according to the output results of the machine learning model, combining the fuzzy query algorithm and the precise query algorithm, and providing the inspection results of different precisions according to the user's needs; wherein, the decision formula is constructed by calculating the cosine similarity of the keyword vector and the data element vector and setting the dynamic adjustment factor, and the judgment rule is used for the user's keyword matching query method, and the decision formula is expressed as: Among them, the confidence of the precise semantics is P precise , the confidence of fuzzy semantics is P vague , keyword vector and the archive metadata vector in the archive Cosine similarity of α is the dynamic adjustment factor.

8. A cloud-based archival information accurate inspection system, characterized in that: include: The data acquisition module is configured to acquire archive information data; A preprocessing module is configured to preprocess the acquired archive information data; The storage module is configured to distribute and store the pre-processed archive information data on multiple cloud nodes based on a distributed cloud storage architecture; A semantic understanding module is configured to perform semantic understanding of the query keywords input by the user based on semantic analysis and machine learning algorithms; The inspection module is configured to provide inspection results of different precisions according to user needs.

9. A computer-readable storage medium storing a plurality of instructions, characterized in that: The instructions are suitable for being loaded by a processor of a terminal device and executing the method according to claim 1 .

10. A terminal device, comprising a processor and a computer-readable storage medium, wherein the processor is used to implement each instruction; and the computer-readable storage medium is used to store multiple instructions, characterized in that: The instructions are suitable for being loaded by a processor and executing the method as claimed in claim 1 .