Intelligent archive information retrieval method based on semantic analysis

Through the semantic analysis method, the problem of insufficient semantic understanding and security in traditional archival information retrieval technology is solved, and more accurate and safe archival information retrieval is achieved.

CN119961435APending Publication Date: 2025-05-09NANJING INST OF RAILWAY TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510114929.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Traditional archival information relies on keyword matching, making it difficult to understand the semantics and contextual relationships of documents, resulting in poor accuracy of search results and difficult to ensure the security of the search information.

Method used

The intelligent search method of archival information based on semantic analysis is adopted. A variety of archival information data are collected by scanning and semantic analysis to obtain search keywords, a search model is constructed based on keywords, and different search verification methods are set according to archival information of different levels to ensure security.

Benefits of technology

Improve the accuracy and security of archival information retrieval, enables faster access to useful information, and ensures the security and compliance of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961435A_ABST
    Figure CN119961435A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent archive information retrieval method based on semantic analysis, and relates to the technical field of semantic analysis, according to the intelligent archive information retrieval method, through data collection, semantic analysis, keyword extraction and model classification and training, the whole process not only improves the retrieval efficiency, but also improves the accuracy of a retrieval result; different retrieval verification modes are set for different types of archive information, sensitive data are effectively protected, and it is ensured that only authorized users can access specific data; manual intervention is reduced through semantic analysis and intelligent model training, the retrieval performance is continuously optimized through machine learning, and the intelligent level of file retrieval is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic analysis, and in particular to an intelligent retrieval method for archive information based on semantic analysis. Background Art

[0002] With the continuous development of digital information technology, digital management of archival information is carried out through digital technology, which saves storage space and labor costs, reduces physical losses, and avoids the loss and damage of archives. When archival information needs to be viewed, it is retrieved through intelligent retrieval technology, while traditional retrieval technology uses keyword retrieval, which only matches the search terms with the text in the document, but has difficulty understanding the semantics and contextual relationship of the document, resulting in poor accuracy of the retrieval results. It is also interfered by a large amount of irrelevant content, making it difficult to screen out useful information, and it is difficult to ensure the security of information after retrieval.

[0003] To sum up, how to improve the accuracy of intelligent retrieval of archival information and ensure the security of the information after retrieval is an urgent problem to be solved and optimized in the intelligent retrieval method of archival information based on semantic analysis. Summary of the invention

[0004] The present invention provides an intelligent retrieval method for archival information based on semantic analysis, which solves the technical problem of how to improve the accuracy of intelligent retrieval of archival information and obtain useful information more quickly.

[0005] In order to solve the above technical problems, the present invention provides an intelligent retrieval method for archival information based on semantic analysis, and the specific technical solution is as follows: Scan and collect various archival information data to obtain archival information data sets; Performing semantic analysis on the archival information data set to obtain archival information retrieval keywords; based on the archival information retrieval keywords, sorting the first letters of each word of the archival information retrieval keywords in sequence to obtain a retrieval keyword sequence list; Based on the archival information data set, the archival information data set is classified and divided to obtain a plurality of archival information retrieval sections; different retrieval verification methods are set according to archival information retrieval sections of different levels to enable the archival retrieval sections to respond safely; The archival information retrieval section and the search keyword sequence list are constructed into a training set to obtain the archival retrieval model; the archival retrieval model outputs a retrieval recognition result representing an archival information data item, and at least one initial letter of a search keyword in the search keyword sequence list is input into the archival retrieval model to output a retrieval result representing archival information data.

[0006] As a further optimization solution of the present invention, a variety of archival information data are scanned and collected to obtain an archival information data set, including: Scan and collect a variety of archival information data to obtain a pre-collected archival information image data set; adjust the contrast, brightness and hue of the image in the pre-collected archival information image data set to obtain a clear scanned image data set; Denoising the image background noise in the clear scanned image data set to obtain a prominent image data set; the prominent image data set indicates that the text display of each image data item in the data set is more prominent and obvious; removing the blurred, ghosted or interfering noise parts of each image data item in the prominent image data set to obtain a clear image data set; The clear image data set is binarized to convert the clear image data into a black and white background color to obtain a binary image data set; the text in the binary image data set is processed by optical character recognition technology to obtain an archival information data set; the text data of the archival information data set is loaded to store the archival information text data to be intelligently retrieved by the archival retrieval model.

[0007] As a further optimization scheme of the present invention, based on the archival information data set, the archival information data set is classified and divided to obtain multiple archival information retrieval sections, including: Classifying and dividing the archival information data set to obtain a plurality of archival information data subsets; respectively grouping the plurality of archival information data subsets into a section, and labeling and annotating the section to obtain a plurality of archival information retrieval sections; The multiple archival information retrieval sections are searched and graded to obtain archival information retrieval sections of different levels; the archival information retrieval sections of different levels include top secret archival information retrieval sections, confidential archival information retrieval sections, internal archival information retrieval sections and public archival information retrieval sections; and specific retrieval sections of archival information data are obtained according to the first letters of the search keywords input in sequence.

[0008] As a further optimization scheme of the present invention, different search verification methods are set according to different levels of archive information search sections to make the archive search section respond safely, including: Based on the specific search section of the archival information data, when the archival information required by the search user is in the public archival information section, the public archival information section stores the archival basic information data and can be directly opened to the public without identity verification or any authority control; When in the internal archive information section, only internal employees can view it, and user identity verification is required; when in the confidential archive information section, further security permission verification is required before access; when in the top secret archive information section, multiple identity verification is required; Based on the access rights of each archival information section, each archival information section responds according to the verification information to obtain confidential archival information data at different levels.

[0009] As a further optimization scheme of the present invention, based on the access rights of each archive information section, each archive information section responds according to the verification information to obtain confidential archive information data of different levels, including: By retrieving the user's identity and authentication, only legitimate users can access the archive information; when the user is successfully authenticated, access rights are controlled based on the user's identity and authority and the level of the archive to determine whether there is authority to access the archive information data corresponding to the high-level archive information section; Different verification methods are set according to different archive levels; and the access permission verification strength is gradually increased to obtain corresponding verification results; according to the verification results, the searching user obtains the corresponding archive information data, so that the archive retrieval model safely responds to the archive information retrieval behavior.

[0010] As a further optimization solution of the present invention, the archive retrieval model includes: The archive information retrieval section and the retrieval keyword sequence list are used to construct a data set to generate structure data, and the structure data is encoded into sequence data to train the archive retrieval model; Inputting the sequence data into the archive retrieval model; the archive retrieval model comprises an input layer, a first hidden layer, a second hidden layer, a third hidden layer and an output layer, transmitting the intermediate representation data of multiple hidden layers to the output layer, and the output layer outputting the retrieval recognition result representing the archive information data item; At least one initial letter of a search keyword in the search keyword sequence list is input into the archive retrieval model, and the output layer outputs the archive information data retrieval result.

[0011] As a further optimization scheme of the present invention, according to the user's archive retrieval and browsing situation, the corresponding archive information after the retrieval is extracted to form the user's query mechanism; through the user's periodic archive retrieval behavior, the high frequency of retrieval users is summarized; by analyzing the high frequency retrieval attribute characteristics of the retrieval users, the archive retrieval model and the query mechanism are interacted.

[0012] As a further optimization solution of the present invention, the query mechanism includes The first search behavior data is obtained by analyzing the keywords, query frequency, search time and search order of the search users for the archive information; the second search behavior data is obtained by analyzing the user's stay time and click count when browsing the archive; By using the first search behavior data and the second search behavior data, a high-frequency search analysis is performed on the search user behavior to extract the user's high-frequency search features; Based on the extracted high-frequency search features, a high-frequency search feature weight is set, which represents the high-frequency query degree of the search user for the archives within the time period, that is, ; In the formula, AF represents the query frequency measurement parameter, T represents the period, ci represents the file information type browsed within the T period, x represents the number of files browsed within the T period, W i Represents the weight of high-frequency retrieval features.

[0013] As a further optimization scheme of the present invention, all high-frequency searches of the archive type are forgotten based on the high-frequency query level to obtain the forgetting factor of the search; Obtain the forgetting factor weight, and add the forgetting factor weight to the archive information with high frequency retrieval by the user; where F(x) represents the forgetting factor weight; t represents the time node from the current time node to the high frequency retrieval feature assignment time node, and f represents the half-life, that is, the high frequency retrieval feature forgetting of the model needs to last at least f days; When the high-frequency search archive information in the archive information already belongs to the searching user, the high-frequency search feature weight of the archive type will be re-acquired to update the high-frequency search feature weight storage; when the high-frequency search archive information does not belong to the searching user, the archive information will be assigned a high-frequency search feature weight, and the search keyword sequence list will be re-sorted according to the size of the high-frequency search feature weight.

[0014] As a further optimization scheme of the present invention, based on the retrieval results output by the archive retrieval model, the retrieved archive information data is compared with the historical archive information data to obtain a comparison result; deviation data is obtained according to the comparison result to form a deviation data set; the deviation data set is iteratively analyzed to obtain interference factors of the deviation data; According to the interference factors of the deviation data, a secondary semantic analysis is performed on each archival information data item in the archival information data set to obtain accurate archival information data; the accurate archival information data is input into the archival retrieval model again for training until the archival retrieval model outputs accurate archival information data with a rapid response.

[0015] The present invention has at least the following beneficial effects: the present invention obtains an archival information data set by scanning and collecting a variety of archival information; by scanning and collecting a variety of archival information, the diversity of archival content can be fully acquired, including multi-dimensional information such as text, pictures, and tables. This provides a rich source of original data for subsequent analysis, classification, and retrieval; it ensures the integrity of the data, avoids information loss, and can more accurately reflect the diversity of archival data; the diversity of the data set helps to build a more comprehensive and accurate retrieval model, and improves the subsequent semantic analysis and retrieval effect.

[0016] The archival information data set is subjected to semantic analysis to obtain archival information retrieval keywords; through semantic analysis, representative and critical keywords can be extracted from a large amount of text data. This process is not just a simple text matching, but to identify potential retrieval targets by understanding the actual meaning of the text; keyword extraction helps to improve retrieval accuracy, ensuring that users can find relevant archives through precise keywords rather than simply relying on word frequency; improve the intelligence of retrieval and reduce the labor cost of manual intervention and keyword selection.

[0017] Sorting keywords alphabetically helps standardize the search process. The sorted keyword sequence is more standardized, making the search process consistent and repeatable; the sorted keyword sequence can improve efficiency, especially when dealing with large amounts of data, and can increase the search response speed; it helps reduce ambiguity, ensure a clear search order, and improve the relevance of search results.

[0018] Based on the classification and division of archival information data sets, multiple archival information retrieval sections are obtained. Classifying data sets can classify different types of archival information and improve retrieval efficiency. For example, by distinguishing archival types such as files, reports, pictures, etc., different retrieval strategies can be used for each type; classified data is easy to process and manage, especially in complex archives, which can help users quickly locate the required retrieval section; through classification processing, more detailed analysis and processing can be carried out, thereby improving the accuracy and pertinence of retrieval.

[0019] Setting different retrieval verification methods for different levels of archival information retrieval sections can enhance the security of retrieval. For example, sensitive information may require higher authentication or permission control; prevent unauthorized access and protect user privacy and sensitive data. This is particularly important for processing archival data involving personal privacy, confidential information or commercial secrets; enhance the compliance of archival management and ensure compliance with relevant laws and regulations, especially in terms of archival protection and data security.

[0020] The archival information retrieval section and the retrieval keyword sequence list are constructed into a training set to obtain the archival retrieval model. Constructing the training set is a key step in building an intelligent archival retrieval model. By combining the retrieval section and keyword sequence with the archival information data, the model can learn how to effectively match the query with the data set. Through training, the retrieval model can be continuously optimized in subsequent practical applications to improve the accuracy and response speed of the retrieval results. The data feedback mechanism in the model training process enables it to have adaptive capabilities, and can adjust algorithms and parameters according to actual usage to improve long-term retrieval performance.

[0021] The archive retrieval model outputs the retrieval recognition results, and searches by inputting a keyword sequence. Through the retrieval model, the retrieval keywords entered by the user can be efficiently identified, and relevant data can be accurately extracted from the archive information to provide instant query results; the retrieval model searches according to the keyword sequence, which improves the matching degree and accuracy of the information. Users can obtain more relevant and more demand-oriented retrieval results; this step can also be continuously improved through intelligent algorithms to improve the model's ability to understand user queries, reduce the amount of query information that users need to enter, and improve user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a flow chart of an intelligent archival information retrieval method based on semantic analysis provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The present application is further described in detail below in conjunction with the accompanying drawings. It is necessary to point out here that the following specific implementation methods are only used to further illustrate the present application and cannot be understood as limiting the scope of protection of the present application. Technical personnel in this field can make some non-essential improvements and adjustments to the present application based on the above application content.

[0024] The specific implementation method of the intelligent archival information retrieval method based on semantic analysis provided in this embodiment is as follows: like Figure 1 As shown, the intelligent retrieval method of archival information based on semantic analysis includes the following steps: Step 11, scanning and collecting a variety of archival information data to obtain an archival information data set; Step 12, performing semantic analysis on the archival information data set to obtain archival information retrieval keywords; based on the archival information retrieval keywords, sorting the first letters of each word of the archival information retrieval keywords in sequence to obtain a retrieval keyword sequence list; Step 13, based on the archival information data set, classify and divide the archival information data set to obtain a plurality of archival information retrieval sections; set different retrieval verification methods according to archival information retrieval sections of different levels to enable the archival retrieval sections to respond safely; Step 14, constructing the archival information retrieval section and the search keyword sequence list into a training set to obtain the archival retrieval model; the archival retrieval model outputs a retrieval recognition result representing the archival information data item, and inputs the first letter of at least one search keyword in the search keyword sequence list into the archival retrieval model to output a retrieval result representing the archival information data item.

[0025] In this embodiment, in the preparation stage in step 11, it is necessary to determine the type, content and format of the collected archive information to ensure that the category, source and format of the archive are clear, such as paper archives, electronic archives, image archives, etc.

[0026] Choose the appropriate scanning equipment according to the type and needs of the archives, such as document scanners, flatbed scanners or dedicated archive scanning equipment. If you are scanning a large number of archives, you may need more efficient batch scanning equipment; Clean up the archives: Make sure the archives are not damaged, folded or have other factors that affect the scanning quality. If necessary, make simple repairs; The archive information types include personnel information archives, confidential document archives or other various information files that need to be archived; Set scanning parameters according to the type and purpose of the file. For example: Resolution: usually use 300DPI (dots per inch) for clear scanning; select black and white, grayscale or color scanning according to the file type; common file formats include PDF, TIFF, JPEG, etc. For text information, it is more common to choose PDF or TIFF format; for image content, you may choose JPEG format.

[0027] Start scanning: Place the document into the scanning device and start scanning. It should be noted that if it is a multi-page document, using an automatic document feeder (ADF) can improve efficiency; if there is no ADF, manually scanning each page is also acceptable.

[0028] After scanning, check the data to ensure the clarity, completeness and accuracy of the scan. For example, check whether there are problems such as missing pages, blurred scans or color differences; Image repair: If problems such as spots, blur or bending occur during the scanning process, you can use image repair software (such as Adobe Photoshop or dedicated document repair tools) to correct them; For paper archives, the brightness and contrast of the scanned image may need to be adjusted to ensure the readability of the document; For scanned image files, you can use OCR software (such as ABBYY FineReader, Tesseract, etc.) for character recognition to convert the text in the image into editable and searchable text data. This can greatly improve the efficiency of subsequent data processing; Although OCR technology is relatively accurate, it may also be misidentified, especially for complex fonts, low-quality scanned images, etc. Manual proofreading is required to ensure the accuracy of text data.

[0029] Organize the scanned archives according to the preset classification standards to ensure that all archives have clear labels, classifications and associated data. For example, use metadata such as "archive number", "title", "creation time" to organize archives; input the organized archive information into the archive management or database to establish an archive information data set. The database can be an SQL database, a NoSQL database, or a dedicated archive management; each archive can be retrieved and managed by entering relevant metadata (such as file name, author, date, etc.). Metadata helps with subsequent archive search and archiving.

[0030] Archive data sets can be stored in local hard disks, network storage (NAS), cloud storage, or enterprise archive management. When choosing a storage solution, you need to consider data security, access efficiency, and long-term preservation requirements; in order to prevent data loss, be sure to set up a regular backup mechanism to ensure data security. You can perform local backup, remote backup, and cloud backup.

[0031] Establish search for archive data sets, so that users can efficiently find archive information. Search can be based on archive content (such as text content recognized by OCR) or metadata (such as archive number, creation time, subject, etc.); set access rights to archives according to actual needs to ensure the security of sensitive or confidential information and prevent unauthorized access.

[0032] Archive the digitized data to long-term storage to ensure data stability and security; select appropriate long-term storage formats and storage media to ensure that the archives can still be read in the next few decades. For example, regularly migrate archive files to new storage to avoid the impact of outdated technology.

[0033] Regularly perform integrity checks on stored data to ensure that no data corruption or loss has occurred.

[0034] Update metadata: As new records are added or existing records are modified, metadata should be regularly updated and maintained to ensure that the records are always up to date.

[0035] Based on step 11, archival information is obtained from various sources (such as paper archives, electronic archives, databases, etc.) through various technical means (such as document scanning, data capture, etc.), and gathered into an archival data set containing diverse data formats and contents; by widely collecting different types of archival information data, it can ensure that the data sources of the archives are diverse and comprehensive, and can cover more fields and needs, providing a rich data foundation for subsequent retrieval and analysis; through step 12, the collected archival information data is subjected to semantic analysis. The goal of semantic analysis is to extract core search keywords from the text, usually including important nouns, phrases, terms, etc., which will become the main keywords for users to search. Through semantic analysis, key information in the archives can be accurately identified, avoiding relying solely on literal matching of keywords, making the search results more intelligent and more relevant. In addition, semantic analysis can also understand synonyms and contextual relationships, thereby improving the flexibility and accuracy of retrieval.

[0036] Step 13 classifies the above-obtained archive information data set by content or other dimensions, such as by subject, time, type, etc., to form multiple search sections. Different search sections can set different search verification methods according to the importance and sensitivity of the archives to ensure the search security of each section; classification can improve the efficiency and accuracy of archive information retrieval, and users can search for specific classification targets to avoid interference from a large amount of irrelevant information. In addition, search sections with different security levels can ensure the security of sensitive or important information and prevent unauthorized access.

[0037] Step 14 constructs the archival information retrieval keyword sequence obtained from step 12 and step 13 together with the classification section data into a training set for machine learning. These data will be used to train the archival retrieval model so that the model can learn how to provide the most relevant archival information data items based on the retrieval keywords entered by the user; by training the machine learning model, the retrieval accuracy and response speed can be gradually improved. The model can learn the correlation between different keywords and archival information data items, and optimize the retrieval results, making the final retrieval more intelligent and efficient. In particular, sorting by the first letter of the keyword can make the organization of the retrieval results more standardized and intuitive.

[0038] The entire process is coordinated through four steps, from information collection, semantic analysis, classification processing to model training, which gradually improves the intelligence and security of archival information retrieval; thereby using semantic analysis to make the retrieval not only rely on literal keywords, but also better understand the actual meaning of the archival content; classification processing and security stratification can optimize the organization of information, so that users can find relevant information more quickly when searching; through different retrieval verification methods, the security of sensitive information is ensured to prevent improper access; through the training of machine learning models, the retrieval results are made more and more accurate, and personalized retrieval results can be provided based on changes in keywords.

[0039] In a preferred embodiment of the present invention, step 11 further includes the following steps: Step 111, scanning and collecting a variety of archival information data to obtain a pre-collected archival information image data set; adjusting the contrast, brightness and hue of the image in the pre-collected archival information image data set to obtain a clear scanned image data set; Step 112, denoising the image background noise in the clear scanned image data set to obtain a prominent image data set; the prominent image data set indicates that the text display of each image data item in the data set is more prominent; removing the blurred, ghosted or interfering noise parts of each image data item in the prominent image data set to obtain a clear image data set; Step 113, binarizes the clear image data set to convert the clear image data into a black and white background color to obtain a binary image data set; uses optical character recognition technology to recognize the text in the binary image data set to obtain an archival information data set; loads the text data of the archival information data set to store the archival information text data for intelligent retrieval by the archival retrieval model.

[0040] In an embodiment of the present invention, in step 111, image data is obtained from paper archives or other sources through a scanning device or other image acquisition technology. After the image data is initially collected, the image quality may not be ideal due to factors such as the scanning device, light, angle, etc. In order to improve the quality of the image, the contrast, brightness and hue of the image are adjusted. These adjustments can enhance the details and clarity of the image, making the text more prominent; by enhancing the contrast of the image, the difference between the text and the background can be made more obvious, thereby improving readability; adjusting the brightness can make the light and dark parts of the image more balanced, ensuring that the text is not overexposed or dim; by adjusting the hue, the image color can be made more balanced to avoid visual unclearness caused by color difference; these image adjustment operations can effectively improve the image quality, ensure that the image is clearer and the text is easier to recognize during subsequent processing, and lay a good foundation for image processing and text extraction in subsequent steps.

[0041] The focus of step 112 is to remove unnecessary background noise in the image, improve the quality of the image, and make the text information more prominent and clear; irrelevant pixels or defects in the image, dust during scanning, paper texture, etc. By using noise removal algorithms (such as median filtering, mean filtering, etc.), these background noises are eliminated, making the useful text in the image clearer; blur, ghosting and interference noise elimination: The image may be blurred, ghosting or other interference noise due to poor scanning quality, paper folding, etc. These interferences will affect the recognition of text. Through image clarity processing (such as deblurring algorithms, sharpening processing, etc.), these influencing factors are eliminated to obtain a clearer and noise-free image; through denoising and eliminating interference, the image quality is significantly improved, and the text is more prominent. This step greatly reduces the text recognition errors caused by image problems and provides high-quality input data for subsequent optical character recognition (OCR) technology.

[0042] Step 113 converts the image into a binary image and extracts the text data therein using OCR technology; binarization is a key step in image processing, which converts all pixel values ​​in the image into black (representing text) and white (representing background). This processing simplifies the image into only two colors, thereby highlighting the text part. Common binarization methods include Otsu algorithm, etc., which can clearly convert the image into a black and white image by setting a suitable threshold; OCR technology recognizes the text in the image through an algorithm, extracts the black characters in the binary image and converts them into digitized text information. These text data can include all text content in the document, such as titles, paragraphs, tables, etc.; through binarization processing, the text in the image is clearly separated and the background is simplified, providing an ideal input form. Using OCR technology to further convert the text in the image into machine-processable text data makes subsequent archival information retrieval more accurate and fast. In particular, during the storage and retrieval process, the digitized text data can be easily processed and queried by the archival retrieval model.

[0043] In a preferred embodiment of the present invention, the above step 13 classifies and divides the archival information data set based on the archival information data set to obtain multiple archival information retrieval sections, including the following steps: Step 131, classifying and dividing the archival information data set to obtain a plurality of archival information data subsets; respectively grouping the plurality of archival information data subsets into a section, and labeling and annotating the section to obtain a plurality of archival information retrieval sections; Step 132, performing retrieval classification processing on the multiple archival information retrieval sections to obtain archival information retrieval sections of different levels; the archival information retrieval sections of different levels include top secret archival information retrieval sections, confidential archival information retrieval sections, internal archival information retrieval sections and public archival information retrieval sections; according to the first letters of the search keywords input in sequence, a specific retrieval section of the archival information data is obtained.

[0044] In the embodiment of the present invention, the original archival information data set is classified according to certain standards or rules through step 131, with the purpose of grouping highly relevant data together to facilitate management and subsequent processing. The classification standard may be based on content, type, subject, time, etc.; through the division, multiple data subsets are finally formed, each subset represents a specific archival category; through the classification and division of archival information, the information becomes more organized, which helps to avoid confusion and improve the manageability of data; the classified archival information subsets can be quickly located in subsequent retrieval, saving retrieval time and avoiding the retrieval of irrelevant data.

[0045] After the classified multiple subsets of archival information are summarized in step 132, each subset is integrated into a "section". Sections are organizational units for easy access and retrieval; each section will have clear label notes, which help to better describe and distinguish the content of the section, for example, the label can be "confidential", "public", "important", etc. Through labels, users can quickly understand the content or sensitivity level of the section; labels and notes make each archival information retrieval section more intuitive and clear, improving the visibility and ease of use of information; adding labels to each section allows users to quickly locate the archival information they need, which is particularly important when facing a large amount of information.

[0046] The archival information retrieval section is further classified to categorize the archives according to their sensitivity. The classification is usually based on the confidentiality level of the archives, such as top secret, confidential, internal, public, etc. Each level represents the security and access rights of the archival information; it contains archival information that is critical to confidentiality and can only be accessed by authorized personnel.

[0047] Confidential archive information retrieval section: contains archives that are sensitive but not involving major interests, and has slightly wider access rights than the top secret section; internal archive information retrieval section: these archives are usually internal management information and are rarely accessed by external personnel; public archive information retrieval section: public information that can be accessed by anyone and usually does not involve sensitive content.

[0048] Through classification, we ensure that each type of archival information is properly protected to avoid unauthorized access or leakage; classification can help administrators more accurately control the access rights of different users or roles to prevent irrelevant personnel from accessing sensitive or confidential data; based on the classification, users can access archival information of a specific level as needed, which also improves query efficiency and avoids unnecessary information interference.

[0049] When users search, they will locate specific archive information search sections according to the first letter of the search keyword they entered. In this way, the search scope can be quickly narrowed down and matching sections can be found; the first letter search can quickly narrow the search scope and prevent users from searching the entire data set in a disorderly manner. For large amounts of data, this method significantly improves efficiency; through simplified search rules (enter only the first letter of the keyword), the user's operation becomes more convenient and the complexity of the operation is reduced.

[0050] In a preferred embodiment of the present invention, in the above step 13, different search verification methods are set according to different levels of archive information search sections to make the archive search section respond safely, and the following steps are also included: Step 133, based on the specific search section of the archival information data, when the archival information required by the search user is in the public archival information section, the public archival information section stores archival basic information data and can be directly opened to the public without identity verification or any authority control; Step 134, when in the internal archive information section, only internal employees can view it, and user identity verification is required; when in the confidential archive information section, further security permission verification is required before access; when in the top secret archive information section, multiple identity verification is required; Step 135 , based on the access rights of each archive information section, each archive information section responds according to the verification information to obtain confidential archive information data of different levels.

[0051] In the embodiment of the present invention, the public archive information section in step 133 contains archive information that does not involve sensitive content, and any user can access it; the storage content of the public archive information section is usually basic information data, which is an open resource, so no identity authentication or special permission control is required. Users only need to perform regular search operations to access relevant information; for archive information that does not involve confidentiality, open access allows users to quickly obtain the required basic information, greatly improving the convenience of use; operations without permission verification will reduce the burden, simplify the user's search process, and are suitable for publicly released data and information; it helps to improve the transparency of archive management, especially in the government or public field, and open public archives can allow the public to obtain basic information related to themselves.

[0052] The internal archive information section in step 134 generally includes archive information related to management information within a company or organization, and is only accessible to employees of the company or organization. Users are required to perform identity authentication, such as entering a user name and password, providing an employee number, or using a corporate certificate, to prove that they are authorized users. The confidential archival information section contains archival information that is sensitive, but does not affect major core interests. Only users with higher permissions can access it; users need to pass some more stringent permission verification, such as entering an additional password, performing fingerprint recognition, or verifying through a mobile phone after passing identity verification, to confirm that they have the authority to access this information.

[0053] The top secret archive information section contains archive information involving the most sensitive content, usually related to the company's core technology or strategic confidential information. Access to such archive information requires multiple authentications. Users must pass multiple authentications, such as password verification, SMS verification code, fingerprint or facial recognition, to ensure that only authorized personnel can access such information; through strict identity verification and permission control, unauthorized access can be effectively avoided, sensitive and confidential information can be protected, and information leakage can be prevented; for confidential and top secret archives, multiple authentication and permission control can help reduce the risk of information leakage and abuse faced by the organization, ensuring that only authorized personnel can view relevant materials; hierarchical permission control ensures that different users obtain corresponding data according to their roles and permissions. Different levels of archive information can be appropriately accessed and configured according to their sensitivity and importance to optimize the management and use of information.

[0054] Step 135 After the user has authenticated his identity, different responses will be made according to the user's authority level and verification results. For example, users who have passed the verification can access the corresponding archive information section. If the authority is insufficient, they cannot access certain sensitive data; when the user is verified and the authority meets the requirements, they can access specific archive sections; when the verification fails or the authority is insufficient, the user will not be able to access the corresponding archive information, and will be given a corresponding error prompt or denied access; accurate response based on the user's verification information ensures that each user can only access the archive section that he has the right to access, further strengthening the precision of information security and management; through a strict response mechanism, the risk of authority abuse and information leakage is avoided, especially in the management of sensitive data, and the security of the archive is maximized; when verifying and responding, reasonable arrangements can be made according to the authority, so that suitable users can smoothly access data, while irrelevant personnel are excluded. This mechanism improves the overall user experience and the standardization of information management.

[0055] In a preferred embodiment of the present invention, the above step 135 further includes the following steps: Step 1351, by retrieving the identity and authentication of the user, only the legitimate user can access the archive information; when the user authentication is successful, based on the user's identity and authority, the access authority is controlled according to the level of the archive to determine whether the user has the authority to advance to the archive information data corresponding to the high-level archive information section; Step 1352 sets different verification methods according to different archive levels; and gradually increases the access permission verification strength to obtain corresponding verification results; based on the verification results, the searching user obtains the corresponding archive information data, so that the archive retrieval model safely responds to the archive information retrieval behavior.

[0056] In the embodiment of the present invention, step 1351 performs identity recognition in some way, such as user name and password, fingerprint recognition, facial recognition, etc., to ensure that the user is legitimate; after verifying the identity, the user's authority will be determined based on the user's identity information, and access rights will be controlled based on the level of the archive. The core of this stage is to determine the legitimacy of the user and provide a basis for subsequent access control; according to the different security levels of the archive (for example, ordinary archives, confidential archives, top confidential archives, etc.), determine whether the user has the authority to access the corresponding archive information data; through an effective identity authentication mechanism, ensure that only legally authenticated users can access the archive data, reducing the risk of illegal access; control access based on the level of archive information to avoid unnecessary information leakage or erroneous access.

[0057] In step 1352, different access verification methods are set for each level according to the different levels of the archives. For example, for ordinary archives, only basic authentication (such as user name and password) may be required; for high-level archives, more stringent authentication methods (such as multi-factor authentication, face recognition, etc.) may be required; during the verification process, the verification strength will be gradually increased to ensure that users who access high-level archive information have higher security guarantees; setting different verification methods for different archive levels makes it more flexible and can balance security and convenience; gradually increasing the verification strength not only improves the security protection of the archives, but also reduces the complexity of user operations and optimizes the user experience.

[0058] In step 1353, based on the user's verification result, it is decided whether to allow the user to access data at a certain archive level. If the verification is successful, the archive information accessible to the user is retrieved, and the corresponding archive content is returned according to the authority; the retrieval process includes querying and extracting the archive database, and may also involve recording the user's access behavior for subsequent auditing and monitoring; ensuring that the user can obtain the archive data they have permission to access after passing strict authentication, preventing unauthorized access; through reasonable authority control and information retrieval, the security of the archive is guaranteed, and the efficiency is improved, ensuring that the data is accurately and timely used by the appropriate users.

[0059] In a preferred embodiment of the present invention, the above-mentioned archive retrieval model includes the following steps: The archive information retrieval section and the retrieval keyword sequence list are used to construct a data set to generate structure data, and the structure data is encoded into sequence data to train the archive retrieval model; Inputting the sequence data into the archive retrieval model; the archive retrieval model comprises an input layer, a first hidden layer, a second hidden layer, a third hidden layer and an output layer, transmitting the intermediate representation data of multiple hidden layers to the output layer, and the output layer outputting the retrieval recognition result representing the archive information data item; At least one initial letter of a search keyword in the search keyword sequence list is input into the archive retrieval model, and the output layer outputs the archive information data retrieval result.

[0060] In an embodiment of the present invention, it is first necessary to extract information about the archives from the archives, and match this information (which may be titles, contents, tags, metadata, etc.) with search keywords. The search keyword sequence is a keyword list generated by the user or by inputting search conditions. This information is organized into a structured data set, and each data item consists of archive information and its corresponding search keyword sequence. These data usually include relevant attributes of the archive (such as document content, classification, keywords, etc.) and a sequence of search terms (keywords) to facilitate subsequent encoding and training; by constructing a structured data set, the originally complex or scattered archive information can be converted into data in a unified format, which is convenient for model understanding and processing; structured data sets are the basis for training deep learning models, which help to improve the learning efficiency and accuracy of the model.

[0061] Structured data usually needs to be encoded, especially when processing text data. Encoding usually involves converting text into digital representations that can be processed by computers. Common practices include word vectors (such as Word2Vec, GloVe), One-Hot encoding, etc. Through this encoding method, archival information and keywords can be converted into vector form to meet the requirements of neural network input; by inputting these sequence data into the neural network (such as multi-layer perceptron, LSTM, etc.), the model learns. The goal of training is to enable the model to predict relevant archival information based on the search keyword sequence and optimize parameters to improve the search effect; encoding data into a sequence enables the neural network to recognize the semantic relationship between texts, thereby improving the search effect; using a deep learning model to train the data can learn useful features from a large amount of historical data and improve the search accuracy.

[0062] The encoded sequence data (such as the search keywords or archive information features entered by the user) will be input into the archive retrieval model for processing. The model is usually a deep neural network (DNN) including multiple levels of hidden layers.

[0063] Model architecture, including: Input layer: receives input data (i.e., vector representation of the search keyword sequence).

[0064] Multiple hidden layers: These layers play a role in abstracting data features layer by layer in the model, helping the model to extract higher-level features from complex data.

[0065] Output layer: Outputs the search results, that is, the relevant archive data predicted by the model based on the input search keywords.

[0066] Feature extraction: Through multiple hidden layers, the model can extract high-order features of the data layer by layer, further improving the retrieval performance, so that the model can adapt to different input formats and retrieval tasks and has a high generalization ability.

[0067] In the multiple hidden layers of a deep neural network, each layer generates an intermediate representation, which is usually passed to the next layer until the final output layer. The output layer generates the final retrieval results based on these intermediate representations; at the output layer, the network will output a predicted value representing the retrieval results of the archive data based on the information passed from the input layer to the hidden layer. This result can be a relevance score or a specific archive item, indicating which archives best meet the user's retrieval intent.

[0068] Through the gradual information transmission of the middle layer, the model can extract the key features of the retrieval information and output more accurate results; through the hierarchical learning of multiple hidden layers, the network can understand complex input data and retrieval requirements, and finally output more accurate retrieval recognition results; the key lies in using the first letter as an additional input feature. By inputting the first letter of the search keyword into the model, it can help the model further understand the intention of the search, especially when polysemous words or keywords are not completely matched, the first letter may play a certain role in distinguishing; the first letter can effectively improve the accuracy of the search, especially when the keyword sequence is ambiguous or vague; using additional information such as the first letter, the response speed and accuracy of the model can be optimized, and the user's search experience can be improved.

[0069] In a preferred embodiment of the present invention, specifically, according to the user's archive retrieval and browsing situation, the corresponding archive information after the retrieval is extracted to form the user's query mechanism; through the user's periodic archive retrieval behavior, the high frequency of retrieval users is summarized; by analyzing the high frequency retrieval attribute characteristics of the retrieval users, the archive retrieval model and the query mechanism are interacted.

[0070] In the embodiment of the present invention, the user's archive retrieval behavior is first recorded and analyzed, including the query conditions entered by the user each time, the frequency of clicking on the search results, the archive content browsed, and the dwell time; the archive information features that the user cares about are extracted from these behavior data, such as archive type, keywords, search time, etc. These data can help understand the user's needs and preferences; based on these search information, a personalized query mechanism for the user will gradually be formed. That is, the user's frequently used search keywords or attributes will be identified, and the most relevant archives will be automatically recommended.

[0071] Improve the accuracy of archive retrieval and reduce the number of query conditions that users need to repeatedly enter; be able to respond to user needs more quickly and improve user experience; automate personalized recommendations and reduce user search time; summarize the high frequency of searches by users through their periodic archive search behavior, monitor and record users' search history, especially those frequently searched archive types or specific keywords; through periodic search behavior analysis, be able to identify high-frequency search items and summarize the archive categories or information content that users frequently query; use statistical methods or machine learning technology to perform pattern recognition on retrieval data and discover users' search patterns; by optimizing these high-frequency search items, improve retrieval efficiency and accuracy, provide users with more accurate recommendations and search results, and reduce the time users spend repeatedly searching for the same or similar information; better adapt to changes in user needs and improve flexibility.

[0072] By analyzing the high-frequency search attribute characteristics of the search users, the archive retrieval model is enabled to interact with the query mechanism. By analyzing the high-frequency search attribute characteristics, the search conditions most commonly used by users are identified, such as: keywords, archive types, date ranges, etc.

[0073] Through these features, the retrieval model can be further optimized. For example, if a certain attribute (such as time or subject) of a certain type of archive is frequently retrieved, the user's query intention can be automatically inferred based on these features, thereby providing more personalized and accurate search results; in addition, the archive retrieval model can be dynamically adjusted to meet the needs of different users, ensuring that the retrieval mechanism always matches the user's behavior; the model is optimized based on the user's high-frequency retrieval behavior to improve the relevance and accuracy of the retrieval.

[0074] It supports seamless connection of user query needs, so that each user's query can get a more accurate and faster response; the enhanced intelligence and adaptive capabilities enable it to still effectively meet the changing needs of users after long-term use.

[0075] In a preferred embodiment of the present invention, the query mechanism described above includes the following steps: The first search behavior data is obtained by analyzing the keywords, query frequency, search time and search order of the search users for archive information; the second search behavior data is obtained by analyzing the user's stay time and click count when browsing archives; By using the first search behavior data and the second search behavior data, a high-frequency search analysis is performed on the search user behavior to extract the user's high-frequency search features; Based on the extracted high-frequency search features, a high-frequency search feature weight is set, which represents the high-frequency query degree of the search user for the archives within the time period, that is, ; In the formula, AF represents the query frequency measurement parameter, T represents the period, ci represents the type of archive information browsed within the T period, x represents the number of archives browsed within the T period, and Wi represents the high-frequency retrieval feature weight.

[0076] In the embodiment of the present invention, the first search behavior data is obtained by analyzing the keywords, query frequency, search time and search order of the search user's search for archive information. By recording the keywords entered by the user each time for search, it reflects the subject or archive category that the user is concerned about; the query frequency of the user for certain keywords or archive information is analyzed. Frequently searched keywords usually mean that the user has a high demand for the information; the user's search time is recorded to understand when the user searches. For example, the search in a specific period of time may be related to the user's work rhythm or demand changes.

[0077] Observe the order in which users perform searches to understand their search paths. For example, some users may search archives in chronological order, while others may search by category or subject.

[0078] By integrating these data, we can construct the first search behavior data to provide a basis for subsequent analysis; thereby capturing the basic patterns of user search behavior, identifying users' frequently used keywords and query habits, building a personalized search model, and predicting in advance the archival content that users may be interested in. We can further provide search suggestions that better meet user needs based on these characteristics, thereby improving user query efficiency.

[0079] By analyzing the time users spend viewing a certain profile. If a profile is viewed for a long time by a user, it may mean that the profile is important to the user or has a high relevance; record the number of times a user clicks on the profile. Multiple clicks may mean that the user has a strong interest in the information in the profile, or has checked certain details multiple times during the viewing process.

[0080] Based on these browsing behaviors, the second retrieval behavior data can be obtained, which complements the query behavior data in the first part and provides an in-depth understanding of the user's attention and interest in the archive content; helps evaluate the attractiveness of the archive and the user's interests, thereby optimizing the display order of the archives; based on the dwell time and number of clicks, the recommended content can be dynamically adjusted to more accurately match user needs; by analyzing the user's browsing behavior, it is possible to identify which archive information is more in line with the user's deep-seated needs.

[0081] Through the first search behavior data and the second search behavior data, a high-frequency search analysis is performed on the search user behavior to extract the user's high-frequency search features; the first search behavior data (such as keywords, query frequency, search order) is combined with the second search behavior data (such as dwell time, number of clicks) to comprehensively analyze all user behaviors; the goal of high-frequency search analysis is to identify the user's commonly used search features. For example, a user frequently searches for a keyword and stays for a long time when clicking on related archives, indicating that the user has a high interest in this type of archive; the extraction of high-frequency features is achieved through cluster analysis, frequency statistics, association rule mining and other technologies, further revealing the user's interest preferences and needs.

[0082] Through comprehensive analysis, we can fully understand the user's behavioral characteristics and thus tap into the user's potential needs; provide more accurate personalized recommendation services and recommend the most relevant files based on the user's high-frequency search characteristics; reduce the interference of irrelevant information and improve retrieval efficiency.

[0083] Based on the extracted high-frequency search features, a weight is assigned to each feature. High-frequency search behaviors (such as repeated searches for a keyword) will be assigned higher weights, while less frequent behaviors will be assigned lower weights. At the same time, user behaviors in different time periods will be tracked, and weights will be dynamically adjusted according to changes in query frequency. For example, if a keyword is frequently searched in a certain period of time, its weight will be increased to reflect the changes in user interests in that period of time; through this weight setting, it can more accurately reflect the user's interest in different archives and their changes in needs in different time periods, so that the search priority can be adjusted according to the user's real-time behavior to ensure that the most relevant archive information is ranked first; provide dynamic personalized services based on behavioral data to further enhance user experience; make the sorting of search results more in line with the user's current interests and needs, and avoid interference from outdated or irrelevant information.

[0084] In a preferred embodiment of the present invention, the query mechanism described above further includes: Based on the high-frequency query level, all high-frequency searches of the archive type are forgotten to obtain the forgetting factor of the search; Obtain the forgetting factor weight, and add the forgetting factor weight to the archive information with high frequency retrieval by the user; where F(x) represents the forgetting factor weight; t represents the time node from the current time node to the high frequency retrieval feature assignment time node, and f represents the half-life, that is, the high frequency retrieval feature forgetting of the model needs to last at least f days; When the high-frequency search archive information in the archive information already belongs to the searching user, the high-frequency search feature weight of the archive type will be re-acquired to update the high-frequency search feature weight storage; when the high-frequency search archive information does not belong to the searching user, the archive information will be assigned a high-frequency search feature weight, and the search keyword sequence list will be re-sorted according to the size of the high-frequency search feature weight.

[0085] In the embodiment of the present invention, the high-frequency search information is "forgotten" according to the high-frequency query degree of the archive type. The "forgetting" here means lowering the search priority of some outdated or no longer frequently queried archive information; by obtaining the "forgetting factor" weight, the "forgetting degree" of each archive information is quantified. This factor increases over time, indicating the gradual fading of past high-frequency queries; through forgetting, archive information that still has high query value can be more effectively screened out, reducing interference with outdated information; avoiding the storage and display of a large amount of outdated and no longer needed information, and improving performance.

[0086] By obtaining the forgetting factor weight and applying it to all archive information with high frequency retrieval, Among them, the calculation of the forgetting factor weight takes into account the time difference between the current time node and the last high-frequency query time node, and combines it with the half-life, that is, how many days are needed to reduce the weight of the retrieval feature by half; making the forgetting process related to time. The longer the time, the lower the relevance and retrieval priority of the archive information, avoiding the inefficiency of query caused by long-term accumulation of information, so that high-frequency retrieval features can be dynamically adjusted according to user behavior patterns to maintain the latest and most relevant information.

[0087] For high-frequency search profile information that already belongs to a certain user, the high-frequency search feature weights of the profile type will be re-obtained and updated for storage. This means that when a user queries a profile multiple times, the relevant weights will be recalculated based on the latest query data. The updated weights ensure that they remain highly sensitive to user behavior and accurately respond to user preferences and needs. Through continuous updates, the relevance of high-frequency search profiles is enhanced, improving the accuracy of retrieval.

[0088] For high-frequency search profile information that does not belong to a certain user, new high-frequency search feature weights will be assigned to the profile information, and the search keyword sequence will be re-sorted according to these weights. Even non-user-related profile information can be sorted according to its current high-frequency search feature weights, thereby ensuring that the search efficiency of both new and old users can be optimized. Through intelligent sorting, high-frequency search profiles related to users can be better pushed, further improving the user experience.

[0089] Through the query mechanism, that is, the high-frequency retrieval feature forgetting mechanism, the intelligent level of archival information management is effectively improved by dynamically adjusting the weight and storage strategy of archival information. Specifically, with the introduction of the forgetting mechanism of high-frequency retrieval archival information, it can avoid irrelevant or outdated information from affecting the query results, thereby improving the response speed and accuracy of the query; each time the user searches, he will only focus on the latest and most relevant archival information, reducing unnecessary retrieval burden; by forgetting the retrieval information, the storage occupancy of useless or infrequently used archives can be effectively reduced, and the use of storage resources can be optimized; in this way, more storage space can be used to process newly emerging important data, thereby improving overall efficiency; updating and re-assigning the feature weights of high-frequency retrieval archives enables the priority of archives to be dynamically adjusted according to the behavior patterns of each user, providing personalized retrieval results.

[0090] This not only improves the user experience, but also helps to understand user needs and behaviors, thereby providing users with content that is more in line with their interests; over time, it can automatically adjust the impact of the forgetting factor to ensure that archival information that has not been queried for a long time does not take up too many resources, while maintaining attention to new information. It can intelligently learn user retrieval behavior and continuously optimize the weight and order of archives according to demand, making it more and more intelligent and adaptive.

[0091] In a preferred embodiment of the present invention, specifically, it also includes: Based on the retrieval results output by the archive retrieval model, the retrieved archive information data is compared with the historical archive information data to obtain a comparison result; deviation data is obtained according to the comparison result to form a deviation data set; the deviation data set is iteratively analyzed to obtain interference factors of the deviation data; According to the interference factors of the deviation data, a secondary semantic analysis is performed on each archival information data item in the archival information data set to obtain accurate archival information data; the accurate archival information data is input into the archival retrieval model again for training until the archival retrieval model outputs accurate archival information data with a rapid response.

[0092] In the embodiment of the present invention, first, the archive retrieval model is used to obtain the current search results, which are the archive information returned when the user queries; the retrieved archive information is compared with the historical archive information data. The historical archive information data may include previous search results and recorded archive information; the purpose of the comparison is to find the difference or deviation between the current search results and the historical data to evaluate the accuracy of the current model; through the comparison, the errors or inaccuracies in the search results can be identified, which helps to find potential problems or deviations and then optimize; this comparison helps to improve the stability of the search, ensure that it can adapt to different data sources and maintain high efficiency.

[0093] According to the comparison results, the deviation data is extracted. Deviation data refers to data items with significant differences between the retrieval results and historical data, which may be model errors or anomalies in historical data; these deviation data are collected and aggregated into a deviation data set to provide basic data for subsequent analysis; aggregation of deviation data helps to identify common problems in the retrieval model and provides a clear direction for subsequent analysis and model adjustment. Through centralized processing of deviation data, different types of errors or deviations can be classified and summarized, improving the efficiency of error location and adjustment.

[0094] Iteratively analyze the biased data set to find interference factors in the data that affect the accuracy of the model. Interference factors may include inconsistency of the data itself, data quality issues, and improper feature selection; this iterative analysis usually uses model training technology in machine learning to gradually identify the main factors that affect retrieval accuracy; iterative analysis can continuously optimize the identification of interference factors through multiple data processing and feedback, making it gradually more adaptable to complex patterns in the data; by finding interference factors, it can better identify which data or features need to be adjusted, thereby reducing the negative impact on the model output.

[0095] According to the interference factors analyzed, a secondary semantic analysis is performed on each data item in the archival information data set. Semantic analysis refers to understanding text data and extracting relevant information through natural language processing (NLP) technology. By deeply understanding the true meaning of each archival data, interference factors are removed and more accurate archival information is extracted; secondary semantic analysis can understand the actual content of the data from a deeper level, avoiding misunderstandings or incorrect extraction; semantic analysis eliminates ambiguity in the data, making the understanding of each archival data more accurate, thereby improving the accuracy of retrieval.

[0096] The accurate archival information data obtained through the secondary semantic analysis will be re-input into the archival retrieval model as new training data; through training on these accurate data, the model can further optimize its retrieval capabilities, adjust its algorithms and parameters, and gradually reduce errors or deviations in retrieval; the training process will be repeated until the output of the model can respond quickly and accurately provide archival information.

[0097] Through continuous training, the output of the model will become more and more accurate, and the retrieval response time for archival information will be significantly reduced; this iterative training process can ensure that the model becomes more and more adapted to actual query needs and data environment over time, further improving the user experience.

[0098] By gradually optimizing the archive retrieval model, a cyclical improvement mechanism is formed. By constantly comparing the retrieval results with historical data, deviations can be discovered and corrected in a timely manner, ensuring that the retrieval effect of the model is increasingly accurate; each iterative analysis and semantic processing makes the model more targeted when facing complex data, thereby improving the accuracy and response speed of the model; analysis of deviation data can help identify and eliminate interference factors that affect retrieval accuracy, thereby improving the accuracy of retrieval results; accurate secondary semantic analysis ensures that the true meaning of each archive data is understood, thereby reducing misjudgments and erroneous retrieval.

[0099] This feedback-based training method enables the archive retrieval model to self-learn and self-improve during use, adapt to different query modes and data environments, and demonstrate strong adaptability; the model can adjust its retrieval strategy based on user needs and feedback over time, thereby providing users with personalized and efficient query services; it can quickly respond to user retrieval needs and provide accurate archival information. This rapid response and high-precision retrieval greatly enhances the user experience; users do not need to repeatedly adjust retrieval conditions or filter out irrelevant information, as it can automatically identify their needs and accurately provide relevant content.

[0100] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation mode. The above-mentioned specific implementation mode is merely illustrative and not restrictive. Under the guidance of this embodiment, ordinary technicians in this field can also make more forms of equivalent embodiments, all of which are within the protection of this embodiment.

Claims

1. An intelligent archival information retrieval method based on semantic analysis, characterized in that: The following steps are involved: Scan and collect various archival information data to obtain archival information data sets; Performing semantic analysis on the archival information data set to obtain archival information retrieval keywords; Based on the archive information search keyword, the first letter of each word of the archive information search keyword is sorted in sequence to obtain a search keyword sequence list; Based on the archival information data set, the archival information data set is classified and divided to obtain a plurality of archival information retrieval sections; different retrieval verification methods are set according to archival information retrieval sections of different levels to enable the archival retrieval sections to respond safely; The archival information retrieval section and the search keyword sequence list are constructed into a training set to obtain the archival retrieval model; the archival retrieval model outputs a retrieval recognition result representing an archival information data item, and at least one initial letter of a search keyword in the search keyword sequence list is input into the archival retrieval model to output a retrieval result representing archival information data.

2. The intelligent archival information retrieval method based on semantic analysis according to claim 1 is characterized in that: Scan and collect a variety of archival information data to obtain archival information data sets, including: Scan and collect a variety of archival information data to obtain a pre-collected archival information image data set; adjust the contrast, brightness and hue of the image in the pre-collected archival information image data set to obtain a clear scanned image data set; Denoising the image background noise in the clear scanned image data set to obtain a prominent image data set; the prominent image data set indicates that the text display of each image data item in the data set is more prominent and obvious; removing the blurred, ghosted or interfering noise parts of each image data item in the prominent image data set to obtain a clear image data set; The clear image data set is binarized to convert the clear image data into a black and white background color to obtain a binary image data set; the text in the binary image data set is processed by optical character recognition technology to obtain an archival information data set; the text data of the archival information data set is loaded to store the archival information text data to be intelligently retrieved by the archival retrieval model.

3. The intelligent archival information retrieval method based on semantic analysis according to claim 2 is characterized in that: Based on the archival information data set, the archival information data set is classified and divided to obtain multiple archival information retrieval sections, including: Classifying and dividing the archival information data set to obtain a plurality of archival information data subsets; respectively grouping the plurality of archival information data subsets into a section, and labeling and annotating the section to obtain a plurality of archival information retrieval sections; The multiple archival information retrieval sections are searched and graded to obtain archival information retrieval sections of different levels; the archival information retrieval sections of different levels include top secret archival information retrieval sections, confidential archival information retrieval sections, internal archival information retrieval sections and public archival information retrieval sections; and specific retrieval sections of archival information data are obtained according to the first letters of the search keywords input in sequence.

4. The intelligent archival information retrieval method based on semantic analysis according to claim 3 is characterized in that: Different search verification methods are set according to different levels of archive information retrieval sections to ensure that the archive retrieval section responds safely, including: Based on the specific search section of the archival information data, when the archival information required by the search user is in the public archival information section, the public archival information section stores the archival basic information data and can be directly opened to the public without identity verification or any authority control; When in the internal archive information section, only internal employees can view it, and user identity verification is required; when in the confidential archive information section, further security permission verification is required before access; when in the top secret archive information section, multiple identity verification is required; Based on the access rights of each archival information section, each archival information section responds according to the verification information to obtain confidential archival information data at different levels.

5. The intelligent archival information retrieval method based on semantic analysis according to claim 4 is characterized in that: Based on the access rights of each archive information section, each archive information section responds according to the verification information to obtain different levels of confidential archive information data, including: By retrieving the user's identity and authentication, only legitimate users can access the archive information; when the user is successfully authenticated, access rights are controlled based on the user's identity and authority and the level of the archive to determine whether there is authority to access the archive information data corresponding to the high-level archive information section; Different verification methods are set according to different archive levels; and the access permission verification strength is gradually increased to obtain corresponding verification results; according to the verification results, the searching user obtains the corresponding archive information data, so that the archive retrieval model safely responds to the archive information retrieval behavior.

6. The intelligent archival information retrieval method based on semantic analysis according to claim 5 is characterized in that: The archive retrieval model includes: The archive information retrieval section and the retrieval keyword sequence list are used to construct a data set to generate structure data, and the structure data is encoded into sequence data to train the archive retrieval model; Inputting the sequence data into the archive retrieval model; the archive retrieval model comprises an input layer, a first hidden layer, a second hidden layer, a third hidden layer and an output layer, transmitting the intermediate representation data of multiple hidden layers to the output layer, and the output layer outputting the retrieval recognition result representing the archive information data item; At least one initial letter of a search keyword in the search keyword sequence list is input into the archive retrieval model, and the output layer outputs the archive information data retrieval result.

7. The intelligent archival information retrieval method based on semantic analysis according to claim 6 is characterized in that: According to the user's archive retrieval and browsing situation, the corresponding archive information after retrieval is extracted to form the user's query mechanism; through the user's periodic archive retrieval behavior, the high frequency of retrieval users is summarized; By analyzing the high-frequency search attribute characteristics of the search users, the archive retrieval model is made to interact with the query mechanism.

8. The intelligent archival information retrieval method based on semantic analysis according to claim 7 is characterized in that: The query mechanism includes The first search behavior data is obtained by analyzing the keywords, query frequency, search time and search order of the search users for the archive information; the second search behavior data is obtained by analyzing the user's stay time and click count when browsing the archive; By using the first search behavior data and the second search behavior data, a high-frequency search analysis is performed on the search user behavior to extract the user's high-frequency search features; Based on the extracted high-frequency search features, a high-frequency search feature weight is set, which represents the high-frequency query degree of the search user for the archives within the time period, that is, ; In the formula, AF represents the query frequency measurement parameter, T represents the period, ci represents the file information type browsed within the T period, x represents the number of files browsed within the T period, W i Represents the weight of high-frequency retrieval features.

9. The intelligent archival information retrieval method based on semantic analysis according to claim 8 is characterized in that: Based on the high-frequency query level, all high-frequency searches of the archive type are forgotten to obtain the forgetting factor of the search; Obtain the forgetting factor weight, and add the forgetting factor weight to the archive information with high frequency retrieval by the user; where F(x) represents the forgetting factor weight; t represents the time node from the current time node to the high frequency retrieval feature assignment time node, and f represents the half-life, that is, the high frequency retrieval feature forgetting of the model needs to last at least f days; When the archive information already belongs to the high-frequency search archive information of the search user, the high-frequency search feature weight of the archive type will be re-acquired to update the high-frequency search feature weight storage; When the high-frequency search archive information does not belong to the searching user, the archive information is assigned a high-frequency search feature weight, and the search keyword sequence list is re-sorted according to the size of the high-frequency search feature weight.

10. The intelligent archival information retrieval method based on semantic analysis according to claim 9 is characterized in that: Based on the retrieval results output by the archive retrieval model, the retrieved archive information data is compared with the historical archive information data to obtain a comparison result; deviation data is obtained according to the comparison result to form a deviation data set; the deviation data set is iteratively analyzed to obtain interference factors of the deviation data; According to the interference factors of the deviation data, a secondary semantic analysis is performed on each archival information data item in the archival information data set to obtain accurate archival information data; the accurate archival information data is input into the archival retrieval model again for training until the archival retrieval model outputs accurate archival information data with a rapid response.

Citation Information

Cited By

  • Data asset sharing method and system

    CN120930913A

  • A data asset sharing method and system

    CN120930913B

  • Archive retrieval result dynamic optimization method based on reinforcement learning

    CN121256144A