Electronic archive intelligent retrieval method and system

By correcting spelling errors and expanding related words through pre-trained language models, combining inverted indexes and semantic similarity models, the problem of insufficient identification of non-text format electronic files by traditional search systems is solved, and efficient and personalized non-text format content recommendations are achieved, which improves retrieval efficiency and user experience.

CN120371936AInactive Publication Date: 2025-07-25SHANDONG YANGMING NETWORK INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510465983.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional keyword search systems lack effective parsing capabilities for non-text format electronic files (such as images, voice in videos, and charts), resulting in a large amount of information being unable to be recognized and retrieved.

Method used

The pre-trained language model is used to correct spelling errors, expand related words, combine inverted indexes and semantic similarity models for correlation reordering, use uploaded information to match non-text electronic files, and personalized recommendations are made through user behavior prediction intentions.

Benefits of technology

It realizes accurate identification and recommendation of non-text format content, and improves the efficiency and user experience of electronic file retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371936A_ABST
    Figure CN120371936A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent retrieval method and system for electronic archives, and belongs to the field of electric digital data processing.The intelligent retrieval method for the electronic archives comprises the following steps that search content input by a user is received, spelling errors are automatically corrected, and related words are expanded; preliminarily screening the candidate text electronic archive set based on a reverse index mechanism to obtain an initial sorting result, performing correlation resorting and extension recall on the initial sorting result through a semantic similarity model, and outputting a final sorting result; and after determining the target text electronic archive in the final sorting result, according to the uploading information of the target text electronic archive, the method has the beneficial effects that the non-text electronic archives are matched through the determined uploading information of the target text electronic archive, and the non-text electronic archives are sorted according to the association strength, so that the efficiency of sorting the non-text electronic archives is improved. The non-text format content can be identified during keyword retrieval; and the tendency of the user is predicted, so that the use experience of the user in downloading or checking the electronic file is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electronic digital data processing, and particularly relates to an intelligent retrieval method and system for electronic files. Background Art

[0002] Electronic files record the business activities and historical processes of institutions or individuals, and rely on information technology to achieve efficient storage, retrieval, and management. They are an important carrier of information assets in the digital age.

[0003] Electronic files include various formats such as documents (e.g., PDF, Word), charts (e.g., Excel, CAD drawings), audio-visual materials (e.g., audio, video), etc. Traditional keyword retrieval mainly targets text content, but lacks effective parsing ability for non-text formats (such as charts in images, voices in videos). For example, voice and image information need to be converted into retrievable content through technologies such as speech-to-text conversion and image recognition. If no preprocessing is performed, the system can only match through file names or manually marked metadata, easily missing the actual content.

[0004] Therefore, a large amount of information in non-text formats corresponding to keyword retrieval cannot be recognized by the retrieval system and needs to be improved. Summary of the Invention

[0005] Based on this, it is necessary to provide an intelligent retrieval method and system for electronic files in view of the above problems.

[0006] An embodiment of the present invention is implemented as follows. An intelligent retrieval method for electronic files includes the following steps:

[0007] Receive the search content input by the user (such as "Contract files of Company A's Business B in 2018"), automatically correct spelling mistakes (such as "file by" → "file"), and expand relevant words (such as "contract" → "agreement", "contractual agreement");

[0008] Based on the inverted index mechanism, initially screen the candidate text electronic file set to obtain the initial sorting result, and then re-sort and expand the recall of the initial sorting result through the semantic similarity model (such as expanding the recall of "termination contract" to "dissolution agreement"), and output the final sorting result;

[0009] After determining the target text electronic file in the final sorting result, match the non-text electronic file according to the upload information of the target text electronic file, and recommend it to the user. The upload information includes upload time, upload user, contract number, and project number.

[0010] In one embodiment, the present invention provides an intelligent retrieval method for electronic files. In the steps of receiving the search content input by the user (such as "Contract files of Company A's Business B in 2018"), automatically correcting spelling mistakes (such as "file press" → "file"), and expanding related words (such as "contract" → "agreement", "contract"), it specifically includes:

[0011] Receive the search content input by the user, analyze the context of the input terms by the pre-trained language model (such as BERT) (such as after "Company A's Business B in 2018" followed by "file press"), match the preset dictionary (such as "file") in combination with the edit distance algorithm, generate candidate words and sort them based on probability (such as the confidence of "file" is higher than that of "file press"); establish a common error mapping table (such as "file press → file") based on the historical search logs to accelerate error correction;

[0012] Extract the synonyms ("contract" → "agreement / contract"), near-synonym scenario words ("sign / terms") and associated entities ("Company A" → "subsidiary company name") of the input terms, and expand the word set through co-occurrence frequency weighted fusion;

[0013] Generate the optimized search content for subsequent retrieval.

[0014] In one embodiment, the present invention provides an intelligent retrieval method for electronic files. In the steps of initially screening the candidate text electronic file set based on the inverted index mechanism to obtain the initial sorting result, and then re-ranking and expanding the recall of the initial sorting result through the semantic similarity model (such as expanding the recall of "termination contract" to "termination agreement") and outputting the final sorting result, it specifically includes:

[0015] Through the inverted index mechanism, quickly match the keywords (such as "contract / agreement / contract") corresponding to the optimized search content in the text electronic files, calculate the word frequency weight using TF-IDF (a statistical method for measuring the importance of words in a document set), and generate the initial sorting in combination with the text electronic file attributes (such as giving priority to recent time); at the same time, support boolean logic (such as "2018 AND Company A") to filter out irrelevant results;

[0016] Encode the query and the candidate text electronic files into vectors through the semantic model (such as Sentence-BERT), calculate the cosine similarity, and improve the ranking of context-related text electronic files (such as "termination contract" matching documents containing "termination clause" but not explicitly named);

[0017] Expand related words based on the knowledge graph or co-occurrence analysis (such as expanding the recall of "termination contract" to "termination agreement"), and recall potential related text electronic files not covered by the inverted index;

[0018] Fuse the inverted index score (keyword matching intensity) and semantic similarity according to weights (such as 6:4), and overlay business rules (such as top - placing high - privilege files) to output the final sorting result.

[0019] In one embodiment, the present invention provides an intelligent retrieval method for electronic files. After determining the target text electronic file in the final sorting result, according to the upload information of the target text electronic file, match non - text electronic files and recommend them to the user. The upload information includes upload time, upload user, contract number, and project number. The steps specifically include:

[0020] After determining the target text electronic file in the final sorting result, according to the number information in the upload information of the target text electronic file, screen files with the same number in the non - text database (such as picture, audio - video table) to obtain the determined non - text electronic files; at the same time, according to the upload user and upload time information in the upload information, supplement potential non - text electronic files to obtain fuzzy non - text electronic files.

[0021] Adopt priority rules to sort the non - text electronic files (determined non - text electronic files + fuzzy non - text electronic files) (such as contract number matching > time proximity > user association); calculate the semantic similarity between the non - text electronic files (such as upload tags) and the target text electronic file through an embedding model to supplement long - tail associated files.

[0022] Sort according to the association strength (exact match > rule inference > semantic extension) to generate a recommendation list (such as contract scan, signing recording), and attach an explanation of the association basis (such as "the same contract number").

[0023] In one embodiment, the present invention provides an intelligent retrieval method for electronic files. After the steps of determining the target text electronic file in the final sorting result, according to the upload information of the target text electronic file, matching non - text electronic files and recommending them to the user, where the upload information includes upload time, upload user, contract number, and project number, it further includes:

[0024] Based on the user's historical behavior (such as download frequency, query stay duration, operation path) and unstructured data (such as the text description of "urgently need to download the contract" in the log), construct hybrid features (such as click density + semantic keywords), and predict the intention tendency through a lightweight integrated model (such as XGBoost + TextCNN).

[0025] If the tendency is to download, link to the background to preload files (such as caching links in advance and verifying permissions), so that the download button and format tags on the front end are highlighted (such as "PDF - Latest Version"); if the tendency is to view, then dynamically render an interactive abstract (such as folding / unfolding key terms), hide irrelevant controls (such as the batch download option), and inject page heat zone analysis (such as automatically focusing on high - frequency browsing areas);

[0026] Continuously optimize the lightweight integration model through A / B testing (also known as a control experiment) and feedback loop (such as the misoperation withdrawal rate).

[0027] In one embodiment, the present invention provides an intelligent electronic file retrieval system, including:

[0028] An error correction and expansion module, which is used to receive the search content input by the user (such as "Contract files of Company A's Business B in 2018"), automatically correct spelling mistakes (such as "dang an" → "dang an"), and expand related words (such as "contract" → "agreement" "contract");

[0029] A result sorting module, which is used to initially screen the candidate text electronic file set based on the inverted index mechanism to obtain the initial sorting result, and then re - sort and expand the recall of the initial sorting result through a semantic similarity model (such as expanding the recall of "terminated contract" to "terminated agreement"), and output the final sorting result;

[0030] A content matching module, which is used to determine the target text electronic file in the final sorting result, and then match non - text electronic files according to the upload information of the target text electronic file and recommend them to the user. The upload information includes upload time, upload user, contract number, and project number.

[0031] In one embodiment, the present invention provides an intelligent electronic file retrieval system, and the error correction and expansion module includes:

[0032] A content error correction unit, which is used to receive the search content input by the user, analyze the context of the user - inputted terms through a pre - trained language model (such as BERT) (such as after "Company A's Business B in 2018" followed by "dang an"), combine with the edit distance algorithm to match a preset dictionary (such as "dang an"), generate candidate words and sort them based on probability (such as the confidence of "dang an" is higher than that of "dang an"); establish a common error mapping table based on historical search logs (such as "dang an → dang an") to accelerate error correction;

[0033] A content expansion unit, which is used to extract synonyms ("contract" → "agreement / contract"), near - synonym scenario words ("sign / clause") and associated entities ("Company A" → "subsidiary company name") of the input terms, and expand the word set through co - occurrence frequency weighted fusion;

[0034] The content generation unit is used to generate optimized search content for subsequent retrieval.

[0035] In one embodiment, the present invention provides an electronic archive intelligent retrieval system, wherein the result sorting module comprises:

[0036] The keyword ranking unit is used to quickly match the text electronic archives corresponding to the keywords (such as "contract / agreement / contract") in the optimized search content through the inverted index mechanism, calculate the word frequency weight using TF-IDF (a statistical method used to measure the importance of words in a document collection), and generate the initial ranking in combination with the text electronic archive attributes (such as time priority); it also supports Boolean logic (such as "2018AND A company") to filter irrelevant results;

[0037] A ranking improvement unit, which is used to encode the query and candidate text electronic archives into vectors through a semantic model (such as Sentence-BERT), calculate cosine similarity, and improve the ranking of context-related text electronic archives (such as "termination of contract" matches documents containing "termination clauses" but not explicitly named);

[0038] Content recall unit, used to expand related words based on knowledge graph or co-occurrence analysis (e.g., expand and recall "termination of agreement" based on "termination of contract"), and recall potential related text electronic archives not covered by the inverted index;

[0039] The sorting result output unit is used to fuse the inverted index score (keyword matching strength) and the semantic similarity according to the weight (such as 6:4), and superimpose business rules (such as placing high-authority files at the top) to output the final sorting result.

[0040] In one embodiment, the present invention provides an electronic archive intelligent retrieval system, wherein the content matching module comprises:

[0041] The upload information matching unit is used to determine the target text electronic archive in the final sorting result, and then select the files with the same number in the non-text database (such as pictures, audio and video tables) according to the number information in the upload information of the target text electronic archive to obtain the determined non-text electronic archive; at the same time, according to the uploading user and uploading time information in the upload information, supplement the potential non-text electronic archive to obtain the fuzzy non-text electronic archive;

[0042] The priority sorting unit is used to sort the non-text electronic archives (definite non-text electronic archives + fuzzy non-text electronic archives) using priority rules (such as contract number matching > time proximity > user association); calculate the semantic similarity between the non-text electronic archives (such as upload tags) and the target text electronic archives through the embedding model, and supplement the long-tail associated files;

[0043] A content recommendation unit, which is used to sort by the association strength (exact match > rule inference > semantic expansion), generate a recommendation list (such as contract scanned copies, signing recordings), and attach an explanation of the association basis (such as "the same contract number").

[0044] In one embodiment, the present invention provides an intelligent electronic file retrieval system, further including:

[0045] An intention prediction module, which is used to construct hybrid features (such as click density + semantic keywords) based on the user's historical behavior (such as download frequency, query stay duration, operation path) and unstructured data (such as the text description of "urgently need to download the contract" in the log), and predict the intention tendency through a lightweight integrated model (such as XGBoost + TextCNN);

[0046] A tendency assistance module, which is used to, if the tendency is to download, link to preload files in the background (such as caching links in advance, verifying permissions), so that the download button and format label on the front end are highlighted (such as "PDF - latest version"); if the tendency is to view, dynamically render an interactive abstract (such as folding / unfolding key terms), hide irrelevant controls (such as batch download options), and inject page hot zone analysis (such as automatically focusing on high - frequency browsing areas);

[0047] A model optimization module, which is used to continuously optimize the lightweight integrated model through A / B testing (also known as a control experiment) and feedback loop (such as the misoperation withdrawal rate).

[0048] Compared with the prior art, the beneficial effects of the present invention are: the present invention matches non - text electronic files through the uploaded information of the target text electronic files that have been determined, sorts the non - text electronic files according to the association strength, so that non - text format content can be recognized during keyword retrieval; predicts the tendency of users, and improves the usage experience when users download or view electronic files. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic diagram of the first - part process of an intelligent electronic file retrieval method provided by an embodiment of the present invention.

[0050] Figure 2 It is a schematic diagram of the process of search content correction and expansion provided by an embodiment of the present invention.

[0051] Figure 3 It is a schematic diagram of the process of sorting search results provided by an embodiment of the present invention.

[0052] Figure 4 It is a schematic diagram of the process of matching non - text electronic files provided by an embodiment of the present invention.

[0053] Figure 5It is a schematic diagram of the second part of the process of an intelligent electronic file retrieval method provided by an embodiment of the present invention.

[0054] Figure 6 It is a schematic diagram of the first part of an intelligent electronic file retrieval system provided by an embodiment of the present invention.

[0055] Figure 7 It is a schematic diagram of an error correction and expansion module provided by an embodiment of the present invention.

[0056] Figure 8 It is a schematic diagram of a result sorting module provided by an embodiment of the present invention.

[0057] Figure 9 It is a schematic diagram of a content matching module provided by an embodiment of the present invention.

[0058] Figure 10 It is a schematic diagram of the second part of an intelligent electronic file retrieval system provided by an embodiment of the present invention. Detailed implementation manners

[0059] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.

[0060] In one embodiment, as Figure 1 shown, an intelligent electronic file retrieval method (applied to an intelligent electronic file retrieval system) includes the following steps:

[0061] Step S1: Receive the search content input by the user (such as "Contract files of Company A's Business B in 2018"), automatically correct spelling mistakes (such as "file by" → "file"), and expand related words (such as "contract" → "agreement" "contract");

[0062] Step S2: Based on the inverted index mechanism, initially screen the candidate text electronic file set to obtain the initial sorting result, and then re-sort and expand the recall of the initial sorting result through the semantic similarity model (such as expanding the recall of "termination contract" to "dissolution agreement"), and output the final sorting result;

[0063] Step S3: After determining the target text electronic file in the final sorting result, match the non-text electronic file according to the upload information of the target text electronic file and recommend it to the user. The upload information includes upload time, upload user, contract number, and project number.

[0064] The entire process realizes the efficient search and precise recommendation of electronic files through intelligent technologies. First, after receiving the search content input by the user, the pre-trained language model (such as BERT) is used to combine context analysis and the historical error mapping table to automatically correct spelling mistakes (such as correcting "dang'an" to "archives"), and expand semantic related words based on the knowledge graph and word vector model (such as expanding "contract" to "agreement", "contract"), generating an optimized multi-dimensional query statement to ensure coverage of the user's potential needs. Subsequently, based on the inverted index mechanism, keywords are quickly matched, and the initial sorting is generated using TF-IDF weights, document attributes (such as time, title), and Boolean logic filtering; at the same time, a semantic model (such as Sentence-BERT) is introduced to vectorize and encode candidate documents, and the ranking of documents that are contextually relevant but do not explicitly match keywords is improved through cosine similarity calculation, and related words are expanded in combination with the knowledge graph (such as recalling "termination of contract" for "dissolution of agreement") to supplement potentially relevant files. Finally, the keyword matching intensity and semantic similarity are fused according to dynamic weights, and business rules (such as permission priorities) are superimposed to output the sorting result. On this basis, the system accurately matches non-text files (such as scanned copies, audio and video with the same number) according to the upload information of the target text file (such as contract number, upload time), and supplements long-tail related content by calculating the metadata similarity through fuzzy association (such as files uploaded by the user during the same period) and the semantic model. Finally, a recommendation list is generated according to the priorities of accurate matching, rule inference, and semantic expansion, and the associated basis is explained (such as "the same contract number"). The entire process realizes a full-link closed loop from input error correction, multi-dimensional search to non-text resource recommendation through semantic understanding, hybrid retrieval, and cross-modal linkage, taking into account retrieval efficiency, semantic depth, and user experience, and effectively improving the intelligent level of file management.

[0065] In one embodiment, as Figure 2 shown, for an intelligent retrieval method of electronic files, in the step S1 of receiving the search content input by the user (such as "contract files of Company A's B business in 2018"), and automatically correcting spelling mistakes (such as "dang'an" → "archives") and expanding related words (such as "contract" → "agreement", "contract"), it specifically includes:

[0066] Step S11, receive the search content input by the user, analyze the context of the user input entry (such as "Company A's B business in 2018" followed by "dang'an") through the pre-trained language model (such as BERT), match the preset dictionary (such as "archives") in combination with the edit distance algorithm, generate candidate words and sort them based on probability (such as the confidence of "archives" is higher than that of "dang'an"); establish a common error mapping table based on the historical search log (such as "dang'an → archives") to accelerate error correction;

[0067] Step S12, extracting synonyms ("contract → agreement / contract"), near-meaning scene words ("signature / clause") and related entities ("Company A → subsidiary name") of the input terms, and expanding the word set by weighted fusion of co-occurrence frequency;

[0068] Step S13, generating optimized search content for subsequent retrieval.

[0069] Steps S11-S13 optimize the user's original search content through intelligent technology. First, a pre-trained language model (such as BERT) is used in combination with context analysis and edit distance algorithm to automatically detect and correct spelling errors (such as "档按" is corrected to "档案"), and a high-frequency error mapping table is constructed based on historical search logs to accelerate the error correction process. Subsequently, the semantic association of the input terms is mined through the knowledge graph and word vector model to extract synonyms (such as "contract→agreement / contract"), scenario synonyms (such as "sign / terms") and entity related words (such as "A company→subsidiary name"), and the word set is expanded based on the weighted fusion of co-occurrence frequency to cover the user's potential intentions. Finally, the core words and expanded words after error correction are integrated into a structured query statement (such as "2018AND(contract OR agreement)AND A company"), taking into account both precise matching and semantic generalization capabilities, and providing highly robust input conditions for subsequent retrieval. It effectively solves the problems of spelling errors, terminology differences and expression diversity, and significantly improves the search recall rate and the fit with user intentions.

[0070] In one embodiment, Figure 3 As shown, an electronic archive intelligent retrieval method, the step S2, preliminarily screens the candidate text electronic archive set based on the inverted index mechanism to obtain the initial sorting result, and then re-sorts and expands the initial sorting result through the semantic similarity model (such as expanding and recalling "termination of agreement" based on "termination of contract"), and outputs the final sorting result step, specifically including:

[0071] Step S21, through the inverted index mechanism, quickly match the text electronic archives corresponding to the keywords (such as "contract / agreement / contract") in the optimized search content, use TF-IDF (a statistical method for measuring the importance of words in a document collection) to calculate the word frequency weight, and combine the text electronic archive attributes (such as time priority) to generate an initial sort; at the same time, support Boolean logic (such as "2018AND A company") to filter irrelevant results;

[0072] Step S22, encoding the query and the candidate text electronic archives into vectors through a semantic model (such as Sentence-BERT), calculating cosine similarity, and improving the ranking of context-related text electronic archives (such as "termination of contract" matches documents containing "termination clause" but not explicitly named);

[0073] Step S23: Based on the knowledge graph or co-occurrence analysis, expand related terms (e.g., expand and recall "terminate the agreement" based on "terminate the contract"), and recall potential relevant text electronic files not covered by the inverted index.

[0074] Step S24: Fuse the inverted index score (keyword matching intensity) and semantic similarity according to weights (such as 6:4), and overlay business rules (such as top-rank high-authority files), and output the final sorting result.

[0075] Steps S21 to S24 achieve efficient and accurate matching of electronic files through multimodal retrieval and sorting technologies. First, based on the inverted index, quickly screen documents directly associated with the optimized keywords (such as "contract / agreement"), and generate an initial sorting by combining TF-IDF weights, document attributes (such as title hit, recent time first), and Boolean logic (such as "2018 AND Company A") to ensure millisecond-level response efficiency; subsequently, encode the query and candidate documents into high-dimensional vectors through a semantic model (such as Sentence-BERT), calculate the cosine similarity, improve the ranking of content with implicit semantic associations (such as "termination clause" matching "terminate the contract"), and supplement long-tail documents not covered by the inverted index based on the knowledge graph to expand related terms (such as "terminate the agreement") to solve the problem of missed detection caused by differences in keyword expressions. Finally, the system fuses the inverted index score (focusing on keyword matching intensity) and semantic similarity (focusing on context relevance) according to dynamic weights (such as 6:4), and at the same time overlays business rules (such as top-rank high-authority files, downgrade sensitive files), and outputs a sorting result that takes into account both efficiency and semantic depth. On the basis of ensuring real-time performance, significantly improve the accuracy in complex scenarios (such as fuzzy queries, polysemous ambiguities), and at the same time support flexible adaptation to industry-specific rules to meet diverse business needs.

[0076] In one embodiment, as Figure 4 shown, for an intelligent retrieval method of electronic files, in step S3, after determining the target text electronic file in the final sorting result, according to the upload information of the target text electronic file, match non-text electronic files and recommend them to the user. The upload information includes upload time, upload user, contract number, and project number. Specifically, it includes:

[0077] Step S31: After determining the target text electronic file in the final sorting result, according to the number information in the upload information of the target text electronic file, screen files with the same number in the non-text database (such as picture, audio / video table) to obtain the determined non-text electronic files; at the same time, according to the upload user and upload time information in the upload information, supplement potential non-text electronic files to obtain fuzzy non-text electronic files.

[0078] Step S32: Using the priority rule, sort the non-text electronic files (determined non-text electronic files + fuzzy non-text electronic files) (e.g., contract number matching > time proximity > user association); calculate the semantic similarity between the non-text electronic files (such as uploaded tags) and the target text electronic file through the embedding model, and supplement the long-tail associated files.

[0079] Step S33: Sort according to the association strength (exact match > rule inference > semantic extension) to generate a recommendation list (such as contract scans, signing recordings), and attach an explanation of the association basis (such as "the same contract number").

[0080] Steps S31 to S33 achieve seamless integration of text and non-text electronic files through multi-modal association and intelligent recommendation technology. Based on the upload information of the target text file (such as contract number, project number, upload user, and time), accurately screen files with the same number (such as contract scans) in the non-text database (such as picture, audio, and video tables) as the core recommended content. At the same time, supplement fuzzy-associated non-text files (such as meeting minutes or on-site photos) according to the user's operation habits (such as upload records of the same account) and time proximity (such as files uploaded within ±3 days) to form a "determined + potential" candidate set. On this basis, adopt a hybrid strategy of rules and models to optimize the recommendation quality: ensure that high-confidence results are displayed first through the priority rule (contract number matching > time proximity > user association), and use the embedding model to calculate the semantic similarity between non-text metadata (such as file tags, description texts) and the target text to mine long-tail content that is not explicitly associated (such as signing recordings without a number but with a matching description). Finally, sort the recommendation list according to the association strength hierarchy of "exact match > rule inference > semantic extension", attach an interpretable association basis (such as "the same contract number", "upload time proximity"), and at the same time support user interaction feedback (such as clicking "irrelevant" to adjust the weight). It breaks the retrieval barrier between text and non-text files, improves the efficiency of cross-modal information integration, and helps users quickly obtain a complete business evidence chain.

[0081] In one embodiment, as Figure 5 shown, for an intelligent retrieval method of electronic files, after step S3, after determining the target text electronic file in the final sorting result, match the non-text electronic file according to the upload information of the target text electronic file and recommend it to the user. The upload information includes upload time, upload user, contract number, and project number. The steps further include:

[0082] Step S4: Based on the user's historical behavior (such as download frequency, query stay duration, operation path) and unstructured data (such as the text description of "urgently need to download the contract" in the log), construct a hybrid feature (such as click density + semantic keywords), and predict the intention tendency through a lightweight integrated model (such as XGBoost + TextCNN).

[0083] Step S5, if the tendency is to download, preload files in the background in association (such as caching links in advance and verifying permissions), so that the download button and format tags on the front end are highlighted (such as "PDF - Latest Version"); if the tendency is to view, dynamically render an interactive abstract (such as folding / unfolding key terms), hide irrelevant controls (such as batch download options), and inject page hot zone analysis (such as automatically focusing on high - frequency browsing areas);

[0084] Step S6, continuously optimize the lightweight integration model through A / B testing (also known as a controlled experiment) and feedback loop (such as the rate of withdrawal for incorrect operations).

[0085] Steps S4 to S6 achieve a personalized closed - loop for the search experience through user intention recognition and dynamic interaction optimization. Based on the user's historical behavior (such as download frequency, page stay duration, operation path) and unstructured log data (such as text descriptions like "urgently need to download"), construct hybrid features (such as combining click density and semantic keywords), and use a lightweight integration model (such as XGBoost + TextCNN) to predict the user's intention tendency in real - time. If the tendency is to download, the background pre - loads the file link in association and verifies permissions, and the download button and format tags on the front end are highlighted (such as "PDF - Latest Version"); if the tendency is to view, dynamically render an interactive abstract (such as folding / unfolding key terms), hide redundant controls (such as batch download options), and focus on the high - frequency browsing area of the page. At the same time, the system compares the differences in user behavior (such as download conversion rate, rate of withdrawal for incorrect operations) through A / B testing (such as the control group retaining the original interaction logic), and combines feedback data (such as search terms manually corrected by the user) to continuously optimize the model weights and policy rules (such as adjusting the feature fusion ratio), forming a closed - loop link of "prediction → response → verification → iteration". This process drives the model to adaptively evolve based on user behavior, reduces interaction friction while improving operation efficiency, and ensures high availability and user satisfaction of the search service in dynamic scenarios.

[0086] In one embodiment, as Figure 6 shown, an intelligent electronic file retrieval system includes:

[0087] An error correction and expansion module 1, which is used to receive the search content input by the user (such as "Contract files of Company A's Business B in 2018"), automatically correct spelling mistakes (such as "dang an" → "file"), and expand relevant words (such as "contract" → "agreement", "contractual agreement");

[0088] A result sorting module 2, which is used to initially screen a candidate set of text electronic files based on the inverted index mechanism to obtain an initial sorting result, and then re - sort and expand the recall of the initial sorting result through a semantic similarity model (such as expanding the recall of "termination contract" to "dissolution agreement"), and output the final sorting result;

[0089] A content matching module 3, which is used to determine the target text electronic file in the final sorting result, and match non-text electronic files according to the upload information of the target text electronic file, and recommend them to the user. The upload information includes upload time, upload user, contract number, and project number.

[0090] In the result sorting module 2, cross-document type correlation analysis can also be added. For example, signature handwriting (image) or signing recording (audio) in scanned documents can be identified based on the contract number, and a semantic correlation channel between text and non-text can be established through a multi-modal embedding model to improve the extended recall coverage rate.

[0091] In one embodiment, as Figure 7 shown, an intelligent electronic file retrieval system, the error correction and expansion module 1 includes:

[0092] A content error correction unit 11, which is used to receive the search content input by the user, analyze the context of the user input term (such as "A company's B business in 2018" followed by "file press") through a pre-trained language model (such as BERT), combine the edit distance algorithm to match a preset dictionary (such as "file"), generate candidate words and sort them based on probability (such as the confidence of "file" is higher than that of "file press"); establish a common error mapping table (such as "file press → file") based on historical search logs to accelerate error correction;

[0093] A content expansion unit 12, which is used to extract synonyms ("contract → agreement / contract"), near-synonym scenario words ("sign / clause") and associated entities ("A company → subsidiary name") of the input term, and expand the word set through co-occurrence frequency weighted fusion;

[0094] A content generation unit 13, which is used to generate optimized search content for subsequent retrieval.

[0095] In the content expansion unit 12, a session state awareness mechanism can also be introduced. For example, after the user continuously searches for "contract termination", the weights of scenario words such as "termination of contract / breach of contract" are automatically strengthened, and restrictive terms (such as "force majeure clause") are supplemented based on the domain knowledge graph (such as company law articles).

[0096] In one embodiment, as Figure 8 shown, an intelligent electronic file retrieval system, the result sorting module 2 includes:

[0097] The keyword ranking unit 21 is used to quickly match the text electronic archives corresponding to the keywords (such as "contract / agreement / contract") in the optimized search content through the inverted index mechanism, calculate the word frequency weight using TF-IDF (a statistical method for measuring the importance of words in a document collection), and generate an initial ranking in combination with the text electronic archive attributes (such as time priority); and also supports Boolean logic (such as "2018AND A company") to filter irrelevant results;

[0098] A ranking improvement unit 22, used to encode the query and the candidate text electronic archives into vectors through a semantic model (such as Sentence-BERT), calculate cosine similarity, and improve the ranking of context-related text electronic archives (such as "termination of contract" matches documents containing "termination clause" but not explicitly named);

[0099] A content recall unit 23 is used to expand associated words based on a knowledge graph or co-occurrence analysis (e.g., to expand and recall “termination of agreement” based on “termination of contract”), and to recall potential related text electronic archives not covered by the inverted index;

[0100] The sorting result output unit 24 is used to fuse the inverted index score (keyword matching strength) and the semantic similarity according to the weight (such as 6:4), and superimpose the business rules (such as placing high-authority files on the top), and output the final sorting result.

[0101] In the keyword ranking unit 21, a near real-time index update mechanism can also be introduced in the inverted index stage, combined with a dynamic knowledge graph (such as changes in contract performance status), to automatically associate the latest data of the business system (such as "termination of contract" with the current invalid status mark) to improve timeliness.

[0102] In one embodiment, Figure 9 As shown, an electronic archive intelligent retrieval system, the content matching module 3 includes:

[0103] The upload information matching unit 31 is used to determine the target text electronic archive in the final sorting result, and then select the files with the same number in the non-text database (such as pictures, audio and video tables) according to the number information in the upload information of the target text electronic archive to obtain the determined non-text electronic archive; at the same time, according to the uploading user and uploading time information in the upload information, supplement the potential non-text electronic archive to obtain the fuzzy non-text electronic archive;

[0104] The priority sorting unit 32 is used to sort the non-text electronic archives (definite non-text electronic archives + fuzzy non-text electronic archives) using priority rules (such as contract number matching > time proximity > user association); calculate the semantic similarity between the non-text electronic archives (such as upload tags) and the target text electronic archives through an embedding model, and supplement the long-tail associated files;

[0105] A content recommendation unit 33, which is used to sort by the strength of association (accurate matching > rule inference > semantic extension), generate a recommendation list (such as contract scanned copies, signing recordings), and attach an explanation of the basis for association (such as "the same contract number").

[0106] In the priority sorting unit 32, real-time business status (such as contract performance exception alerts) can also be integrated, the rule weights can be dynamically adjusted (such as automatic weight increase for non-text files "near expiration"), and domain preference features can be injected based on the user profile (such as the legal affairs role).

[0107] In addition, a sensitive information detection module (such as red seal recognition, personal privacy desensitization) can be added, combined with a dynamic desensitization strategy and a permission-level exposure mechanism to ensure that the recommended content complies with data security specifications.

[0108] In one embodiment, as Figure 10 shown, an intelligent retrieval system for electronic files further includes:

[0109] An intention prediction module 4, which is used to construct hybrid features (such as click density + semantic keywords) based on the user's historical behavior (such as download frequency, query stay duration, operation path) and unstructured data (such as the text description "urgently need to download a contract" in the log), and predict the intention tendency through a lightweight integrated model (such as XGBoost + TextCNN);

[0110] A tendency assistance module 5, which is used to, if the tendency is to download, link to preload files in the background (such as pre-cache links in advance, verify permissions), so that the download button and format tags on the front end are highlighted (such as "PDF - latest version"); if the tendency is to view, dynamically render an interactive summary (such as folding / unfolding of key terms), hide irrelevant controls (such as batch download options), and inject page hot zone analysis (such as automatically focusing on high-frequency browsing areas);

[0111] A model optimization module 6, which is used to continuously optimize the lightweight integrated model through A / B testing (also known as a controlled experiment) and a feedback loop (such as the rate of operation withdrawal due to misoperation).

[0112] In the model optimization module 6, a causal intervention module can also be implanted during result generation, and the implicit causal chain between user operations and document value can be identified through a structural causal model (such as the actual impact of frequently downloaded documents on business decisions), and counterfactual recommendation reasons can be generated (such as "if the confidentiality clause is not added, the risk will increase by 37%").

[0113] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0114] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0115] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

[0116] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0117] In addition, it should be understood that although this specification is described according to implementation manners, not every implementation manner only includes an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other implementation manners that can be understood by those skilled in the art.

Claims

1. An intelligent retrieval method for electronic files, characterized in that, The intelligent retrieval method for electronic files includes the following steps: Receive the search content input by the user, automatically correct spelling mistakes, and expand relevant words; Based on the inverted index mechanism, initially screen the candidate text electronic file set to obtain the initial sorting result, and then re-rank and expand the recall of the initial sorting result through the semantic similarity model, and output the final sorting result; After determining the target text electronic file in the final sorting result, match the non-text electronic file according to the upload information of the target text electronic file and recommend it to the user. The upload information includes upload time, upload user, contract number, and project number.

2. The intelligent retrieval method for electronic files according to claim 1, characterized in that In the step of receiving the search content input by the user, automatically correct spelling mistakes, and expand relevant words, it specifically includes: Receive the search content input by the user, analyze the context of the input terms by the pre-trained language model, match the preset dictionary in combination with the edit distance algorithm, generate candidate words and sort them based on probability; establish a common error mapping table based on the historical search log to accelerate error correction; Extract the synonyms, near-synonym scenario words, and associated entities of the input terms, and expand the word set through co-occurrence frequency weighted fusion; Generate the optimized search content for subsequent retrieval.

3. The intelligent retrieval method for electronic files according to claim 1, characterized in that, In the step of initially screening the candidate text electronic file set based on the inverted index mechanism to obtain the initial sorting result, and then re-rank and expand the recall of the initial sorting result through the semantic similarity model, and output the final sorting result, it specifically includes: Through the inverted index mechanism, quickly match the text electronic files corresponding to the keywords in the optimized search content, calculate the word frequency weight using TF-IDF, and generate the initial sorting in combination with the attributes of the text electronic files; at the same time, support Boolean logic to filter out irrelevant results; Encode the query and the candidate text electronic files into vectors through the semantic model, calculate the cosine similarity, and improve the ranking of the context-related text electronic files; Expand the associated words based on the knowledge graph or co-occurrence analysis, and recall the potentially relevant text electronic files not covered by the inverted index; Fuse the inverted index score and the semantic similarity according to the weight, and superimpose the business rules to output the final sorting result.

4. The intelligent retrieval method for electronic files according to claim 1, characterized in that In the step of determining the target text electronic file in the final sorting result, matching the non-text electronic file according to the upload information of the target text electronic file and recommending it to the user. The upload information includes upload time, upload user, contract number, and project number, it specifically includes: After determining the target text electronic file in the final sorting result, screen the files with the same number in the non-text database according to the number information in the upload information of the target text electronic file to obtain the determined non-text electronic file; at the same time, according to the upload user and upload time information in the upload information, supplement the potential non-text electronic files to obtain the fuzzy non-text electronic files; Adopt the priority rule to sort the non-text electronic files; calculate the semantic similarity between the non-text electronic files and the target text electronic files through the embedding model, and supplement the long-tail associated files; Sort according to the association strength, generate the recommendation list, and attach the description of the association basis.

5. The intelligent retrieval method for electronic files according to any one of claims 1 to 4, characterized in that After the target text electronic archive is determined in the final sorting result, the non-text electronic archive is matched according to the upload information of the target text electronic archive and recommended to the user, and the upload information includes the upload time, the upload user, the contract number, and the project number. The following steps are also included: Based on user historical behavior and unstructured data, hybrid features are constructed to predict intention tendency through a lightweight integrated model; If you prefer to download, the backend will preload the file, so that the download button and format label are highlighted on the frontend. If you prefer to view, the interactive summary will be dynamically rendered, irrelevant controls will be hidden, and page hot zone analysis will be injected. Continuously optimize the lightweight integration model through A / B testing and feedback loop.

6. An intelligent retrieval system for electronic files, characterized in that, The electronic archive intelligent retrieval system comprises: The error correction expansion module is used to receive the search content input by the user, automatically correct spelling errors, and expand related words; The result sorting module is used to preliminarily screen the candidate text electronic archive collection based on the inverted index mechanism to obtain the initial sorting result, and then re-sort and expand the initial sorting result through the semantic similarity model to output the final sorting result; The content matching module is used to match non-text electronic archives and recommend them to users based on the upload information of the target text electronic archives after determining the target text electronic archives in the final sorting results. The upload information includes the upload time, uploading user, contract number, and project number.

7. The intelligent electronic file retrieval system according to claim 6, characterized in that The error correction extension module includes: The content error correction unit is used to receive the search content input by the user, analyze the context of the user input term through the pre-trained language model, match the preset dictionary with the edit distance algorithm, generate candidate words and sort them based on probability; establish a common error mapping table based on historical search logs to accelerate error correction; Content expansion unit, used to extract synonyms, near-meaning scene words and related entities of input terms, and expand the word set by weighted fusion of co-occurrence frequency; The content generation unit is used to generate optimized search content for subsequent retrieval.

8. The intelligent electronic file retrieval system according to claim 6, wherein The result sorting module includes: The keyword ranking unit is used to quickly match the text electronic archives corresponding to the keywords in the optimized search content through the inverted index mechanism, calculate the word frequency weight using TF-IDF, and generate the initial ranking based on the text electronic archive attributes; it also supports Boolean logic to filter irrelevant results; A ranking improvement unit, used for encoding the query and candidate text electronic archives into vectors through semantic modules, calculating cosine similarity, and improving the ranking of context-related text electronic archives; A content recall unit, which is used to expand associated words based on knowledge graph or co-occurrence analysis and recall potential related text electronic archives not covered by the inverted index; The sorting result output unit is used to fuse the inverted index score and the semantic similarity according to the weight, and superimpose the business rules to output the final sorting result.

9. The intelligent electronic file retrieval system according to claim 6, wherein The content matching module includes: An upload information matching unit, which is used to determine the target text electronic file in the final sorting result, and then screen for files with the same number in the non-text database according to the number information in the upload information of the target text electronic file to obtain the determined non-text electronic file; at the same time, supplement potential non-text electronic files according to the upload user and upload time information in the upload information to obtain fuzzy non-text electronic files; A priority sorting unit, which is used to sort the non-text electronic files by adopting priority rules; calculate the semantic similarity between the non-text electronic files and the target text electronic file through an embedding model, and supplement long-tail associated files; A content recommendation unit, which is used to sort by association strength, generate a recommendation list, and attach an explanation of the association basis.

10. The intelligent electronic file retrieval system according to any one of claims 6 to 9, characterized in that, The electronic file intelligent retrieval system further includes: An intention prediction module, which is used to construct hybrid features based on the user's historical behavior and unstructured data, and predict the intention tendency through a lightweight integrated model; A tendency assistance module, which is used to, if the tendency is to download, link to preload files in the background, so that the download button and format label are highlighted at the front end; if the tendency is to view, dynamically render an interactive abstract, hide irrelevant controls, and inject page hot zone analysis; A model optimization module, which is used to continuously optimize the lightweight integrated model through A / B testing and feedback loop.

Citation Information

Cited By

  • Intelligent search engine system and method based on NLP and vector hybrid retrieval

    CN121029791A

  • An intelligent search engine system and method based on hybrid NLP and vector retrieval

    CN121029791B

  • Classification and storage method, device and system for electronic archives

    CN121092713A

  • Medical consumable semantic vectorization matching method and system fusing registry number constraint

    CN121166759A

  • Fusion registration certificate number constraint medical consumable semantic vectorization matching method and system

    CN121166759B