Retrieval method and system based on document database

By constructing data documents, keyword analysis, field importance evaluation and index construction in the document database, the problem of inefficiency of traditional data retrieval methods is solved, fast and accurate data retrieval and information extraction are achieved, and user experience and work efficiency are improved.

CN120216670APending Publication Date: 2025-06-27HUANENG INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510397183.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional data retrieval methods are inefficient when processing large-scale unstructured or semi-structured data, and it is difficult to meet users' requirements for real-time and efficiency.

Method used

By constructing data documents and storing them in the document database, keyword analysis and field importance evaluation are carried out, index relationships are constructed, and query conditions are formulated according to search needs to match data, and finally a search information report is generated.

Benefits of technology

It improves the accessibility and maintainability of data, realizes fast and accurate data retrieval, generates search information reports that meet user needs, and improves user experience and work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216670A_ABST
    Figure CN120216670A_ABST
Patent Text Reader

Abstract

The invention discloses a retrieval method and system based on a document database, and the method comprises the steps: obtaining data needing to be retrieved, constructing a data document according to a document form, and storing the data document in the document database; determining fields and field information contained in the data document, and extracting keywords in each field; performing importance evaluation on each field based on the keyword and the field information, and determining a key field in the fields according to an evaluation result; constructing an index relationship for the key field, and determining a retrieval demand condition in the retrieval demand; formulating a retrieval query condition based on the retrieval demand condition, performing data information matching in a document database according to the retrieval query condition, and taking a successfully matched data document as a retrieval result; and performing information extraction on the retrieval result, analyzing and sorting to generate a retrieval information report. According to the retrieval method based on the document database, the technologies of document storage, document retrieval, index construction and the like are combined, and the retrieval efficiency and the retrieval accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data retrieval, and particularly to a retrieval method and system based on a document database. Background Art

[0002] A document database is a database system for storing and managing document-type data, which focuses on processing unstructured or semi-structured data such as text, images, audio, etc. In the field of information retrieval, the retrieval method based on a document database is an important technical means for effectively retrieving information of interest to users from a large-scale document library.

[0003] However, traditional retrieval methods are usually applied to storage methods of unstructured data or data with low structuring degree, such as plain text files or simple database tables, resulting in low retrieval efficiency and difficult information extraction, leading to an inflexible and inefficient data storage and retrieval mechanism. Moreover, when traditional retrieval methods process large-scale data, due to the lack of index support and efficient search algorithms, the query speed is slow, making it difficult to meet the requirements of users for real-time performance and efficiency. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a retrieval method and system based on a document database, including: Obtain the data to be retrieved, construct a data document according to the document form for the data to be retrieved, and store the constructed data document in the document database; Determine the fields and field information included in the data document in the document database, perform keyword analysis and extraction on the fields included in the data document, and determine the keywords in each field; Evaluate the importance of each field based on the keywords and field information, and determine the key fields in the field according to the evaluation results; Build an index relationship for the key fields in the document database, determine the retrieval requirements, and determine the retrieval requirement conditions based on the retrieval requirements; Formulate a retrieval query condition based on the retrieval requirement conditions, perform data information matching in the document database according to the retrieval query condition, and use the data document with successful matching as the retrieval result; Extract information from the retrieval result, analyze and organize the extracted information, and generate a retrieval information report.

[0005] Further, the determining the fields and field information included in the data document in the document database, performing keyword analysis and extraction on the fields included in the data document, and determining the keywords in each field includes: Determine the fields and field information included in the data document in the document database, and extract candidate keywords from the fields using natural language processing technology; Determine the term frequency and inverse document frequency of each candidate keyword, and calculate the product of the term frequency and the inverse document frequency to obtain the term frequency-inverse document frequency value of each candidate keyword; Determine the candidate keywords with term frequency-inverse document frequency values higher than the preset threshold as the keywords in each field.

[0006] Furthermore, the importance of each field is evaluated based on the keywords and field information, and the key fields in the field are determined according to the evaluation results, including: Determine the field information corresponding to each field, where the field information includes the query frequency and the number of queries, and determine the term frequency-inverse document frequency value of the keywords corresponding to each field; Determine the information importance value of the field information based on the query frequency and the number of queries, and respectively evaluate and obtain the first importance evaluation value of the keyword and the second importance evaluation value of the field information for the term frequency-inverse document frequency value of the keyword and the information importance value of the field information; Add the first importance evaluation value and the second importance evaluation value to calculate the comprehensive importance evaluation value of the field, and determine the field with the comprehensive importance evaluation value greater than the preset evaluation value in each field as the key field.

[0007] Furthermore, the information importance value of the field information is determined based on the query frequency and the number of queries, and the calculation formula of the information importance value of the field information is: L = α * M + β * N; where L is the information importance value of the field information, α is the preset first weight coefficient, M is the query frequency of the field, β is the preset second weight coefficient, and N is the number of queries of the field.

[0008] Furthermore, the retrieval requirement is determined, and the retrieval requirement conditions are determined based on the retrieval requirement, including: Determine the user's retrieval requirement, and split the retrieval requirement into multiple retrieval requirement conditions. The retrieval requirement conditions include the retrieval content, the retrieval range, and the retrieval time.

[0009] Furthermore, the retrieval query conditions are formulated based on the retrieval requirement conditions, and data information matching is performed in the document database according to the retrieval query conditions, and the data documents with successful matching are used as the retrieval results, including: Write the corresponding retrieval content query statement, retrieval range query statement, and retrieval time query statement based on the retrieval content, retrieval range, and retrieval time respectively, and form the retrieval query conditions according to the retrieval content query statement, retrieval range query statement, and retrieval time query statement; Input the retrieval query conditions into the document database. The document database performs a retrieval operation according to the retrieval query conditions, and determines the data documents in the document database that match the retrieval content as the retrieval results.

[0010] Furthermore, information extraction is performed on the retrieval results, and the extracted information is analyzed and sorted to generate a retrieval information report, including: Perform information extraction on the retrieval results, and perform preprocessing on the extracted information. The preprocessing includes word segmentation, stop word removal, and part-of-speech tagging; Identify key information from the preprocessed information, and clean and deduplicate the extracted key information; Classify the cleaned key information, and fill the classified key information according to the report content classification to generate a retrieval information report.

[0011] The present invention also provides a retrieval system based on a document database, including: An acquisition module, configured to acquire data to be retrieved, construct data documents according to the document form of the data to be retrieved, and store the constructed data documents in the document database; An extraction module, configured to determine the fields and field information included in the data documents in the document database, perform keyword analysis and extraction on the fields included in the data documents, and determine the keywords in each field; An evaluation module, configured to perform importance evaluation on each field based on the keywords and field information, and determine the key fields in the fields according to the evaluation results; A determination module, configured to construct an index relationship for the key fields in the document database, determine the retrieval requirements, and determine the retrieval requirement conditions based on the retrieval requirements; A matching module, configured to formulate retrieval query conditions based on the retrieval requirement conditions, perform data information matching in the document database according to the retrieval query conditions, and use the successfully matched data documents as the retrieval results; A generation module, configured to perform information extraction on the retrieval results, and analyze and sort the extracted information to generate a retrieval information report.

[0012] Compared with the prior art, the retrieval method and system based on a document database according to an embodiment of the present invention have the beneficial effects that: By constructing data documents and storing them in the document database, the present invention realizes the structured storage and management of data, and improves the accessibility and maintainability of data; Through keyword analysis and field importance evaluation, the present invention realizes the extraction and analysis of information in data documents, and helps users understand and utilize data more quickly; The present invention realizes the acceleration and optimization of the retrieval operation by constructing an index relationship and formulating retrieval requirement conditions, improving the retrieval efficiency and user experience; Based on keyword fields and retrieval requirement conditions, the present invention realizes accurate data retrieval and information extraction, generates a retrieval information report that meets the user's needs, and helps the user quickly obtain the required information; By analyzing and sorting out the information of the retrieval results, the present invention can discover the laws and trends hidden in the data and provide data-driven decision support for the user; The entire process of the present invention can be partially or fully automated. With the help of technologies such as artificial intelligence, the processing efficiency and accuracy are improved; By integrating various technical effects, the present invention provides users with higher-quality retrieval and information processing services, enhancing the user experience and work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a schematic flow structure diagram of a retrieval method based on a document database in an embodiment of the present invention; Figure 2 is a schematic composition diagram of a retrieval system based on a document database in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] The following further describes in detail the specific embodiments of the present application in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0015] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the platform or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present application.

[0016] The terms "", "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "", "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, "a plurality" means two or more.

[0017] In the description of the present application, it should be noted that, unless otherwise clearly specified and limited, the terms "installation", "connection", and "linkage" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0018] As Figure 1 shown, in the embodiment of the present application, a retrieval method based on a document database is provided, including: S100: Obtain the data to be retrieved, construct a data document according to the document form for the data to be retrieved, and store the constructed data document in the document database; S200: Determine the fields and field information included in the data document in the document database, perform keyword analysis and extraction on the fields included in the data document, and determine the keywords in each field; S300: Perform importance evaluation on each field based on the keywords and field information, and determine the key fields in the fields according to the evaluation results; S400: Construct an index relationship for the key fields in the document database, determine the retrieval requirements, and determine the retrieval requirement conditions based on the retrieval requirements; S500: Formulate a retrieval query condition based on the retrieval requirement conditions, perform data information matching in the document database according to the retrieval query condition, and use the successfully matched data document as the retrieval result; S600: Extract information from the retrieval result, analyze and organize the extracted information, and generate a retrieval information report.

[0019] Furthermore, through constructing and storing the data document in the document database, the present invention realizes the structured storage and management of data, improving the accessibility and maintainability of data; through keyword analysis and field importance evaluation, the present invention realizes the extraction and analysis of information in the data document, helping users to understand and utilize data more quickly; through constructing an index relationship and formulating retrieval requirement conditions, the present invention realizes the acceleration and optimization of retrieval operations, improving the retrieval efficiency and user experience; based on the key fields and retrieval requirement conditions, the present invention realizes accurate data retrieval and information extraction, generates a retrieval information report that meets the user's needs, and helps users quickly obtain the required information; through the information analysis and organization of the retrieval result, the present invention can discover the laws and trends hidden in the data, providing data-driven decision-making support for users; the entire process of the present invention can be partially or fully automated, with the help of technologies such as artificial intelligence, improving the processing efficiency and accuracy; by integrating various technical effects, the present invention provides users with a better retrieval and information processing service, enhancing the user experience and work efficiency.

[0020] In an embodiment of the present application, a retrieval method based on a document database is provided. It determines the fields and field information included in the data documents in the document database, and performs keyword analysis and extraction on the fields included in the data documents to determine the keywords in each field, including: determining the fields and field information included in the data documents in the document database, and using natural language processing technology to extract candidate keywords from the fields; determining the term frequency and inverse document frequency of each candidate keyword, and calculating the product of the term frequency and the inverse document frequency to obtain the term frequency-inverse document frequency value of each candidate keyword; determining the candidate keywords with a term frequency-inverse document frequency value higher than a preset threshold as the keywords in each field.

[0021] Specifically, in the document database, the data documents contain various fields, such as title, author, date, content, etc. Each field contains specific types of information and is an important part of the document data. Using natural language processing technology, candidate keywords can be extracted from each field. The term frequency (TF) represents the frequency of a certain word appearing in a document, and the inverse document frequency (IDF) represents the importance of the word in the entire document collection. The product of these two values is the TF-IDF value, which is used to measure the importance of a word in a document. Comparing the calculated term frequency-inverse document frequency value with the preset threshold, the candidate keywords higher than the threshold are determined as the keywords in each field, which helps to identify the most important information units in the document. Through natural language processing technology, this step can automatically extract keywords from document fields, reducing the workload of manual annotation, improving the processing efficiency and accuracy. Using the TF-IDF calculation method can better measure the importance of keywords in a document, helping to identify the most representative keywords in each field and improving the accuracy of information extraction. By setting a threshold, the most important keywords can be screened out, reducing the interference of irrelevant information and making the finally extracted keywords more practical and applicable. This process helps users better understand the document content, thereby improving the comprehensibility and information utilization value of the document, and providing strong support for subsequent retrieval, analysis, and decision-making. Through natural language processing technology and automated calculation, the entire process can be partially or fully automated, improving the data processing efficiency and accuracy and reducing the need for human intervention. Generally speaking, this step combines natural language processing, text mining, and information retrieval technologies, can effectively extract key information from the document database, improves the intelligence and automation level of data processing, and provides users with better data analysis and application services.

[0022] In an embodiment of the present application, a retrieval method based on a document database is provided. The method evaluates the importance of each field based on keywords and field information, and determines the key fields in the field according to the evaluation results, including: determining the field information corresponding to each field, where the field information includes the query frequency and the number of queries, and determining the term frequency-inverse document frequency value of the keywords corresponding to each field; determining the information importance value of the field information based on the query frequency and the number of queries, and respectively evaluating and obtaining the first importance evaluation value of the keywords and the second importance evaluation value of the field information; adding the first importance evaluation value and the second importance evaluation value to obtain the comprehensive importance evaluation value of the field, and determining the fields with the comprehensive importance evaluation value greater than the preset evaluation value as the key fields.

[0023] Specifically, in addition to the content, each field also has information such as the query frequency and the number of queries, which reflect the degree of attention and importance of the user to the field; based on information such as the query frequency and the number of queries, the information importance value of the field information is determined, which reflects the importance of the field in the user's query; evaluating and obtaining the term frequency-inverse document frequency value of the keywords and the information importance value of the field information provides a basis for subsequent importance evaluation; adding the first importance evaluation value of the keywords and the second importance evaluation value of the field information to obtain the comprehensive importance evaluation value of the field, which is used to determine the importance ranking of the fields; determining the fields with the comprehensive importance evaluation value greater than the preset evaluation value as the key fields, and these fields have importance and priority in data retrieval and analysis. This step can more comprehensively evaluate the importance of the fields by comprehensively considering the importance of the keywords and the information importance value of the field information, which helps to determine the most valuable and concerned content in the data; according to information such as the query frequency and the number of queries, the user's needs and concerns can be better understood, so as to optimize the data display and retrieval results and improve user satisfaction; by identifying the key fields, the data can be optimized and managed targeted, improving the usability and application effect of the data, and providing better data services for users; this process combines technologies such as natural language processing, data analysis, and user behavior analysis, enabling intelligent processing and decision-making of data, and improving the efficiency and effect of data management and utilization; by determining the key fields, customized data retrieval and analysis services can be provided for users, meeting the personalized needs of users, and enhancing the quality and value of data services. Generally speaking, this step combines data analysis, user behavior analysis, and information retrieval technologies, enabling a better understanding of user needs, optimizing data display, and improving data utilization effect, and providing more intelligent and personalized data services and decision support for users.

[0024] In an embodiment of the present application, a retrieval method based on a document database is provided. The information importance value of field information is determined based on the query frequency and the number of queries. The calculation formula for the information importance value of field information is as follows: L = α * M + β * N; where L is the information importance value of the field information, α is a preset first weight coefficient, M is the query frequency of the field, β is a preset second weight coefficient, and N is the number of queries of the field.

[0025] Specifically, the value of L represents the importance degree of a specific field in the entire dataset, which is obtained by comprehensively considering the query frequency and the number of queries; α and β are preset weight coefficients, and these two coefficients are used to adjust the influence degree of the query frequency and the number of queries on the information importance value of the field information, and can be set according to specific situations; M is the query frequency of the field, which represents the frequency of a specific field appearing in the query, that is, the ratio of the number of times being queried to the total number of queries, reflecting the general importance of the field in the query process; N is the number of queries of the field, which represents the specific number of times a specific field is queried, reflecting the degree of attention paid to this field in the actual query; α * M reflects the contribution of the field query frequency to the information importance value. The higher the query frequency, the larger this part of the value, and the greater the importance of the field. β * N reflects the contribution of the field query number to the information importance value. The more the number of queries, the larger this part of the value, indicating that this field is frequently queried and is thus more important. By calculating the information importance value in this step, the importance degree of the field in the dataset can be quantified, which helps to prioritize the processing of important information; by adjusting the weight coefficients of α and β, the influence degrees of the query frequency and the number of queries can be flexibly adjusted according to actual needs and scenarios; according to the information importance value of the field information, the data can be optimized and managed targeted, and the important information can be utilized and displayed more effectively; a quantification method is provided to help decision-makers identify and focus on important field information, providing more powerful support for decision-making; by identifying important fields, the presentation manner of the query results can be optimized, improving the query efficiency and user experience. Through such calculation of the information importance value, the importance of the field information can be evaluated more scientifically, providing strong support for data processing and decision-making, and improving the efficiency and quality of information management.

[0026] In an embodiment of the present application, a retrieval method based on a document database is provided. The retrieval requirements are determined, and the retrieval requirement conditions are determined based on the retrieval requirements, including: determining the retrieval requirements of the user, and splitting the retrieval requirements into multiple retrieval requirement conditions. The retrieval requirement conditions include retrieval content, retrieval scope, and retrieval time.

[0027] Specifically, by communicating with the user, the retrieval requirements of the user are determined, that is, which information the user hopes to retrieve from the database or document; decomposing the user's retrieval requirements into multiple retrieval requirement conditions, including aspects such as retrieval content, retrieval scope, and retrieval time, helps to more precisely locate the user's needs and improve the accuracy and efficiency of retrieval; the retrieval content is the keywords, phrases, or topics that the user hopes to retrieve and is used to describe the information content that the user hopes to find; the retrieval scope specifies the range of documents, databases, or data sets that the user hopes to retrieve, helping to narrow the retrieval scope and improve the retrieval efficiency; the retrieval time condition is used to limit the time range that the user hopes to retrieve, helping the user to find information within a specific time period. This step can more accurately understand the user's needs, improve the precision and accuracy of retrieval, by splitting the user's retrieval requirements into multiple conditions; splitting the retrieval requirements into multiple conditions helps to quickly locate the user's needs, narrow the retrieval scope, improve the retrieval efficiency, and save the user's time; setting the retrieval conditions can help filter out irrelevant information and only present the content that the user is interested in, improving the relevance and usability of the information; by understanding the user's retrieval requirements, more personalized and customized retrieval services can be provided to the user to meet the different needs and preferences of the user; splitting the user's retrieval requirements into multiple conditions helps to perform data analysis and mining, understand the user's behavior and demand trends, and provide a basis for data management and service optimization. By splitting the user's retrieval requirements into multiple conditions, the user's needs can be better understood and met, the efficiency and precision of information retrieval can be improved, and a better retrieval experience and service can be provided to the user.

[0028] In an embodiment of the present application, a retrieval method based on a document database is provided. The retrieval query conditions are formulated based on the retrieval requirement conditions, and data information matching is performed in the document database according to the retrieval query conditions. The data documents that match successfully are used as retrieval results, including: respectively writing corresponding retrieval content query statements, retrieval scope query statements, and retrieval time query statements based on the retrieval content, retrieval scope, and retrieval time, and forming retrieval query conditions according to the retrieval content query statements, retrieval scope query statements, and retrieval time query statements; inputting the retrieval query conditions into the document database, and the document database performs a retrieval operation according to the retrieval query conditions, and determines the data documents in the document database that match the retrieval content as retrieval results.

[0029] Specifically, a query statement is written based on the keywords or phrases input by the user to match relevant content in the database; a specific range or table to be retrieved in the document database is specified to ensure that the retrieval operation targets a specific data set; a query statement is written according to the time range specified by the user to limit the retrieval results within a specific time period; the retrieval content query statement, the retrieval range query statement, and the retrieval time query statement are combined into a complete retrieval query condition to ensure that the retrieval operation can accurately locate the user's needs; the composed retrieval query condition is input into the document database, and the database system performs a retrieval operation according to the condition to filter out the data documents that match the retrieval content; the document database performs a retrieval operation according to the query condition, determines the data documents that match the retrieval content in the database as the retrieval results, and can be sorted according to the relevance and presented to the user. By constructing an accurate retrieval query condition in this step, it can be ensured that the retrieval results highly match the user's needs, improving the accuracy and precision of the retrieval; using the retrieval query condition to perform a retrieval operation can quickly locate the data documents that meet the conditions, saving the user's time and improving the retrieval efficiency; according to the retrieval content and time range specified by the user, a personalized retrieval service is provided to ensure that the user obtains the data documents that meet the requirements; by performing a retrieval operation through the database, rapid retrieval and screening of a large amount of data are realized, which helps with data management and optimization; providing accurate retrieval results can improve the user experience, enabling the user to more easily find the required information and enhancing the user-friendliness of the system. By combining the retrieval content, range, and time together to construct a complete retrieval query condition, efficient and accurate information retrieval can be achieved, enhancing the user experience, meeting the user's personalized retrieval needs, and at the same time providing important support for data management and decision-making.

[0030] In an embodiment of the present application, a retrieval method based on a document database is provided. The method extracts information from the retrieval results, analyzes and organizes the extracted information, and generates a retrieval information report, including: extracting information from the retrieval results and performing preprocessing on the extracted information. The preprocessing includes word segmentation, stop word removal, and part-of-speech tagging; identifying key information from the preprocessed information and cleaning and deduplicating the extracted key information; classifying the cleaned key information and filling the classified key information into the report content categories to generate a retrieval information report.

[0031] Specifically, data documents that meet the search criteria are extracted from the document database as the information source to be processed; the text information is segmented according to words to form a word sequence for subsequent processing; common stop words that are irrelevant to information retrieval, such as "的", "和", etc., are removed to reduce noise interference; the part of speech of each word is marked to facilitate subsequent information analysis and processing; through text analysis technology, key information in the text is identified and extracted, such as important nouns, verb phrases, etc.; the extracted key information is cleaned, duplicate information and noise are removed, and the accuracy and completeness of the information are ensured; the cleaned key information is classified according to the pre-set classification standards, and the information is classified into corresponding categories; the classified key information is filled in according to the report content classification, and a structured retrieval information report is generated for user access and analysis. This step improves the accuracy of information extraction from the original text through preprocessing and key information identification, ensuring the quality of the extracted information; cleaning and deduplication can improve the quality and consistency of information, avoid repeated information interference and noise interference, and make the information more readable and credible; through classification and filling, the information is presented in a structured manner in the report, making it easier to manage and analyze, and improving the utilization value of the information; the automated information extraction, cleaning and classification process greatly improves the efficiency of processing large amounts of information, saving manpower and time costs; the generated retrieval information report can accurately reflect the key information in the retrieval results, provide users with clear and concise information display, and facilitate users to quickly obtain the required information. By extracting, preprocessing, identifying key information, cleaning, classifying and generating reports for the retrieval results, the efficiency and accuracy of information processing can be effectively improved, and a structured and clear information display can be provided to users, helping users to better understand and use the retrieval results.

[0032] like Figure 2As shown in the figure, in an embodiment of the present application, a retrieval system based on a document database is provided, including: an acquisition module, configured to acquire data to be retrieved, construct a data document according to the document form for the data to be retrieved, and store the constructed data document in the document database; an extraction module, configured to determine the fields and field information included in the data document in the document database, perform keyword analysis and extraction on the fields included in the data document, and determine the keywords in each field; an evaluation module, configured to evaluate the importance of each field based on the keywords and field information, and determine the key fields in the field according to the evaluation result; a determination module, configured to construct an index relationship for the key fields in the document database, determine the retrieval requirements, and determine the retrieval requirement conditions based on the retrieval requirements; a matching module, configured to formulate a retrieval query condition based on the retrieval requirement conditions, perform data information matching in the document database according to the retrieval query condition, and use the successfully matched data document as the retrieval result; a generation module, configured to extract information from the retrieval result, analyze and organize the extracted information, and generate a retrieval information report.

[0033] In summary, the embodiment of the present invention provides a retrieval method and system based on a document database, which includes: acquiring data to be retrieved, constructing a data document according to the document form and storing it in the document database; determining the fields and field information included in the data document, and extracting the keywords in each field; evaluating the importance of each field based on the keywords and field information, and determining the key fields in the field according to the evaluation result; constructing an index relationship for the key fields, and determining the retrieval requirement conditions in the retrieval requirements; formulating a retrieval query condition based on the retrieval requirement conditions, performing data information matching in the document database according to the retrieval query condition, and using the successfully matched data document as the retrieval result; extracting, analyzing and organizing the information from the retrieval result to generate a retrieval information report. The retrieval method based on the document database of the present invention combines technologies such as document storage, document retrieval, and index construction, and can improve the retrieval efficiency and accuracy.

[0034] Finally, it should be noted that: Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

[0035] The above is only an example of the present invention, but it should not limit the scope of the present invention. Any structural changes made in accordance with the present invention, as long as they do not deviate from the essence of the present invention, should be regarded as falling within the scope of protection of the present invention and being restricted. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process and related explanations of the above-described platform can refer to the corresponding process in the foregoing platform embodiments and will not be repeated here.

[0036] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, platform, article, or device / platform that comprises a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent in these processes, platforms, articles, or devices / platforms.

[0037] So far, the technical solution of the present invention has been described in combination with the further embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

[0038] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.

Claims

1. A retrieval method based on a document database, characterized in that: include: Obtain the data to be retrieved, construct a data document based on the document form of the data to be retrieved, and store the constructed data document in a document database; Determine the fields and field information contained in the data document in the document database, and perform keyword analysis and extraction on the fields contained in the data document to determine the keywords in each field; Evaluate the importance of each field based on keywords and field information, and determine the key fields in the field based on the evaluation results; Construct index relationships for key fields in the document database, determine retrieval requirements, and determine retrieval requirement conditions based on the retrieval requirements; Formulate search query conditions based on search requirements, match data information in the document database according to the search query conditions, and use the successfully matched data documents as search results; Extract information from the search results, analyze and organize the extracted information, and generate a search information report.

2. A retrieval method based on a document database according to claim 1, characterized in that: The determining of the fields and field information contained in the data document in the document database, and performing keyword analysis and extraction on the fields contained in the data document to determine the keywords in each field includes: Determine the fields and field information contained in the data documents in the document database, and use natural language processing technology to extract candidate keywords from the fields; Determine the word frequency and inverse document frequency of each candidate keyword, and calculate the product of the word frequency and the inverse document frequency to obtain the word frequency-inverse document frequency value of each candidate keyword; The candidate keywords whose word frequency-inverse document frequency value is higher than a preset threshold are determined as keywords in each field.

3. A retrieval method based on a document database according to claim 2, characterized in that: The importance of each field is evaluated based on the keywords and field information, and the key fields in the field are determined according to the evaluation results, including: Determine the field information corresponding to each field, wherein the field information includes the query frequency and the query times, and determine the word frequency-inverse document frequency value of the keyword corresponding to each field; Determine the information importance value of the field information based on the query frequency and the number of queries, and evaluate the word frequency-inverse document frequency value of the keyword and the information importance value of the field information respectively to obtain a first importance evaluation value of the keyword and a second importance evaluation value of the field information; The first importance evaluation value and the second importance evaluation value are added to obtain a comprehensive importance evaluation value of the field, and the field whose comprehensive importance evaluation value is greater than a preset evaluation value is determined as a key field.

4. A retrieval method based on a document database according to claim 3, characterized in that: The information importance value of the field information is determined based on the query frequency and the query times. The calculation formula of the information importance value of the field information is: L = α*M + β*N; Among them, L is the information importance value of the field information, α is the preset first weight coefficient, M is the query frequency of the field, β is the preset second weight coefficient, and N is the number of queries for the field.

5. A retrieval method based on a document database according to claim 3, characterized in that: The step of determining the search requirement and determining the search requirement conditions based on the search requirement includes: Determine the user's search needs and split the search needs into multiple search requirement conditions, which include search content, search scope and search time.

6. A retrieval method based on a document database according to claim 5, characterized in that: The step of formulating a search query condition based on the search requirement condition, matching data information in the document database according to the search query condition, and taking the successfully matched data document as the search result includes: Based on the search content, search scope and search time, respectively write corresponding search content query statements, search scope query statements and search time query statements, and form search query conditions according to the search content query statements, search scope query statements and search time query statements; The search query conditions are input into the document database, and the document database performs a search operation according to the search query conditions, and determines the data documents in the document database that match the search content as the search results.

7. A retrieval method based on a document database according to claim 6, characterized in that: The information extraction from the search results, and the analysis and organization of the extracted information to generate a search information report include: Extract information from the search results and preprocess the extracted information, including word segmentation, stop word removal and part-of-speech tagging; Identify key information from the preprocessed information, and clean and deduplicate the extracted key information; The cleaned key information is classified, and the classified key information is filled in according to the report content classification to generate a search information report.

8. A retrieval system based on a document database, characterized in that: include: The acquisition module is used to acquire the data to be retrieved, construct a data document according to the document form of the data to be retrieved, and store the constructed data document in the document database; The extraction module is used to determine the fields and field information contained in the data document in the document database, and perform keyword analysis and extraction on the fields contained in the data document to determine the keywords in each field; An evaluation module is used to evaluate the importance of each field based on keywords and field information, and determine the key fields in the field according to the evaluation results; A determination module is used to construct an index relationship for key fields in a document database, determine retrieval requirements, and determine retrieval requirement conditions based on the retrieval requirements; The matching module is used to formulate search query conditions based on the search requirement conditions, and match data information in the document database according to the search query conditions, and take the successfully matched data documents as the search results; The generation module is used to extract information from the search results, analyze and organize the extracted information, and generate a search information report.