Knowledge base document searching method and apparatus, electronic device, and storage medium

By combining hybrid tag expansion and hierarchical storage optimization with lightweight local model-generated summaries, this approach addresses the issues of disconnect between hot and cold tiering strategies and business needs, insufficient tag generation coverage, and poor AI semantic search performance in traditional document search systems. This results in more efficient storage resource utilization and more accurate search results.

CN120821830BActive Publication Date: 2025-12-23JINAN DALU ELECTROMECHANICAL CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511332929.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-12-23
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Traditional document search systems suffer from problems such as a disconnect between the Elasticsearch hot and cold document tiering strategy and business needs, insufficient tag generation coverage, and poor performance of AI semantic search.

Method used

The system uses hybrid tag expansion to generate composite tags, combines chunked processing and lightweight local models to generate summaries, calculates popularity scores in real time, performs tiered storage based on popularity scores and access frequency, dynamically adjusts SSD and HDD migration thresholds, and combines multi-dimensional fields to calculate comprehensive matching scores for sorting.

Benefits of technology

It achieves matching of hot and cold tiering strategy with business scenarios, improves the coverage of tag generation and the effect of artificial intelligence semantic search, and optimizes storage resource utilization and search efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821830B_ABST
    Figure CN120821830B_ABST
Patent Text Reader

Abstract

The application provides a knowledge base document search method and device, electronic equipment and storage medium, and relates to the technical field of data processing, in which, during hierarchical storage, the heat score of the uploaded document is used, the heat score of the uploaded document considers not only the access frequency of the uploaded document, but also the business scenario, that is, the hot and cold hierarchical strategy is closely related to the business scenario, in addition, the composite tag of the uploaded document generated includes core keywords, semantic synonyms and structured business tags, and the coverage is comprehensive, when the abstract is generated, a lightweight local model is used, the network is not needed, the required computing power and hardware conditions are not high, and the effect of artificial intelligence semantic search is good.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a knowledge base document search method and device, electronic equipment and storage medium. BACKGROUND

[0002] The current mainstream document search system usually adopts the following technologies:

[0003] Elasticsearch (ES) dominance: As a distributed search engine, ES is widely used in full-text retrieval scenarios, and its inverted index, tokenizer (such as IK Analyzer) and distributed shard mechanism support PB-level data with second-level response.

[0004] Auxiliary relational database: MySQL and other relational databases are usually used to store structured metadata (such as file attributes, permission information), and maintain data consistency with ES through timed synchronization or message queues (such as Kafka).

[0005] Horizontal expansion architecture: ES realizes linear expansion through shard and replica mechanisms.

[0006] The traditional document search has the following technical defects:

[0007] ES cold and hot layering strategy is out of touch with business needs: The traditional ES shard scheme only relies on access frequency and does not consider business scenario priority (such as technical documents need to be kept on SSD for a long time). For example, in a certain enterprise knowledge base, conference files are important and need to be retrieved in time, but due to less access frequency, they are mistakenly migrated to HDD, resulting in increased search delay;

[0008] Insufficient coverage of label generation: synonym library needs to be maintained manually; semantic fragmentation, TF-IDF cannot identify equivalent words (such as "AI" and "artificial intelligence"), resulting in missing "AI" files when searching for "artificial intelligence".

[0009] When performing artificial intelligence semantic search, if using a cloud API interface, the cost of calling is high and depends on the network; if using a local model, the computing power and hardware conditions are required to be high.

[0010] In summary, the traditional document search has the technical problems of ES cold and hot layering strategy being out of touch with business needs, insufficient coverage of label generation, and poor effect of artificial intelligence semantic search. SUMMARY

[0011] Therefore, the present application aims to provide a knowledge base document search method and device, electronic equipment and storage medium to alleviate the technical problems of the disconnection between ES hot and cold layering strategy and business needs, insufficient label coverage, and poor artificial intelligence semantic search effect in traditional document search.

[0012] In a first aspect, the embodiments of the present application provide a knowledge base document search method, comprising:

[0013] Performing hybrid label expansion on the uploaded document to generate a composite label of the uploaded document, wherein the composite label comprises a core keyword, a semantic synonym and a structured business label;

[0014] Generating a first-level summary of the uploaded document by using block processing and a lightweight local model, and automatically downgrading to extract a second-level summary of the uploaded document by using a preset rule after the first-level summary generation fails;

[0015] Real-time calculation of the hotness score of the uploaded document, wherein the hotness score is related to the access frequency and business scenario of the uploaded document;

[0016] Automatic setting of the SSD migration threshold and the HDD migration threshold according to the median of the access frequency of the full library document in a preset time period, and hierarchical storage of the uploaded document according to the hotness score, the SSD migration threshold and the HDD migration threshold;

[0017] Obtaining a search request, and calculating the comprehensive matching score between the search term in the search request and each uploaded document based on the composite label, the first-level summary or the second-level summary, the hotness score and the hierarchical storage of each uploaded document, and then displaying each uploaded document in descending order of the comprehensive matching score.

[0018] Further, performing hybrid label expansion on the uploaded document comprises:

[0019] Extracting the core keyword of the uploaded document by using TF-IDF;

[0020] Generating the semantic synonym of the core keyword by using the Word2Vec model;

[0021] Extracting the structured business label of the uploaded document based on document name regular matching;

[0022] Taking the core keyword, the semantic synonym and the structured business label as the composite label.

[0023] Further, after the first-level summary generation fails, automatically downgrading to extract the second-level summary of the uploaded document by using a preset rule, comprising:

[0024] The first 200 characters and the last 100 characters of the uploaded document are taken as the secondary abstract of the uploaded document.

[0025] Further, the heat score of the uploaded document is calculated in real time, including:

[0026] The heat score of the uploaded document is calculated in real time by using the heat score calculation formula HeatScore=(the number of visits of the uploaded document in the last 3 days × 0.7 + the number of historical visits of the uploaded document × 0.3)× 0.6 + the preset business weight × 0.4.

[0027] Further, the SSD migration threshold is 1.5 times of the median of the number of visits of all library documents in a preset time period, and the HDD migration threshold is 0.5 times of the median of the number of visits of all library documents in a preset time period, and the uploaded document is stored in layers according to the heat score of the uploaded document, the SSD migration threshold and the HDD migration threshold, including:

[0028] If the heat score of the uploaded document is greater than the SSD migration threshold, the uploaded document is stored in the SSD;

[0029] If the heat score of the uploaded document is less than the HDD migration threshold, the uploaded document is stored in the HDD.

[0030] Further, the comprehensive matching score of the search word in the search request and each uploaded document is calculated based on the composite label of each uploaded document, the first abstract or the second abstract of each uploaded document, the heat score, and the layered storage, including:

[0031] The basic score of the search word and each uploaded document is calculated according to the content of each uploaded document and the search word;

[0032] The label score of the search word and each uploaded document is calculated according to the composite label of each uploaded document and the search word;

[0033] The abstract score of the search word and each uploaded document is calculated according to the first abstract or the second abstract of each uploaded document and the search word;

[0034] The heat bonus of each uploaded document is calculated according to the heat score of each uploaded document and the layered storage.

[0035] According to the comprehensive matching score calculation formula, a comprehensive matching score of the search word and each of the uploaded documents is calculated, wherein ω1, ω2, ω3 and ω4 represent preset weights.

[0036] Further, a basic score of the search word and each of the uploaded documents is calculated according to the content of each of the uploaded documents and the search word, including:

[0037] According to the basic score calculation formula, the basic score of the search word and each of the uploaded documents is calculated as follows: basic score = min(100, (the relevance score of the content of the uploaded document and the search word));

[0038] According to the composite tag of each of the uploaded documents and the search word, a tag score of the search word and each of the uploaded documents is calculated, including:

[0039] According to the tag score calculation formula, the tag score of the search word and each of the uploaded documents is calculated as follows: tag score = min(100, the number of core keywords matched by the search word × 2 + the number of semantic synonyms and structured business tags matched by the search word × 1);

[0040] According to the first-level abstract or the second-level abstract of each of the uploaded documents and the search word, an abstract score of the search word and each of the uploaded documents is calculated, including:

[0041] If it is a first-level abstract, according to the first abstract score calculation formula, the abstract score of the search word and each of the uploaded documents is calculated as follows: abstract score = min(100, the number of keywords in the search word contained in the abstract × 20);

[0042] If it is a second-level abstract, according to the second abstract score calculation formula, the abstract score of the search word and each of the uploaded documents is calculated as follows: abstract score = min(100, the length of the abstract / 200*80);

[0043] According to the hot degree score of each of the uploaded documents and the hierarchical storage, a hot degree addition of each of the uploaded documents is calculated, including:

[0044] If the uploaded document is stored in the SSD, the hot degree addition of the uploaded document is a first preset value;

[0045] If the uploaded document is stored in the HDD, the hot degree addition of the uploaded document is a second preset value;

[0046] If the uploaded document is not stored in the SSD and the HDD, a heat bonus of each uploaded document is calculated according to a heat bonus calculation formula: heat bonus = 50 + 49 x (HeatScore-HDD migration threshold value) / (SSD migration threshold value-HDD migration threshold value), wherein HeatScore represents the heat score.

[0047] In a second aspect, the embodiments of the present application further provide a search device for a knowledge base document, comprising:

[0048] A composite tag generation unit is configured to perform hybrid tag expansion on the uploaded document to generate a composite tag of the uploaded document, wherein the composite tag comprises a core keyword, a semantic synonym and a structured business tag.

[0049] An abstract extraction unit is configured to generate a first-level abstract of the uploaded document by using block processing and a lightweight local model, and automatically downgrade to extract a second-level abstract of the uploaded document by using a preset rule after the first-level abstract generation fails.

[0050] A calculation unit is configured to calculate a heat score of the uploaded document in real time, wherein the heat score is related to an access frequency and a business scenario of the uploaded document.

[0051] A hierarchical storage unit is configured to automatically set an SSD migration threshold value and an HDD migration threshold value according to a median of an access frequency of all library documents in a preset time period calculated daily, and perform hierarchical storage on the uploaded document according to the heat score of the uploaded document, the SSD migration threshold value and the HDD migration threshold value.

[0052] A comprehensive matching score calculation unit is configured to obtain a search request, and calculate a comprehensive matching score between a search word in the search request and each uploaded document based on the composite tag, the first-level abstract or the second-level abstract, the heat score and the hierarchical storage of each uploaded document, and then display each uploaded document in descending order of the comprehensive matching score.

[0053] In a third aspect, the embodiments of the present application further provide an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method of any one of the first aspect when executing the computer program.

[0054] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, wherein the computer readable storage medium stores machine executable instructions, and the machine executable instructions cause the processor to execute the method of any one of the first aspect when the machine executable instructions are invoked and executed by the processor.

[0055] In the embodiment of the present application, a knowledge base document search method is provided, which comprises: performing hybrid label expansion on an uploaded document to generate a composite label of the uploaded document, wherein the composite label comprises a core keyword, a semantic synonym and a structured business label; generating a first-level summary of the uploaded document by using block processing and a lightweight local model, and automatically degrading to extract a second-level summary of the uploaded document by using a preset rule after the generation of the first-level summary fails; calculating a hotness score of the uploaded document in real time, wherein the hotness score is related to the access frequency and the business scenario of the uploaded document; automatically setting an SSD migration threshold and an HDD migration threshold according to the median of the access times of all library documents in a preset time period calculated daily, and performing hierarchical storage on the uploaded document according to the hotness score, the SSD migration threshold and the HDD migration threshold of the uploaded document; obtaining a search request, and calculating a comprehensive matching score between a search term in the search request and each uploaded document based on the composite label, the first-level summary or the second-level summary, the hotness score and the hierarchical storage of each uploaded document, and then displaying each uploaded document in descending order of the comprehensive matching score. As can be seen from the above description, in the knowledge base document search method of the present application, hierarchical storage is realized according to the hotness score of the uploaded document, the hotness score of the uploaded document not only considers the access frequency of the uploaded document, but also considers the business scenario, that is, the cold and hot hierarchical strategy is closely related to the business scenario, in addition, the generated composite label of the uploaded document includes core keywords, semantic synonyms and structured business labels, which has comprehensive coverage, when generating the summary, a lightweight local model is used, which does not need to rely on the network, and the required computing power and hardware conditions are also not high, the effect of artificial intelligence semantic search is good, and the technical problems of the traditional document search, such as the disconnection between the ES cold and hot hierarchical strategy and the business demand, the insufficient coverage of the label generation, and the poor effect of artificial intelligence semantic search, are solved. BRIEF DESCRIPTION OF DRAWINGS

[0056] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0057] Figure 1 A flowchart of a knowledge base document search method provided in the embodiment of the present application;

[0058] Figure 2 A schematic diagram of a knowledge base document search device provided in the embodiment of the present application;

[0059] Figure 3 A schematic diagram of an electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0060] The technical solutions of the present application will be described in detail below with reference to the embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0061] The search of the traditional document exists the technical problems that the ES cold and hot layering strategy is out of touch with the business demand, the coverage of the generated label is insufficient, and the effect of the artificial intelligence semantic search is poor.

[0062] Therefore, in the search method of the knowledge base document, the hierarchical storage is realized according to the hot degree score of the uploaded document. The hot degree score of the uploaded document not only considers the access frequency of the uploaded document, but also considers the business scenario, that is, the cold and hot layering strategy is closely related to the business scenario. In addition, the generated composite label of the uploaded document includes core keywords, semantic synonyms and structured business labels, and the coverage is comprehensive. When generating the abstract, a lightweight local model is used, which does not need to rely on the network, and the required computing power and hardware conditions are not high, and the effect of the artificial intelligence semantic search is good.

[0063] In order to facilitate the understanding of the present embodiment, first, a search method of a knowledge base document disclosed in the present embodiment is introduced in detail.

[0064] Embodiment one:

[0065] According to the embodiment of the present application, a search method of a knowledge base document is provided. It should be noted that the steps shown in the flowchart can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0066] Figure 1 The flowchart of the search method of the knowledge base document according to the embodiment of the present application is shown in FIG. 1, which includes the following steps: Figure 1

[0067] Step S102, performing mixed label expansion on the uploaded document to generate a composite label of the uploaded document, wherein the composite label includes core keywords, semantic synonyms and structured business labels;

[0068] ​Specifically, the content of the document is read and analyzed immediately after uploading, and the basic information of the uploaded document is stored in the MySQL database and the ES database. Then the content of the uploaded document is read, and the mixed label (i.e. mixed label expansion) is extracted. The process of mixed label expansion will be described in detail below.

[0069] In step S104, a first-level summary of the uploaded document is generated by using a block processing and a lightweight local model, and if the generation of the first-level summary fails, a second-level summary of the uploaded document is automatically extracted by using a preset rule;

[0070] Specifically, the first-level summary of the uploaded document is generated by using a divide-and-conquer strategy (long text blocking, i.e. block processing) and a lightweight local model (Ollama+qwen3:4b), which solves the problem of high network dependence by using network API interface and solves the problem of high computing power and hardware conditions by using lightweight model. If the extraction of the first-level summary fails (no extraction or extraction timeout), the rule engine extracts the second-level summary of the uploaded document by using a preset rule, realizes multi-level fault tolerance, and improves search coverage.

[0071] In step S106, the heat score of the uploaded document is calculated in real time, wherein the heat score is related to the access frequency and business scenario of the uploaded document.

[0072] In step S108, the median of the access frequency of all documents in a preset time period is calculated daily to automatically set the SSD migration threshold and the HDD migration threshold, and the uploaded document is stored in layers according to the heat score, the SSD migration threshold and the HDD migration threshold.

[0073] Specifically, the median of the access frequency of all documents in a preset time period is calculated daily at dawn, and the above-mentioned preset time period can be the last 30 days. Then, the SSD migration threshold and the HDD migration threshold are automatically set according to the median of the access frequency, and the uploaded document is stored in layers according to the heat score, the SSD migration threshold and the HDD migration threshold.

[0074] In step S110, the search request is obtained, and the comprehensive matching score between the search words in the search request and each uploaded document is calculated based on the composite label, the first-level summary or the second-level summary, the heat score and the layered storage of each uploaded document. Then, each uploaded document is sorted and displayed in descending order of the comprehensive matching score.

[0075] Specifically, ES realizes accurate sorting by multi-dimensional field collaborative calculation, and embeds function_score function to give 1.2 times score bonus to high heat files (HeatScore>800), which effectively improves search efficiency.

[0076] In the embodiment of the present application, a knowledge base document search method is provided, comprising: performing hybrid label expansion on the uploaded document to generate a composite label of the uploaded document, wherein the composite label comprises a core keyword, a semantic synonym and a structured business label; generating a first-level summary of the uploaded document by using block processing and a lightweight local model, and automatically degrading to extract a second-level summary of the uploaded document by using a preset rule after the first-level summary fails to be generated; calculating a hotness score of the uploaded document in real time, wherein the hotness score is related to the access frequency and the business scenario of the uploaded document; automatically setting an SSD migration threshold and an HDD migration threshold according to the median of the access times of all library documents in a preset time period calculated daily, and storing the uploaded document in layers according to the hotness score, the SSD migration threshold and the HDD migration threshold of the uploaded document; obtaining a search request, and calculating a comprehensive matching score of the search words in the search request and each uploaded document based on the composite label, the first-level summary or the second-level summary, the hotness score and the layered storage of each uploaded document, and then displaying each uploaded document in descending order of the comprehensive matching score. As can be seen from the above description, in the knowledge base document search method of the present application, the layered storage is realized according to the hotness score of the uploaded document, the hotness score of the uploaded document not only considers the access frequency of the uploaded document, but also considers the business scenario, that is, the cold and hot layering strategy is closely related to the business scenario, in addition, the generated composite label of the uploaded document includes core keywords, semantic synonyms and structured business labels, the coverage is comprehensive, when generating the summary, the lightweight local model is used, which does not need to rely on the network, and the required computing power and hardware conditions are also not high, the effect of artificial intelligence semantic search is good, and the technical problems of the traditional document search, such as the disconnection between the ES cold and hot layering strategy and the business demand, the insufficient coverage of label generation, and the poor effect of artificial intelligence semantic search, are solved.

[0077] The above describes the knowledge base document search method of the present application briefly, and the specific contents involved therein are described in detail below.

[0078] In an optional embodiment of the present application, the uploaded document is subjected to hybrid label expansion, which specifically comprises the following steps:

[0079] (1) The core keyword of the uploaded document is extracted by using TF-IDF;

[0080] Specifically, the core keyword (such as "artificial intelligence") of the uploaded document is extracted by using TF-IDF.

[0081] (2) The semantic synonym of the core keyword is generated by using Word2Vec model;

[0082] Specifically, the semantic synonym (such as "AI") of the core keyword (such as "artificial intelligence") is generated by using Word2Vec model.

[0083] (3) Extracting the structured business label of the uploaded document based on document name regular matching;

[0084] Specifically, the structured business label (such as “R&D Department_Technical Document”) of the uploaded document is extracted based on document name regular matching.

[0085] (4) Taking the core keyword, semantic synonym and structured business label as a composite label.

[0086] Specifically, the composite label can improve semantic coverage.

[0087] In an optional embodiment of the present application, after the first-level abstract generation fails, the second-level abstract of the uploaded document is automatically extracted by using a preset rule, and the specific steps include:

[0088] The first 200 characters and the last 100 characters of the uploaded document are taken as the second-level abstract of the uploaded document.

[0089] In an optional embodiment of the present application, the heat score of the uploaded document is calculated in real time, and the specific steps include:

[0090] The heat score of the uploaded document is calculated in real time by using the heat score calculation formula HeatScore= (the access frequency of the uploaded document in the last 3 days x 0.7 + the historical access frequency of the uploaded document x 0.3) x 0.6 + the preset business weight x 0.4.

[0091] Specifically, the business weight (1-10, for example: the weight of financial report = 9, the weight of temporary log = 2) is mapped in real time based on the MySQL document classification field (such as department and type). The above business weight is manually defined and saved in the database dictionary, and can be manually maintained and adjusted through the system management interface. The document information change is monitored through the Binlog log of MySQL, and is synchronized to ES.

[0092] The heat score of the uploaded document is calculated in real time by using the heat score calculation formula HeatScore= (the access frequency of the uploaded document in the last 3 days x 0.7 + the historical access frequency of the uploaded document x 0.3) x 0.6 + the preset business weight x 0.4. After each document change, the heat score is written to ES.

[0093] In an optional embodiment of the present application, the SSD migration threshold is 1.5 times the median of the access frequency of all library documents in a preset time period, and the HDD migration threshold is 0.5 times the median of the access frequency of all library documents in a preset time period. The uploaded document is stored in layers according to the heat score of the uploaded document, the SSD migration threshold and the HDD migration threshold, and the specific steps include:

[0094] (1) If the hotness score of the uploaded document is greater than the SSD migration threshold, the uploaded document is stored to the SSD;

[0095] (2) If the hotness score of the uploaded document is less than the HDD migration threshold, the uploaded document is stored to the HDD.

[0096] In an optional embodiment of the present application, based on the composite label, the primary abstract or the secondary abstract, the hotness score of each uploaded document, and the comprehensive matching score of the search words in the search request and each uploaded document, the specific steps include the following steps:

[0097] (1) The basic score of the search words and each uploaded document is calculated according to the content of each uploaded document and the search words;

[0098] Specifically, the basic score of the search words and each uploaded document is calculated according to the basic score calculation formula: basic score = min(100, (the relevance score of the content of the uploaded document and the search words)).

[0099] The relevance score of the content of the uploaded document and the search words is directly obtained by the BM25 algorithm of the ES (with built-in normalization).

[0100] (2) The label score of the search words and each uploaded document is calculated according to the composite label of each uploaded document and the search words;

[0101] Specifically, the label score of the search words and each uploaded document is calculated according to the label score calculation formula: label score = min(100, the number of core keywords matched by the search words x 2 + the number of semantic synonyms and structured business labels matched by the search words x 1).

[0102] The core keywords are labels directly matched such as titles / keywords.

[0103] (3) The abstract score of the search words and each uploaded document is calculated according to the primary abstract or the secondary abstract of each uploaded document and the search words;

[0104] Specifically, if it is a primary abstract, the abstract score of the search words and each uploaded document is calculated according to the first abstract score calculation formula: abstract score = min(100, the number of keywords in the search words contained in the abstract x 20); if it is a secondary abstract, the abstract score of the search words and each uploaded document is calculated according to the second abstract score calculation formula: abstract score = min(100, the length of the abstract / 200*80).

[0105] The meaning of 20 is that the business data analysis shows that a high-quality AI abstract contains an average of 4-6 keywords, so it is set to 20 points per word, that is, 5 words can reach the full score. Further adjustment can be made according to business needs.

[0106] The above 200 words are the longest original text summary standard, and the quality of AI summary is usually better than manual summary, so it is fair to limit the highest score of manual summary to 80. Further adjustments can be made according to business needs.

[0107] (4) Calculate the heat bonus of each uploaded document according to the heat score and hierarchical storage of each uploaded document;

[0108] Specifically, if the uploaded document is stored in SSD, the heat bonus of the uploaded document is a first preset value; if the uploaded document is stored in HDD, the heat bonus of the uploaded document is a second preset value; if the uploaded document is not stored in SSD and HDD, the heat bonus of each uploaded document is calculated according to the heat bonus calculation formula heat bonus = 50 + 49 × (HeatScore - HDD migration threshold) / (SSD migration threshold - HDD migration threshold), wherein HeatScore represents the heat score.

[0109] The first preset value is 100, and the second preset value is 0. The meaning of 100 is that the file reaching the SSD storage standard is considered as a high-heat file and obtains the highest bonus. The meaning of 0 is that the cold file has no bonus to avoid invalid file interference in sorting. The meaning of 50 and 49 is that the 50 points are the bottom value to ensure that the active file obtains the basic bonus, and the 49 points are the adjustment value to provide fine differentiation in the active interval (50-99 points). Business logic: make the file with high storage location promotion potential obtain higher bonus.

[0110] (5) Calculate the comprehensive matching score of the search word and each uploaded document according to the comprehensive matching score calculation formula comprehensive matching score = base score × ω1 + label score × ω2 + summary score × ω3 + heat bonus × ω4, wherein ω1, ω2, ω3, and ω4 represent preset weights.

[0111] The above ω1 can be 40%, ω2 can be 30%, ω3 can be 25%, and ω4 can be 5%.

[0112] An example of Elasticsearch query is as follows:

[0113] "script_score": {

[0114] "script": {

[0115] "source": """

[0116] / / Get base data

[0117] double base = Math.min(100, _score);

[0118] double tags = Math.min(100, doc['core_tags'].value*2 + doc['related_tags'].value);

[0119] / / Calculate summary score

[0120] double summary;

[0121] if (doc['abstract_type'].value == 'AI') {

[0122] summary = Math.min(100, doc['keywords_count'].value * 20);

[0123] } else {

[0124] summary = Math.min(100, doc['abstract_length'].value / 200.0 * 80);

[0125] }

[0126] / / Get storage tier threshold parameters (read from ES global parameters index)

[0127] double ssdThreshold = params.ssd_threshold;

[0128] double hddThreshold = params.hdd_threshold;

[0129] double heatScore = doc['heat'].value;

[0130] / / Calculate heat bonus

[0131] double heat;

[0132] if (heatScore>= ssdThreshold) {

[0133] heat = 100; / / SSD file highest bonus

[0134] } else if (heatScore>= hddThreshold) {

[0135] / / Linear mapping (50-99 score range)

[0136] double range = ssdThreshold - hddThreshold;

[0137] double position = (heatScore - hddThreshold) / range;

[0138] heat = 50 + (49 * position); / / 50 base + 49 delta

[0139] } else {

[0140] heat = 0; / / No bonus for HDD files

[0141] }

[0142] / / Combined matching score (preserve original weights)

[0143] return base*0.4 + tags*0.3 + summary*0.25 + heat*0.05;

[0144] """,

[0145] "params": {

[0146] / / Pass in daily updated threshold parameters via API

[0147] "ssd_threshold": 2000, / / Example value, actual value is obtained from global parameters

[0148] "hdd_threshold": 200 / / Example value, actual value is obtained from global parameters

[0149] }

[0150] }

[0151] }

[0152] The method of the present application has the following effects:

[0153] 1. Dynamic optimization of storage resources: by combining dynamic threshold algorithm (i.e. SSD migration threshold and HDD migration threshold) with business weight, the system can intelligently identify high-frequency access documents and preferentially allocate high-performance storage resources, while automatically migrating low-active documents to low-cost storage. This mechanism effectively balances storage efficiency and access performance, reducing resource waste;

[0154] 2. Enhanced semantic understanding: The composite tag strategy integrates keyword extraction, synonym expansion, and structured tag parsing, significantly improving the semantic coverage of professional terms, polysemous words, and business scenarios, making search results more comprehensive and accurate;

[0155] 3. Improved robustness of summary generation: A multi-level degradation strategy is adopted, which prioritizes the generation of structured summaries through a lightweight local model. When the model encounters anomalies, it automatically switches to the key information extraction rules (i.e., preset rules) to ensure that effective summary content can still be output even in extreme cases.

[0156] 4. Dynamic Search Ranking Adaptation: Based on a multi-dimensional scoring model (i.e., comprehensive matching score) that considers document popularity, content relevance, and business rules, search results can dynamically respond to real-time access changes, prioritizing the display of high-value documents and improving user search efficiency.

[0157] Advantages compared to existing technologies:

[0158] 1. Dynamic storage tiering

[0159] Compared to traditional fixed threshold solutions, this invention dynamically adjusts the migration threshold using a rolling median, avoiding storage misjudgments caused by sudden traffic spikes or long-term inefficient documents. For example, traditional solutions require manual threshold adjustment to cope with access fluctuations, while this invention reduces manual intervention through adaptive algorithmic optimization.

[0160] 2. Mixed Tag Expansion

[0161] Compared to single-semantic model solutions, existing technologies rely on a single model (such as BERT) to generate tags, which lacks sufficient support for professional terminology and business rules. This invention combines a rule engine with a pre-trained model, ensuring both general semantic understanding and accurate extraction of structured business tags. For example, the filename "2023Q4_R&D Report" can be parsed to extract tags such as "department, time, and type," which is difficult to achieve with traditional solutions relying solely on word segmentation.

[0162] 3. Multi-level summary generation

[0163] Compared to single-path summarization schemes, existing technologies (such as pure model generation) are prone to failure in processing long texts or complex formats and lack degradation strategies. This invention ensures that valid summaries are output from technical documents to scanned files through chunked processing and rule-based degradation. For example, when model processing times out, the system automatically extracts the core content of the first paragraph of the file instead of directly returning an empty result.

[0164] 4. Dynamic rating search

[0165] Compared to static weighted ranking schemes, traditional schemes use fixed formulas to calculate ranking scores, making it difficult to reflect changes in file popularity in real time. This invention dynamically couples access popularity with search scores, enabling documents experiencing sudden surges in popularity to quickly improve their ranking priority, while static schemes require waiting for scheduled task updates, resulting in significant delays.

[0166] Analysis of the mechanism by which the effect occurs:

[0167] The core logic of storage optimization is that the rolling median threshold identifies genuine high-frequency access needs by statistically analyzing the recent access distribution, avoiding interference from individual extreme values ​​in migration decisions. For example, when a document experiences a surge in short-term access, a dynamic threshold can quickly trigger migration, while a fixed threshold requires manual calibration.

[0168] The effectiveness of semantic expansion is guaranteed: TF-IDF filters core terms from the content, Word2Vec expands synonyms to solve the problem of expression diversity, and regular expression rules parse filenames to supplement business context. The three work together to cover the three layers of information: content, semantics, and business. For example, when searching for "AI", the system simultaneously matches the "artificial intelligence" tag and technical documents containing "machine learning".

[0169] The rationale behind the summary fault tolerance design: When model generation fails, the degradation strategy prioritizes extracting the first paragraph of the text or metadata (such as filenames and tags) rather than simply truncating random fragments. For example, the core objective of a technical document is usually located at the beginning, and truncating the first paragraph can retain more than 80% of the key information;

[0170] Real-time responsiveness of dynamic scoring: HeatScore's real-time updates are dynamically linked to search scores, ensuring that highly popular documents immediately impact their ranking weight after changes in visitor volume. For example, a technical report experiencing a surge in visitor volume due to a sudden event can see its search ranking rise to the top within minutes.

[0171] Example 2:

[0172] This invention also provides a knowledge base document search device, which is mainly used to execute the knowledge base document search method provided in Embodiment 1 of this invention. The following is a detailed description of the knowledge base document search device provided in this invention.

[0173] Figure 2 This is a schematic diagram of a knowledge base document search device according to an embodiment of the present invention, such as... Figure 2 As shown, the device mainly includes: a composite tag generation unit 10, a summary extraction unit 20, a calculation unit 30, a hierarchical storage unit 40, and a comprehensive matching score calculation unit 50, wherein:

[0174] The composite tag generation unit is configured to perform hybrid tag expansion on the uploaded document to generate a composite tag of the uploaded document, wherein the composite tag comprises a core keyword, a semantic synonym and a structured business tag.

[0175] The abstract extraction unit is configured to generate a first-level abstract of the uploaded document by using a block processing and a lightweight local model, and automatically downgrade to extract a second-level abstract of the uploaded document by using a preset rule when the first-level abstract fails to be generated.

[0176] The computing unit is configured to calculate a hotness score of the uploaded document in real time, wherein the hotness score is related to an access frequency and a business scenario of the uploaded document.

[0177] The hierarchical storage unit is configured to automatically set an SSD migration threshold and an HDD migration threshold according to a median of the access times of all the documents in a preset time period calculated daily, and to store the uploaded document according to the hotness score, the SSD migration threshold and the HDD migration threshold.

[0178] The comprehensive matching score calculation unit is configured to obtain a search request, and to calculate a comprehensive matching score between a search word in the search request and each uploaded document based on the composite tag, the first-level abstract or the second-level abstract, the hotness score and the hierarchical storage of each uploaded document, and to display each uploaded document in descending order of the comprehensive matching score.

[0179] In the embodiment of the present application, a knowledge base document searching device is provided, comprising: performing hybrid label expansion on an uploaded document to generate a composite label of the uploaded document, wherein the composite label comprises: a core keyword, a semantic synonym and a structured business label; generating a first-level summary of the uploaded document by using block processing and a lightweight local model, and automatically degrading to extract a second-level summary of the uploaded document by using a preset rule after the generation of the first-level summary fails; calculating a heat score of the uploaded document in real time, wherein the heat score is related to the access frequency and the business scenario of the uploaded document; automatically setting an SSD migration threshold and an HDD migration threshold according to the median of the access times of all library documents in a preset time period calculated daily, and storing the uploaded document in layers according to the heat score, the SSD migration threshold and the HDD migration threshold of the uploaded document; obtaining a search request, and calculating the comprehensive matching scores of the search words in the search request and each uploaded document based on the composite label, the first-level summary or the second-level summary, the heat score and the layered storage of each uploaded document, and then displaying each uploaded document in descending order of the comprehensive matching scores. As can be seen from the above description, in the knowledge base document searching device of the present application, the layered storage is realized according to the heat score of the uploaded document, the heat score of the uploaded document not only considers the access frequency of the uploaded document, but also considers the business scenario, that is, the cold and hot layering strategy is closely related to the business scenario, in addition, the generated composite label of the uploaded document includes core keywords, semantic synonyms and structured business labels, the coverage is comprehensive, when generating the summary, the lightweight local model is used, which does not need to rely on the network, and the required computing power and hardware conditions are also not high, the effect of artificial intelligence semantic search is good, and the technical problems of the traditional document searching, such as the disconnection between the ES cold and hot layering strategy and the business demand, the insufficient coverage of the label generation, and the poor effect of artificial intelligence semantic search, are solved.

[0180] Optionally, the composite label generation unit is further configured to: extract the core keyword of the uploaded document by using TF-IDF; generate the semantic synonym of the core keyword by using a Word2Vec model; extract the structured business label of the uploaded document based on document name regular matching; and take the core keyword, the semantic synonym and the structured business label as the composite label.

[0181] Optionally, the summary extraction unit is further configured to: take the first 200 characters and the last 100 characters of the uploaded document as the second-level summary of the uploaded document.

[0182] Optionally, the calculation unit is further configured to: calculate the heat score of the uploaded document in real time by using a heat score calculation formula HeatScore= (the access times of the uploaded document in the last 3 days x 0.7 + the historical access times of the uploaded document x 0.3) x 0.6 + a preset business weight x 0.4.

[0183] Optionally, the SSD migration threshold is 1.5 times of the median of the access times of all library documents in a preset time period, the HDD migration threshold is 0.5 times of the median of the access times of all library documents in a preset time period, and the hierarchical storage unit is further configured to: if the heat score of the uploaded document is greater than the SSD migration threshold, store the uploaded document in the SSD; and if the heat score of the uploaded document is less than the HDD migration threshold, store the uploaded document in the HDD.

[0184] Optionally, the comprehensive matching score calculation unit is further configured to: calculate a basic score of the search word and each uploaded document according to the content of each uploaded document and the search word; calculate a label score of the search word and each uploaded document according to the composite label of each uploaded document and the search word; calculate a summary score of the search word and each uploaded document according to the first-level summary or the second-level summary of each uploaded document and the search word; calculate a heat bonus of each uploaded document according to the heat score of each uploaded document and the hierarchical storage; and calculate a comprehensive matching score of the search word and each uploaded document according to a comprehensive matching score calculation formula: comprehensive matching score = basic score × ω1 + label score × ω2 + summary score × ω3 + heat bonus × ω4, wherein ω1, ω2, ω3 and ω4 represent preset weights.

[0185] Optionally, the comprehensive matching score calculation unit is further configured to: calculate a basic score of the search word and each uploaded document according to a basic score calculation formula: basic score = min(100, (the relevance score of the content of the uploaded document and the search word)); calculate a label score of the search word and each uploaded document according to a label score calculation formula: label score = min(100, the number of core keywords matched by the search word × 2 + the number of semantic synonyms and structured business labels matched by the search word × 1); calculate a summary score of the search word and each uploaded document according to a first summary score calculation formula: summary score = min(100, the number of keywords contained in the summary × 20) if the summary is a first-level summary; calculate a summary score of the search word and each uploaded document according to a second summary score calculation formula: summary score = min(100, the length of the summary / 200 * 80) if the summary is a second-level summary; take a first preset value as the heat bonus of the uploaded document if the uploaded document is stored in the SSD; take a second preset value as the heat bonus of the uploaded document if the uploaded document is stored in the HDD; and calculate a heat bonus of each uploaded document according to a heat bonus calculation formula: heat bonus = 50 + 49 × (HeatScore - HDD migration threshold) / (SSD migration threshold - HDD migration threshold) if the uploaded document is not stored in the SSD and the HDD, wherein HeatScore represents the heat score.

[0186] The device provided in the embodiments of the present application has the same implementation principle and technical effects as the foregoing method embodiments, and for brevity, the part not mentioned in the device embodiment part can be referred to the corresponding content in the foregoing method embodiments.

[0187] As Figure 3 shown, an electronic device 600 provided by an embodiment of the present application includes a processor 601, a memory 602, and a bus. The memory 602 stores machine readable instructions executable by the processor 601. When the electronic device is running, the processor 601 communicates with the memory 602 through the bus. The processor 601 executes the machine readable instructions to perform the steps of the knowledge base document searching method described above.

[0188] Specifically, the memory 602 and the processor 601 can be general memory and processor, which are not specifically limited herein. When the processor 601 runs the computer program stored in the memory 602, the processor 601 can execute the knowledge base document searching method described above.

[0189] The processor 601 can be an integrated circuit chip having a processing capability of signals. In the implementation process, each step of the method described above can be completed by the integrated logic circuit of hardware in the processor 601 or the instructions in the form of software. The processor 601 described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory, an electrically erasable programmable memory, a register, or other mature storage medium in the art. The storage medium is located in the memory 602, and the processor 601 reads the information in the memory 602 and combines the hardware to complete the steps of the method described above.

[0190] Corresponding to the search method of the knowledge base document, the embodiment of the application further provides a computer readable storage medium, the computer readable storage medium stores computer executable instructions, when the computer executable instructions are called and run by a processor, the computer executable instructions cause the processor to run the steps of the search method of the knowledge base document.

[0191] The search device for the knowledge base document provided by the embodiment of the application can be specific hardware on a device or software or firmware installed on the device, etc. The device provided by the embodiment of the application has the same implementation principle and generated technical effects as the foregoing method embodiments, and for the sake of brief description, the part not mentioned in the device embodiment part can be referred to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that, for the sake of convenience and brevity of description, the specific working process of the system, device and unit described above can be referred to the corresponding process in the foregoing method embodiments, which will not be described here.

[0192] In the embodiments provided in the application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some communication interfaces, devices or units, and can be electrical, mechanical or other forms.

[0193] For another example, the flowcharts and block diagrams in the drawings show the possible implementation architecture, function and operation of the devices, methods and computer program products according to the embodiments of the application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different order from that shown in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system for executing the specified functions or actions, or can be implemented by a combination of special hardware and computer instructions.

[0194] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e. may be located in one place, or may be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0195] In addition, the functional units in the embodiments provided in the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.

[0196] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the part of the present application that essentially contributes to the prior art or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the knowledge base document search method described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0197] It should be noted that: similar reference numbers and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings, in addition, the terms "first", "second", "third" and the like are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0198] Finally, it should be noted that: the above-described embodiments are only specific embodiments of the present application, used to illustrate the technical solutions of the present application, and not to limit them, the protection scope of the present application is not limited thereto, although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any skilled person in the art can modify or easily think of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed by the present application, or make equivalent replacement to part of the technical features; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application. All should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for searching knowledge base documents, characterized in that, include: Perform hybrid tag expansion on the uploaded document to generate composite tags for the uploaded document, wherein the composite tags include: core keywords, semantic synonyms and structured business tags; The upload document is generated using a chunked processing and lightweight local model. If the generation of the first-level summary fails, it is automatically downgraded to a second-level summary of the upload document extracted using preset rules. The popularity score of the uploaded document is calculated in real time, wherein the popularity score is related to the access frequency and business scenario of the uploaded document; The SSD migration threshold and HDD migration threshold are automatically set based on the median number of accesses to all documents in the database within a preset time period calculated daily. The uploaded documents are then stored in tiers based on their popularity score, the SSD migration threshold, and the HDD migration threshold. A search request is obtained, and a comprehensive matching score between the search terms in the search request and each uploaded document is calculated based on the composite tags of each uploaded document, the first-level summary or the second-level summary, the popularity score, and the hierarchical storage. Then, the uploaded documents are sorted and displayed in descending order of the comprehensive matching score. This includes performing mixed tag expansion on uploaded documents, including: The core keywords of the uploaded document were extracted using TF-IDF. The Word2Vec model was used to generate semantic synonyms for the core keywords. Extract structured business tags from the uploaded documents based on document name regular expression matching; The core keywords, semantic synonyms, and structured business tags are used as the composite tags; The automatic downgrade to extracting a second-level summary of the uploaded document using preset rules after the first-level summary generation fails includes: The first 200 characters and the last 100 characters of the uploaded document are used as the secondary summary of the uploaded document; The real-time calculation of the popularity score of the uploaded document includes: The heatscore is calculated in real time using the formula HeatScore = (Number of visits to the uploaded document in the last 3 days × 0.7 + Number of visits to the uploaded document in the past × 0.3) × 0.6 + Preset business weight × 0.

4. Wherein, the SSD migration threshold is 1.5 times the median number of accesses to all documents within a preset time period, and the HDD migration threshold is 0.5 times the median number of accesses to all documents within a preset time period. The uploaded documents are stored in tiers based on their popularity score, the SSD migration threshold, and the HDD migration threshold, including: If the popularity score of the uploaded document is greater than the SSD migration threshold, then the uploaded document will be stored on the SSD. If the popularity score of the uploaded document is less than the HDD migration threshold, the uploaded document will be stored in the HDD. The comprehensive matching score between the search terms in the search request and each of the uploaded documents is calculated based on the composite tags of each uploaded document, the first-level summary or the second-level summary, the popularity score, and the hierarchical storage, including: Calculate the basic score between the search term and each uploaded document based on the content of each uploaded document and the search term; Calculate the tag score between the search term and each uploaded document based on the composite tag of each uploaded document and the search term; Calculate the summary score of the search term and each uploaded document based on the first-level summary or the second-level summary of each uploaded document and the search term; The popularity bonus for each uploaded document is calculated based on its popularity score and the hierarchical storage. The comprehensive matching score is calculated using the formula: Comprehensive Matching Score = Base Score × ω1 + Tag Score × ω2 + Summary Score × ω3 + Popularity Bonus × ω4. This formula calculates the comprehensive matching score between the search term and each uploaded document, where ω1, ω2, ω3, and ω4 represent preset weights.

2. The method according to claim 1, characterized in that, Based on the content of each uploaded document and the search terms, a basic score is calculated between the search terms and each uploaded document, including: The basic score is calculated based on the formula: Basic Score = min(100, (Relevance score between the content of the uploaded document and the search term)). The basic scores of the search term and each uploaded document are calculated. Calculate the tag score between the search term and each uploaded document based on the composite tag of each uploaded document and the search term, including: The tag score is calculated based on the following formula: Tag Score = min(100, number of core keywords matched by the search term × 2 + number of semantic synonyms and structured business tags matched by the search term × 1). The tag score of the search term and each of the uploaded documents is calculated using this formula. Calculate the summary score of the search term and each uploaded document based on the first-level summary or the second-level summary of each uploaded document and the search term, including: If it is a first-level summary, the summary score is calculated based on the first summary score using the formula: summary score = min(100, number of keywords in the search terms contained in the summary × 20). The summary scores of the search terms and each of the uploaded documents are then calculated. If it is a second-level summary, the summary score is calculated based on the second summary score using the formula: summary score = min(100, summary length / 200*80) to calculate the summary score of the search term and each of the uploaded documents; Based on the popularity score of each uploaded document and the hierarchical storage, a popularity bonus is calculated for each uploaded document, including: If the uploaded document is stored on an SSD, then the popularity of the uploaded document is increased by a first preset value; If the uploaded document is stored in an HDD, then the popularity of the uploaded document is increased by a second preset value; If the uploaded document is not stored on SSD and HDD, the heat score is calculated according to the heat score calculation formula: Heat Score = 50 + 49 × (HeatScore - HDD migration threshold) / (SSD migration threshold - HDD migration threshold), where HeatScore represents the heat score.

3. A search device for knowledge base documents, characterized in that, include: A composite tag generation unit is used to perform hybrid tag expansion on the uploaded document and generate composite tags for the uploaded document, wherein the composite tags include: core keywords, semantic synonyms and structured business tags; The abstract extraction unit is used to generate a first-level abstract of the uploaded document using block processing and a lightweight local model, and automatically downgrades to extracting a second-level abstract of the uploaded document using preset rules if the generation of the first-level abstract fails. A calculation unit is used to calculate the popularity score of the uploaded document in real time, wherein the popularity score is related to the access frequency and business scenario of the uploaded document; The tiered storage unit is used to automatically set SSD migration thresholds and HDD migration thresholds based on the median number of accesses to all documents in the database within a preset time period calculated daily, and to perform tiered storage of the uploaded documents based on the popularity score of the uploaded documents, the SSD migration thresholds, and the HDD migration thresholds. The comprehensive matching score calculation unit is used to obtain search requests and calculate the comprehensive matching score between the search terms in the search request and each of the uploaded documents based on the composite tags of each uploaded document, the first-level summary or the second-level summary, the popularity score, and the hierarchical storage. Then, the uploaded documents are sorted and displayed in descending order of the comprehensive matching score. The composite tag generation unit is further configured to: extract core keywords from the uploaded document using TF-IDF; generate semantic synonyms for the core keywords using the Word2Vec model; extract structured business tags from the uploaded document based on document name regular expression matching; and use the core keywords, the semantic synonyms, and the structured business tags as the composite tag. The abstract extraction unit is further configured to: use the first 200 characters and the last 100 characters of the uploaded document as a secondary abstract of the uploaded document; The calculation unit is also used to: calculate the heat score of the uploaded document in real time using the formula HeatScore = (number of visits to the uploaded document in the last 3 days × 0.7 + number of visits to the uploaded document in the past × 0.3) × 0.6 + preset business weight × 0.

4. Wherein, the SSD migration threshold is 1.5 times the median number of accesses to all documents in the database within a preset time period, and the HDD migration threshold is 0.5 times the median number of accesses to all documents in the database within a preset time period. The tiered storage unit is further configured to: if the popularity score of the uploaded document is greater than the SSD migration threshold, then store the uploaded document to the SSD; if the popularity score of the uploaded document is less than the HDD migration threshold, then store the uploaded document to the HDD. The comprehensive matching score calculation unit is further configured to: calculate the basic score between the search term and each uploaded document based on the content of each uploaded document and the search term; calculate the tag score between the search term and each uploaded document based on the composite tags of each uploaded document and the search term; calculate the summary score between the search term and each uploaded document based on the first-level summary or the second-level summary of each uploaded document and the search term; calculate the popularity bonus of each uploaded document based on the popularity score of each uploaded document and the hierarchical storage; and calculate the comprehensive matching score between the search term and each uploaded document based on the formula: comprehensive matching score = basic score × ω1 + tag score × ω2 + summary score × ω3 + popularity bonus × ω4, where ω1, ω2, ω3, and ω4 represent preset weights.

4. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 2.

5. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the method of any one of claims 1 to 2.

Citation Information

Patent Citations

  • Quick retrieval method for documents in knowledge base, application server and computer readable storage medium

    CN108038096A