A file query method, device and electronic equipment
Patent Information
- Application Number
- CN202610636462.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-09
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]随着公司业务扩张,制度文件数量倍增,此时,由于制度文件增加,导致在传统知识库中手动查找制度文件的运维人员的检索效率显著降低,同时,运维人员在传统知识库中基于关键词检索制度文件时,传统知识库无法根据运维人员输入的关键词精准匹配到专业术语,从而容易检索到无关结果
[0016]第五方面,本申请实施例提供了一种计算机程序产品,包括计算机程序,所述计算机程序存储在计算机可读存储介质中;当电子设备的处理器从所述计算机可读存储介质读取所述计算机程序时,所述处理器执行所述计算机程序,使得所述电子设备执行上述文件查询方法。
Smart Images

Figure CN122838530A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a document retrieval method, apparatus, and electronic device. Background Technology
[0002] In a digital operations and maintenance (O&M) system, the work of data center operators and O&M personnel relies on a large number of standardized policy documents. When the number of policy documents is limited, they are typically fragmented and stored in a traditional knowledge base for management. When needed, O&M personnel can find the required documents by manually searching or by performing simple keyword searches.
[0003] As the company's business expands, the number of policy documents increases exponentially. This leads to a significant decrease in the retrieval efficiency for operations and maintenance personnel manually searching for policy documents in the traditional knowledge base. Furthermore, when operations and maintenance personnel search for policy documents in the traditional knowledge base based on keywords, the traditional knowledge base cannot accurately match the professional terms entered by the keywords, thus easily resulting in irrelevant results.
[0004] In conclusion, when there are many policy documents, improving the efficiency and accuracy of policy document retrieval is an urgent issue to be addressed. Summary of the Invention
[0005] This application provides a document retrieval method, apparatus, and electronic device to improve the retrieval efficiency and accuracy of institutional documents when there are many documents.
[0006] In a first aspect, embodiments of this application provide a document retrieval method, including: Retrieve the query statement entered by the query object; Based on the query statement, after performing a hybrid retrieval on the intelligent knowledge base, multiple query results and the similarity of each query result to each retrieval method are obtained; the hybrid retrieval includes multiple retrieval methods; the intelligent knowledge base includes: knowledge blocks obtained by dynamically slicing the file to be stored according to the internal hierarchical structure of each file to be stored; For each query result, based on the preset weight of each retrieval method, multiple similarities of the query results are weighted and fused, and an initial query result is determined from the multiple query results according to the weighted fusion result; The query statement is categorized by intent, and the initial query result is supplemented with contextual information based on the categorization results to obtain the final query result.
[0007] Secondly, embodiments of this application provide a document retrieval device, including: The retrieval unit is used to retrieve the query statement input by the query object; The retrieval unit is used to perform a hybrid retrieval on the intelligent knowledge base based on the query statement, and obtain multiple query results and the similarity of each query result to each retrieval method; the hybrid retrieval includes multiple retrieval methods; the intelligent knowledge base includes: knowledge blocks obtained by dynamically slicing the file to be stored according to the internal hierarchical structure of each file to be stored; An initial query unit is used to perform weighted fusion of multiple similarities of each query result based on a preset weight for each retrieval method, and determine an initial query result from the multiple query results based on the weighted fusion result. The final query unit is used to classify the intent of the query statement and supplement the initial query result with contextual information based on the classification result to obtain the final query result.
[0008] In some embodiments, the apparatus further includes a storage unit, which stores each of the files to be stored in the intelligent knowledge base in the following manner: The file to be stored is subjected to noise filtering and format normalization processing; Identify the chapter structure of the processed file to be stored, and obtain at least one chapter title in the processed file to be stored; For each chapter title, if the chapter title belongs to a preset independent chapter title, then the chapter content corresponding to the chapter title is taken as a knowledge block; if the chapter title does not belong to a preset independent chapter title, then the chapter content is dynamically sliced based on the length of the chapter content corresponding to the chapter title to obtain multiple knowledge blocks. The knowledge blocks are stored in the intelligent knowledge base.
[0009] In some embodiments, the storage unit is specifically used for: Obtain the length of the chapter content corresponding to the chapter title; when the length does not reach a preset length threshold, the chapter content corresponding to the chapter title is treated as a knowledge block; when the length reaches the preset length threshold, based on the preset length threshold and according to the natural boundaries of the chapter content, the chapter content corresponding to the chapter title is split into multiple knowledge blocks containing complete semantics; the length of the chapter content corresponding to each knowledge block does not exceed the preset length threshold.
[0010] In some embodiments, the storage unit is further configured to: For each knowledge block, a corresponding chapter identifier is generated for the knowledge block based on the chapter title to which the knowledge block belongs; a character count identifier is generated for the knowledge block based on the number of characters contained in the knowledge block; and a unique identifier is generated for the knowledge block. Each knowledge block, along with its corresponding chapter identifier, character count identifier, and unique identifier, is stored in the intelligent knowledge base.
[0011] In some embodiments, the hybrid retrieval includes: semantic retrieval and keyword retrieval; the retrieval unit is specifically used for: Based on the query statement, semantic retrieval is performed on the intelligent knowledge base to obtain a preset number of first query results and the vector similarity of the first query results to the semantic retrieval; and Based on the query statement, keyword retrieval is performed on the intelligent knowledge base to obtain a preset number of second query results and the word similarity of the second query results to the keyword retrieval.
[0012] In some embodiments, the retrieval unit performs semantic retrieval on the intelligent knowledge base in the following manner: The query statement is converted into a query vector, and the knowledge blocks in the intelligent knowledge base are converted into knowledge block vectors; Determine the vector similarity between the query vector and each knowledge block vector; Based on the similarity of each vector, a preset number of first query results are determined from the intelligent knowledge base; In some embodiments, the retrieval unit performs keyword retrieval on the intelligent knowledge base in the following manner: Obtain at least one query keyword from the query statement; For each knowledge block, perform the following steps: For each query keyword, determine the frequency of occurrence of the query keyword in the knowledge block and the inverse document frequency of the query keyword in each knowledge block; The knowledge block is subjected to length normalization processing to obtain the length of the knowledge block after length normalization. Based on the occurrence frequency, the inverse document frequency, and the knowledge block length, the sub-similarity between the knowledge block and the query keyword is determined; The sum of the similarities of each sub-sub ... Based on the similarity of each word, a preset number of second query results are determined from the intelligent knowledge base.
[0013] In some embodiments, the final query unit is specifically used for: The intent of the query statement is categorized to obtain the intent category to which the query statement belongs; The initial query results are supplemented with contextual information based on the intent category; Based on a preset prompt word template, a target prompt word template is generated in conjunction with the query statement. The initial query result, after supplementing with context information, is then filled into the target prompt word template to obtain the final query result. The final query result contains the filename of the file to be stored corresponding to the query statement, the chapter title in the file to be stored, and the chapter content of the file to be stored. The prompt word template includes: file requirements, chapter title requirements, and chapter content requirements.
[0014] Thirdly, embodiments of this application provide an electronic device, including: Memory, used to store program instructions; The processor is used to call the program instructions stored in the memory and execute the above-mentioned file query method according to the obtained program instructions.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the aforementioned file query method.
[0016] Fifthly, embodiments of this application provide a computer program product, including a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the above-described file query method.
[0017] This application provides a file query method, apparatus, and electronic device. This application overcomes the limitations of traditional fixed-length slices by dynamically slicing the file to be stored according to its internal hierarchical structure before storage. By combining the internal hierarchical structure of the file to be stored, it avoids semantic breaks and solves the fragmentation problem of the file to be stored in existing solutions.
[0018] Compared to the simple keyword-based retrieval in existing solutions, the embodiments of this application can effectively improve the accuracy of the initial multiple query results by using a hybrid retrieval method that includes multiple retrieval methods. Furthermore, by weighting and fusing each query result, a solid foundation is laid for further determining the final query result from these multiple query results, thereby significantly improving the accuracy and efficiency of the overall retrieval system.
[0019] This application's embodiments solve the problem that existing solutions cannot adapt query results to complex operation and maintenance scenarios by classifying query statements by intent and supplementing the initial query results with contextual information based on the classification results. Attached Figure Description
[0020] Figure 1This is a schematic diagram of an application scenario provided by an embodiment of this application; Figure 2 A flowchart illustrating the implementation of a document retrieval method provided in this application embodiment; Figure 3 A flowchart illustrating the implementation of another document query method provided in this application embodiment; Figure 4 This is a schematic diagram of the structure of an electronic device for a file query method according to an embodiment of this application; Figure 5 This is a schematic diagram of the hardware structure of an electronic device using an embodiment of this application; Figure 6 This is a schematic diagram of the hardware structure of a computing device according to an embodiment of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] The following describes some of the concepts involved in the embodiments of this application.
[0023] 1. Intelligent Knowledge Base: This is a knowledge management system that integrates artificial intelligence technologies such as natural language processing, semantic understanding, and large language models. It not only efficiently stores and organizes structured or unstructured institutional documents, operational procedures, and other professional knowledge, but also accurately understands the complex query intent of the querying object, automatically identifies the business scenario behind the question, and precisely locates relevant information within massive amounts of knowledge. It supports logical correlation analysis across documents and regulations, such as dynamically linking maintenance processes with risk management clauses, and generates clear, accurate, and traceable answers based on authoritative original texts, thereby helping the querying object quickly obtain actionable decision-making basis.
[0024] 2. Keyword retrieval: This refers to a method of information retrieval where the query object inputs one or more representative words or phrases (i.e., keywords) to search for content containing these keywords in a document collection, database, or knowledge base, and returns relevant results. This method mainly sorts the results based on the frequency, position, or degree of matching of keywords in the text.
[0025] 3. Semantic retrieval: This is a retrieval technology based on natural language understanding. It not only focuses on the literal form of words in the query, but also emphasizes understanding the underlying semantic intent and contextual meaning. It searches for semantically related content in a knowledge base or document collection, rather than simply results containing the same keywords. By utilizing word vectors, embedding models, or large language models, semantic retrieval can identify synonyms, near-synonyms, conceptual associations, and domain-specific terminology differences, thereby significantly improving the accuracy, relevance, and recall of retrieval.
[0026] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.
[0027] like Figure 1 The diagram shown is an application scenario illustration of an embodiment of this application. The application scenario diagram includes a terminal device 110 and a server 120.
[0028] The query subject enters a query statement on the client of terminal device 110, and server 120 obtains the query statement entered by the query subject from terminal device 110. Based on the query statement, server 120 performs a mixed search on the intelligent knowledge base to obtain multiple query results and the similarity of each query result to each search method. For each query result, server 120 performs weighted fusion of multiple similarities of the query results based on the preset weight of each search method, and determines the initial query result from the multiple query results according to the weighted fusion result. Server 120 classifies the query statement by intent, and supplements the initial query result with contextual information according to the classification result to obtain the final query result, and returns the final query result to terminal device 110 so that the query subject can view it on the client of terminal device 110.
[0029] In one alternative implementation, the terminal device 110 and the server 120 can communicate via a communication network.
[0030] In one alternative implementation, the communication network is a wired network or a wireless network.
[0031] The document query method provided by the exemplary embodiments of this application will be described below with reference to the accompanying drawings and the application scenarios described above. It should be noted that the application scenarios described above are only shown to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way.
[0032] The document query method provided in this application embodiment can be applied to intelligent knowledge base systems that include data layers, processing layers, and service layers.
[0033] Data layer: In a digital operation and maintenance system, the data layer is used to store the source files of unprocessed operation and maintenance policy documents, covering categories such as equipment management, fault handling, and risk control. Processing layer: Used to implement file storage and output semantically complete and searchable knowledge blocks; Service layer: Provides file query services, responds to query requests from query objects (such as operations and maintenance personnel), and returns the final query results.
[0034] The document query method provided in this application will be further described below with reference to specific embodiments and the intelligent knowledge base system with the above structure.
[0035] Figure 2 A flowchart of a document query method provided in an embodiment of this application is shown, such as... Figure 2 As shown, the method may include the following steps S201~S204: S201: Obtain the query statement input by the query object.
[0036] S202: Based on the query statement, after performing a mixed search on the intelligent knowledge base, multiple query results are obtained, as well as the similarity of each query result to each search method.
[0037] The hybrid retrieval includes multiple retrieval methods, such as semantic retrieval and keyword retrieval; the intelligent knowledge base includes knowledge blocks obtained by dynamically slicing the files to be stored according to their internal hierarchical structure.
[0038] First, the dynamic slicing and storage method of the files to be stored in the embodiments of this application will be introduced.
[0039] This application embodiment takes into account that existing solutions use fixed-length slices when storing files to be stored, which destroys the semantic integrity of the chapters and clauses within the file. For example, if a clause on risk level classification is split, the query object cannot obtain the complete level determination criteria. Therefore, this application embodiment preserves the semantic integrity of the internal hierarchical structure of chapters and clauses as much as possible when slicing files to be stored.
[0040] In one optional implementation, each file to be stored is stored in the intelligent knowledge base in the following manner: The system performs noise filtering and format normalization on the file to be stored; obtains at least one chapter title from the processed file; for each chapter title, if the chapter title belongs to a preset independent chapter title, the chapter content corresponding to the chapter title is taken as a knowledge block; if the chapter title does not belong to a preset independent chapter title, the chapter content is dynamically sliced based on the length of the chapter content corresponding to the chapter title to obtain multiple knowledge blocks; and stores the knowledge blocks in the intelligent knowledge base.
[0041] Specifically, in an intelligent knowledge base system, firstly, the processing layer is compatible with formats such as Portable Document Format (PDF), JavaScript Object Notation (JSON), and plain text, securely parsing the files to be stored. For example, if the input file to be stored is already a dictionary structure, it prioritizes extraction from common fields (such as data1, content, and text); if the input file to be stored is a JSON string, it attempts to parse and extract the main text to ensure that a processable string is ultimately obtained. For files of different formats, the actual text content is extracted through multiple layers of attempts.
[0042] Secondly, the processing layer extracts metadata (such as file number, version, effective date, etc.) from the files to be stored using automated scripts, providing a basis for subsequent data traceability. For example, key metadata, including file number (e.g., OP-2023-001), file version (e.g., V2.1), effective date (e.g., June 2023), and file title (matching Chinese titles containing keywords such as "system," "regulation," and "method"), can be automatically identified within the first 1000 characters of the file to be stored.
[0043] Furthermore, to improve the efficiency of subsequent noise filtering and format normalization, compared to processing the entire file to be stored directly, this embodiment can paginate the file before processing, facilitating simultaneous processing of multiple pages. For example, this embodiment can split the file into multiple pages based on page numbers (or fixed lengths), while retaining the pagination information. The processing layer can divide the file into logical pages based on common page number markers (such as "page 5" or "5 / 20"); if a valid page number cannot be identified, it can be roughly segmented into segments of approximately 2000 characters each, and a virtual page number can be assigned to each segment for easier subsequent page-by-page management.
[0044] Finally, the processing layer filters out noisy content, fixes cross-line text and formatting issues, standardizes text formatting (such as unifying list symbols), and outputs clean paginated text and complete metadata. For example, it cleans each page of text line by line, removing typical interfering elements: fixed watermarks (such as internal controlled documents), lines consisting solely of numbers / symbols (such as page numbers in footers, separators, etc.), and residual HyperText Markup Language (HTML) tags. It also standardizes whitespace characters and bullet points (such as...). The code is uniformly converted to -), improving text cleanliness. Meta-information is completed and standardized, and the total number of pages and characters is calculated. If the file number or title of the file to be stored is not extracted, a unique identifier (ID) is generated based on the content hash and assigned a default name (e.g., "Unnamed System_OP_abc12345") to ensure that each file to be stored has complete and usable metadata.
[0045] After performing the above processing on the file to be stored, in order to facilitate processing, the file to be stored is first paginated. Therefore, before performing subsequent slicing, the processed file to be stored needs to be merged into a complete and continuous full text.
[0046] When splitting the processed file to be stored, since this embodiment splits it according to its internal storage structure, the processing layer first needs to identify the chapter structure of the file to be stored. The chapter structure of the processed file to be stored can be identified by "chapter title regular expression". For example, the chapter title and its level number in the processed file to be stored can be automatically identified by regular expression (such as matching the format "3.2 Device Reset Process"). The start and end positions (character index) of each chapter in the processed file to be stored are recorded to form a structured chapter list.
[0047] To address the fragmentation problem of knowledge blocks in existing solutions, this application adopts a multi-rule slicing strategy: independent chapters are retained as a whole, short chapters are completely split, and long chapters are split according to semantic boundaries.
[0048] Specifically, in this application embodiment, independent chapter titles (such as table of contents, change records, etc.) can be preset. If a chapter title belongs to a preset independent chapter title, regardless of the length of the chapter content, it will not be sliced. Instead, each independent chapter will be stored as a knowledge block in the intelligent knowledge base to ensure that its information integrity is not compromised.
[0049] If the chapter title is not a preset independent chapter title, the processed file to be stored will be further sliced in the following way: Get the length of the chapter content corresponding to the chapter title; if the length does not reach the preset length threshold, treat the chapter content corresponding to the chapter title as a knowledge block; if the length reaches the preset length threshold, based on the preset length threshold and according to the natural boundaries of the chapter content, split the chapter content corresponding to the chapter title into multiple knowledge blocks containing complete semantics; the length of the chapter content corresponding to the knowledge block does not exceed the preset length threshold.
[0050] Specifically, in this embodiment, the processing layer, based on the internal hierarchical structure of the processed file to be stored, adopts a multi-rule slicing strategy (such as retaining independent chapters as a whole, completely splitting short chapters, and splitting long chapters according to semantic boundaries) to generate knowledge blocks that do not exceed a preset length threshold.
[0051] The preset length threshold can be an empirical value, such as 900 characters or 600 characters. For chapter content exceeding the preset length threshold, the length of the chapter content of its split knowledge blocks will not exceed the preset length threshold. Taking a preset length threshold of 900 characters as an example, the length of the split knowledge blocks can be 500-800 characters. To avoid semantic breaks, in this embodiment, there are overlapping characters between adjacent knowledge blocks. For example, 80 characters can be set to overlap between adjacent knowledge blocks. Alternatively, there can be overlapping sentences between adjacent knowledge blocks. For example, two sentences can be set to overlap between adjacent knowledge blocks to avoid sentence breaks caused by character-level overlap. This overlap setting is more suitable for institutional documents with large differences in sentence length (such as short sentences like "Notes" and long sentences like "Operating Instructions"), which can maintain sentence-level semantic integrity.
[0052] Natural boundaries refer to the naturally existing locations in a text that can clearly divide semantic units or information blocks, such as paragraph boundaries, periods, or semantic breakpoints.
[0053] If the chapter title is not a preset independent chapter title, but the length of the chapter content exceeds the preset length threshold, it can be finely segmented according to natural paragraph boundaries, periods, or semantic breakpoints, cutting the long chapter into multiple knowledge blocks. However, each knowledge block also contains complete semantics, avoiding mechanical truncation that could lead to the breakage of key clauses.
[0054] Furthermore, this application embodiment can also employ dynamic segmentation based on sentence embedding clustering. By generating the embedding vector of each sentence through a pre-trained model (Sentence-BERT) that generates semantic sentence vectors, semantically similar sentences are grouped into a knowledge block using K-means clustering (the K value is dynamically adjusted according to the length of the file to be stored). The clustering distance threshold is set to 0.6 (adapting to the semantic density of the sentences). This slicing method is suitable for non-standardized files to be stored with unclear chapter structures.
[0055] After performing the above-mentioned slicing, the knowledge blocks in this application embodiment can also be marked in the following way: For each knowledge block, a corresponding chapter identifier is generated based on the chapter title to which the knowledge block belongs; a character count identifier is generated based on the number of characters contained in the knowledge block; a unique identifier is generated for the knowledge block; and each knowledge block, along with its corresponding chapter identifier, character count identifier, and unique identifier, is stored in the intelligent knowledge base.
[0056] Specifically, the processing layer also needs to generate metadata for each knowledge block, such as a unique ID, the chapter to which it belongs, and the number of characters, to ensure that the knowledge blocks obtained after slicing not only meet the length requirements but also retain the hierarchical semantics of the file to be stored.
[0057] This application embodiment stores the file to be stored in the intelligent knowledge base through the above-described slicing method. The multi-rule dynamic slicing strategy breaks through the limitations of traditional fixed-length slicing. Combining the internal hierarchical structure of the chapter-clause of the file to be stored, a multi-rule slicing mechanism is designed to retain independent chapters as a whole, completely split short chapters, and split the semantic boundaries of long chapters. At the same time, the semantic break of core content is avoided by designing the character overlap of adjacent knowledge blocks, thus solving the text fragmentation problem in existing solutions.
[0058] After storing each file slice to be stored into the intelligent knowledge base in the above manner, in an optional implementation, this application embodiment performs the following query for each query statement.
[0059] Based on the query statement, semantic retrieval is performed on the intelligent knowledge base to obtain a preset number of first query results and the vector similarity of the semantic retrieval corresponding to the first query results; and based on the query statement, keyword retrieval is performed on the intelligent knowledge base to obtain a preset number of second query results and the word similarity of the keyword retrieval corresponding to the second query results.
[0060] Specifically, after the query object inputs a query statement (such as "manufacturer maintenance personnel qualification requirements"), the processing layer simultaneously performs semantic retrieval (vector similarity matching) and keyword retrieval (BM25 keyword matching). The semantic retrieval and keyword retrieval are introduced below respectively.
[0061] In one optional implementation, the embodiments of this application perform semantic retrieval on the intelligent knowledge base in the following manner: The query statement is transformed into a query vector, and the knowledge blocks in the intelligent knowledge base are transformed into knowledge block vectors; the vector similarity between the query vector and each knowledge block vector is determined; based on the vector similarity, a preset number of first query results are determined from the intelligent knowledge base.
[0062] Specifically, during the retrieval process in the intelligent knowledge base, the processing layer first converts the natural language query statement input by the query object (such as "manufacturer maintenance personnel qualification requirements") into a high-dimensional numerical vector, called the query vector, through an embedding model. At the same time, each knowledge block in the intelligent knowledge base has also been converted into a corresponding knowledge block vector and stored in the vector index library.
[0063] Next, the service layer calculates the vector similarity (such as cosine similarity) between the query vector and all knowledge block vectors in the knowledge base, thereby measuring the semantic matching degree between each knowledge block and the query statement. The higher the vector similarity, the more likely the knowledge block is to contain the information needed by the query object.
[0064] Finally, the service layer sorts the vector similarity scores from high to low and selects a preset number (e.g., Top 20) of relevant knowledge blocks as the first query result.
[0065] In one optional implementation, keyword retrieval of the intelligent knowledge base is performed in the following manner: Retrieve at least one query keyword from the query statement; For each knowledge block, perform the following steps: For each query keyword, determine the frequency of occurrence of the query keyword in the knowledge block and the inverse document frequency of the query keyword in each knowledge block; perform length normalization on the knowledge block to obtain the length of the normalized knowledge block; based on the frequency of occurrence, the inverse document frequency, and the length of the knowledge block, determine the sub-similarity between the knowledge block and the query keyword; and use the sum of all sub-similarity as the word similarity between the knowledge block and the query statement. Based on the similarity of each word, a preset number of second query results are determined from the intelligent knowledge base.
[0066] Specifically, the service layer extracts one or more representative query keywords from the query statement of the query object (for example, extracting core words such as "device", "reset", and "check" from "What items need to be checked before resetting the device?").
[0067] Subsequently, for each knowledge block in the knowledge base, the processing layer will process these keywords one by one: For each query keyword, its Term Frequency (TF) in the knowledge block is calculated, which is how many times the query keyword appears in the current knowledge block. At the same time, the Inverse Document Frequency (IDF) of the query keyword in all knowledge blocks of the entire intelligent knowledge base is calculated to measure the distinctiveness of the query keyword. The less common the query keyword, the higher its IDF value and the greater its contribution to relevance. Since longer knowledge blocks are naturally more likely to contain more query keywords, in order to avoid length bias, the service layer will also perform length normalization processing on the knowledge block. That is, based on its actual number of characters or words, its weight influence is compressed according to a standard formula to obtain a normalized effective length factor.
[0068] Based on this, using TF, IDF, and length normalization factor, and following classic relevance models such as BM25, the sub-similarity between the query keyword and the current knowledge block is calculated. By summing the sub-similarity values corresponding to all query keywords, the comprehensive word similarity between the knowledge block and the entire query statement can be obtained.
[0069] Finally, the service layer sorts all knowledge blocks from high to low based on word similarity and selects a preset number (e.g., Top 20) of knowledge blocks as the second query result.
[0070] In addition, in the embodiments of this application, besides using a hybrid retrieval method that includes semantic retrieval and keyword retrieval to determine the first query result and the second query result, dense retrieval and TF-IDF can also be used. The dense retrieval model generates a dense vector of the file to be stored, and TF-IDF calculates the keyword weights. This scheme is more suitable for long text semantics.
[0071] S203: For each query result, based on the preset weight of each retrieval method, multiple similarities of the query results are weighted and fused, and the initial query result is determined from multiple query results according to the weighted fusion result.
[0072] Obviously, there are some duplicate query results in the first query results and the second query results determined in this application embodiment. These duplicate query results have both vector similarity and word similarity. However, for non-duplicate query results, there is only one kind of similarity. Therefore, in this application embodiment, in order to perform subsequent weighted fusion, the word similarity of the first query result is set to a fixed value, such as 0, and the vector similarity of the second query result is set to a fixed value, such as 0.
[0073] In one optional implementation, the embodiments of this application employ the following method to weightedly fuse multiple similarities for each query result: For each query result, the vector similarity and word similarity are weighted and fused based on the first preset weight of semantic retrieval and the second preset weight of full-text retrieval to obtain the fused similarity; the initial query result is determined from multiple query results based on the fused similarity corresponding to each query result.
[0074] Specifically, for each query result, the service layer comprehensively considers the matching degree of both semantic retrieval and full-text retrieval. Semantic retrieval assesses deep semantic connections by calculating vector similarity between the query and the result, while full-text retrieval measures the surface vocabulary matching degree through word similarity. To balance the influence of these two retrieval methods, the service layer assigns a first preset weight to semantic retrieval and a second preset weight to full-text retrieval. Based on these weights, the service layer performs a weighted fusion of vector similarity and word similarity to generate a unified fused similarity for each query result. This fused similarity comprehensively reflects the overall relevance of the result at both the semantic and lexical levels. Finally, based on the fused similarity corresponding to all query results, the service layer determines the most relevant initial query result through methods such as sorting or threshold filtering.
[0075] The first and second preset weights are empirical values, and their sum is 1. For example, the first preset weight can be set to 0.8, the second preset weight to 0.2, and the fusion similarity to 0.8. Vector similarity +0.2 Word similarity. Then, the query results are sorted according to the fusion similarity, and the query result with the highest fusion similarity is used as the initial query result.
[0076] Furthermore, if the embodiments of this application employ dense retrieval and TF-IDF, generating a dense vector of the file to be stored through a dense retrieval model and calculating keyword weights through TF-IDF, then the fused similarity can be used as the dense retrieval score. 0.7+ TF-IDF score 0.3.
[0077] In addition, in addition to determining the initial query result through a weighted fusion method, the embodiments of this application can also input the first query result and the second query result into a small language reordering model, calculate the query-candidate set cross-attention score of each query result, sort each query result according to the cross-attention score, and select the initial query result from them.
[0078] In the above embodiments, this application embodiment addresses the characteristic of dense terminology in operation and maintenance scenarios by constructing a dynamic hybrid retrieval method that combines semantic retrieval and keyword retrieval. Semantic retrieval employs a fine-tuned embedding model adapted to operation and maintenance terminology, while keyword retrieval optimizes the BM25 algorithm to strengthen the weight of high-frequency operation and maintenance keywords such as "fault" and "emergency," achieving better accuracy than traditional single retrieval.
[0079] S204: Classify the intent of the query statement and supplement the initial query results with contextual information based on the classification results to obtain the final query results.
[0080] This application embodiment addresses the problem of incomplete answers and fragmented information caused by insufficient understanding of the deep intent of the query object in existing solutions. By combining intent recognition with contextual information supplementation, the service layer can automatically aggregate and organize relevant information for different types of queries, thereby generating a complete, logically coherent, and directly usable answer unit, realizing the leap from returning a list of relevant documents to providing structured answers.
[0081] In one optional implementation, embodiments of this application supplement the initial query results with contextual information in the following manner: The query statement is categorized by intent to obtain the intent category to which the query statement belongs; contextual information is added to the initial query results based on the intent category; a target prompt word template is generated based on a preset prompt word template and the query statement; the initial query results with added contextual information are filled into the target prompt word template to obtain the final query results, so that the final query results include the file name of the file to be stored corresponding to the query statement, the chapter title in the file to be stored, and the chapter content of the file to be stored; the prompt word template includes: file requirements, chapter title requirements, and chapter content requirements.
[0082] Specifically, this application embodiment can categorize query statements into several types, such as process consultation, fault handling, and terminology explanation, using an intent recognition model. For multi-turn interaction scenarios (such as "What records need to be made after fault handling is completed?"), it automatically supplements the contextual information of the initial query results. Furthermore, it calls a large language model and combines it with preset target prompt word templates for different intent categories under the operation and maintenance scenario to force the generated final query results to cite specific institutional clauses, generating final query results that fit the original text of the institutional clauses and indicating the source of the citation.
[0083] For example, the target prompt template could be "Based on Article {section} of the Maintenance Management System, answer: {query}", then the final query result could be "Article 4.2 of the Maintenance Management System requires that the qualifications of the manufacturer's maintenance personnel comply with the Supplier Management System".
[0084] In addition, embodiments of this application may also employ a lightweight text convolutional neural network (TextCNN) model to replace the intent recognition model for intent classification. By extracting local semantic features of the query statement through TextCNN, the intent category is output. This solution is suitable for real-time question answering scenarios with high response speed requirements.
[0085] This application embodiment can also use a few-sample prompt words to replace the three-segment tracing template. Two examples with clearly labeled sources of the relevant regulations are added to the prompt words (e.g., "Example 1: Query 'Sick Leave Procedure,' the answer must cite Article 3.1 of the 'Attendance Management System'"), guiding the large language model to automatically learn the prompt word template for answering and tracing the regulations. This solution is suitable for scenarios where regulations are frequently updated, eliminating the need to manually modify the version information in the prompt words. The large language model can adapt to new regulations content through the examples.
[0086] In the above embodiments, this application adopts an intent-driven intelligent interaction mechanism, which can automatically supplement contextual information in multi-round interactions, solving the problem that existing solutions cannot adapt to complex operation and maintenance scenarios. By setting a three-stage traceability template, the generated content is forced to cite specific institutional clauses and is marked with document version and chapter title, ensuring the traceability and credibility of the generated content and avoiding answers that deviate from the original text of the institutional rules in existing solutions.
[0087] like Figure 3 The diagram shown is an implementation flowchart of another document query method provided in this application embodiment. Figure 3 The middle data layer stores the source files of unprocessed operation and maintenance policy documents. In the automated workflow of the processing layer, the document cleaning module is responsible for noise filtering and format standardization of the files to be stored; the dynamic slicing module is responsible for slicing the processed files; the vector generation module is responsible for converting query statements into query vectors and knowledge blocks in the intelligent knowledge base into knowledge block vectors; and the service layer generates the final query results through a hybrid retrieval module, an intent classification module, and a context information supplementation module.
[0088] Based on the same inventive concept, this application also provides a document retrieval device, such as... Figure 4 As shown, the document query device 4000 includes: The acquisition unit 4001 is used to acquire the query statement input by the query object; The retrieval unit 4002 is used to perform a hybrid retrieval on the intelligent knowledge base based on the query statement, and obtain multiple query results and the similarity of each query result to each retrieval method; the hybrid retrieval includes multiple retrieval methods; the intelligent knowledge base includes: knowledge blocks obtained by dynamically slicing the files to be stored according to the internal hierarchical structure of each file to be stored; The initial query unit 4003 is used to perform weighted fusion of multiple similarities of each query result based on the preset weight of each retrieval method for each query result, and determine the initial query result from multiple query results based on the weighted fusion result. The final query unit 4004 is used to classify the intent of the query statement and supplement the initial query result with contextual information based on the classification result to obtain the final query result.
[0089] In some embodiments, the device further includes a storage unit 4005, which stores each file to be stored in the intelligent knowledge base in the following manner: Perform noise filtering and format normalization on the files to be stored; Identify the chapter structure of the processed file to be stored, and obtain at least one chapter title in the processed file to be stored; For each chapter title, if the chapter title belongs to a preset independent chapter title, the chapter content corresponding to the chapter title is treated as a knowledge block; if the chapter title does not belong to a preset independent chapter title, the chapter content is dynamically sliced based on the length of the chapter content corresponding to the chapter title to obtain multiple knowledge blocks. Store knowledge blocks in an intelligent knowledge base.
[0090] In some embodiments, the storage unit 4005 is specifically used for: Get the length of the chapter content corresponding to the chapter title; if the length does not reach the preset length threshold, treat the chapter content corresponding to the chapter title as a knowledge block; if the length reaches the preset length threshold, based on the preset length threshold and according to the natural boundaries of the chapter content, split the chapter content corresponding to the chapter title into multiple knowledge blocks containing complete semantics; the length of the chapter content corresponding to the knowledge block does not exceed the preset length threshold.
[0091] In some embodiments, the storage unit is further configured to: For each knowledge block, generate a corresponding chapter identifier for the knowledge block based on the chapter title to which the knowledge block belongs; generate a character count identifier for the knowledge block based on the number of characters contained in the knowledge block; and generate a unique identifier for the knowledge block. Each knowledge block, along with its corresponding chapter identifier, character count identifier, and unique identifier, is stored in the intelligent knowledge base.
[0092] In some embodiments, hybrid retrieval includes semantic retrieval and keyword retrieval; the retrieval unit 4002 is specifically used for: Based on the query statement, semantic retrieval is performed on the intelligent knowledge base to obtain a preset number of first query results and the vector similarity of the semantic retrieval corresponding to the first query results; and Based on the query statement, keyword retrieval is performed on the intelligent knowledge base to obtain a preset number of second query results and the word similarity of the second query results to the corresponding keyword retrieval.
[0093] In some embodiments, the retrieval unit 4002 performs semantic retrieval on the intelligent knowledge base in the following manner: Transform query statements into query vectors, and transform knowledge blocks in the intelligent knowledge base into knowledge block vectors; Determine the vector similarity between the query vector and each knowledge block vector; Based on the similarity of each vector, a preset number of first query results are determined from the intelligent knowledge base; In some embodiments, the retrieval unit 4002 specifically performs keyword retrieval on the intelligent knowledge base in the following manner: Retrieve at least one query keyword from the query statement; For each knowledge block, perform the following steps: For each query keyword, determine the frequency of occurrence of the query keyword in the knowledge block and the inverse document frequency of the query keyword in each knowledge block; Perform length normalization on the knowledge block to obtain the length of the knowledge block after length normalization; The sub-similarity between the knowledge block and the query keywords is determined based on the frequency of occurrence, inverse document frequency, and knowledge block length. The sum of the similarities of each sub-sub ... Based on the similarity of each word, a preset number of second query results are determined from the intelligent knowledge base.
[0094] In some embodiments, the final query unit 4004 is specifically used for: Classify the intent of the query statement to obtain the intent category to which the query statement belongs; Supplement the initial query results with contextual information based on the intent category; Based on a preset prompt template, a target prompt template is generated by combining the query statement. The initial query results, after supplementing with contextual information, are then filled into the target prompt template to obtain the final query results. This ensures that the final query results include the filename of the file to be stored, the chapter title in the file to be stored, and the chapter content of the file to be stored. The prompt template includes: file requirements, chapter title requirements, and chapter content requirements.
[0095] Based on the same inventive concept, this application also provides an electronic device. In this embodiment, the structure of the electronic device can be as follows: Figure 5 As shown, it includes a memory 501, a communication module 503, and one or more processors 502.
[0096] The memory 501 is used to store computer programs executed by the processor 502. The memory 501 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0097] Memory 501 may be volatile memory, such as random-access memory (RAM); memory 501 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 501 may be any other medium capable of carrying or storing a desired computer program having the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 501 may be a combination of the above-mentioned memories.
[0098] Processor 502 may include one or more central processing units (CPUs) or digital processing units, etc. Processor 502 is used to implement the above-described file query method when calling computer programs stored in memory 501.
[0099] The communication module 503 is used to communicate with terminal devices and other servers.
[0100] This application embodiment does not limit the specific connection medium between the memory 501, communication module 503, and processor 502 described above. This application embodiment... Figure 5 The memory 501 and the processor 502 are connected via a bus 504, and the bus 504 is in Figure 5 The diagram uses thick lines to describe the connections between other components; these are for illustrative purposes only and should not be considered limiting. The 504 bus can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 5 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.
[0101] The memory 501 stores a computer storage medium, which stores computer-executable instructions for implementing the file query method of this application embodiment. The processor 502 executes the aforementioned file query method. Based on the same inventive concept, this application embodiment provides a computer-readable storage medium, the computer program product including: computer program code, which, when run on a computer, causes the computer to execute any of the file query methods discussed above. Since the principle of the problem solved by the aforementioned computer-readable storage medium is similar to that of the file query method, the implementation of the aforementioned computer-readable storage medium can be referred to the implementation of the method, and repeated details will not be elaborated further.
[0102] The following reference Figure 6 To describe a computing device 600 according to this embodiment of the present application. Figure 6 The computing device 600 is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0103] like Figure 6 The computing device 600 is manifested in the form of a general-purpose computing device. The components of the computing device 600 may include, but are not limited to: at least one processing unit 601, at least one storage unit 602, and a bus 603 connecting different system components (including storage unit 602 and processing unit 601).
[0104] Bus 603 represents one or more of several bus structures, including a memory bus or memory controller, peripheral bus, processor, or local bus using any of the various bus structures.
[0105] Storage unit 602 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 621 and / or cache memory 622, and may further include read-only memory (ROM) 623.
[0106] Storage unit 602 may also include a program / utility 625 having a set (at least one) of program modules 624, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0107] The computing device 600 can also communicate with one or more external devices 604 (e.g., keyboard, pointing device, etc.), and with one or more devices that enable a query object to interact with the computing device 600, and / or with any device that enables the computing device 600 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 605. Furthermore, the computing device 600 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 606. Figure 6 As shown, network adapter 606 communicates with other modules for computing device 600 via bus 603. It should be understood that, although not shown in the figures, other hardware and / or software modules may be used in conjunction with computing device 600, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0108] This application also provides a computer program product. The methods in this application can be implemented, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in this application are executed, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a query target device, a core network device, an OAM (Operational Information Management) system, or other programmable devices.
[0109] A computer-readable storage medium can be an implementation of a computer program product. In other words, this application also provides a computer-readable storage medium that includes a computer program, which, when executed by a processor, implements any of the file query methods described above.
[0110] The computer program or instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions may be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium may be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; or an optical medium, such as a digital video optical disc; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0111] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0112] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0113] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0114] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0115] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A file retrieval method, characterized in that, The method includes: Retrieve the query statement entered by the query object; Based on the query statement, after performing a hybrid retrieval on the intelligent knowledge base, multiple query results and the similarity of each query result to each retrieval method are obtained; the hybrid retrieval includes multiple retrieval methods; the intelligent knowledge base includes: knowledge blocks obtained by dynamically slicing the file to be stored according to the internal hierarchical structure of each file to be stored; For each query result, based on the preset weight of each retrieval method, multiple similarities of the query results are weighted and fused, and an initial query result is determined from the multiple query results according to the weighted fusion result; The query statement is categorized by intent, and the initial query result is supplemented with contextual information based on the categorization results to obtain the final query result.
2. The method as described in claim 1, characterized in that, Each of the files to be stored is stored in the intelligent knowledge base in the following manner: The file to be stored is subjected to noise filtering and format normalization processing; Identify the chapter structure of the processed file to be stored, and obtain at least one chapter title in the processed file to be stored; For each chapter title, if the chapter title belongs to a preset independent chapter title, then the chapter content corresponding to the chapter title is taken as a knowledge block; if the chapter title does not belong to a preset independent chapter title, then the chapter content is dynamically sliced based on the length of the chapter content corresponding to the chapter title to obtain multiple knowledge blocks. The knowledge blocks are stored in the intelligent knowledge base.
3. The method as described in claim 2, characterized in that, The process involves dynamically slicing the chapter content based on the length of the chapter content corresponding to the chapter title to obtain multiple knowledge blocks, including: Obtain the length of the chapter content corresponding to the chapter title; when the length does not reach a preset length threshold, the chapter content corresponding to the chapter title is treated as a knowledge block; when the length reaches the preset length threshold, based on the preset length threshold and according to the natural boundaries of the chapter content, the chapter content corresponding to the chapter title is split into multiple knowledge blocks containing complete semantics; the length of the chapter content corresponding to each knowledge block does not exceed the preset length threshold.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: For each knowledge block, a corresponding chapter identifier is generated for the knowledge block based on the chapter title to which the knowledge block belongs; a character count identifier is generated for the knowledge block based on the number of characters contained in the knowledge block; and a unique identifier is generated for the knowledge block. Each knowledge block, along with its corresponding chapter identifier, character count identifier, and unique identifier, is stored in the intelligent knowledge base.
5. The method as described in claim 1, characterized in that, The hybrid retrieval includes semantic retrieval and keyword retrieval; after performing a hybrid retrieval on the intelligent knowledge base based on the query statement, multiple query results and the similarity of each query result to each retrieval method are obtained, including: Based on the query statement, semantic retrieval is performed on the intelligent knowledge base to obtain a preset number of first query results and the vector similarity of the first query results to the semantic retrieval; and Based on the query statement, keyword retrieval is performed on the intelligent knowledge base to obtain a preset number of second query results and the word similarity of the second query results to the keyword retrieval.
6. The method as described in claim 5, characterized in that, Semantic retrieval of the intelligent knowledge base is performed using the following methods: The query statement is converted into a query vector, and the knowledge blocks in the intelligent knowledge base are converted into knowledge block vectors; Determine the vector similarity between the query vector and each knowledge block vector; Based on the similarity of each vector, a preset number of first query results are determined from the intelligent knowledge base.
7. The method as described in claim 5, characterized in that, Keyword retrieval of the intelligent knowledge base can be performed using the following methods: Obtain at least one query keyword from the query statement; For each knowledge block, perform the following steps: For each query keyword, determine the frequency of occurrence of the query keyword in the knowledge block and the inverse document frequency of the query keyword in each knowledge block; The knowledge block is subjected to length normalization processing to obtain the length of the knowledge block after length normalization. Based on the occurrence frequency, the inverse document frequency, and the knowledge block length, the sub-similarity between the knowledge block and the query keyword is determined; The sum of the similarities of each sub-sub ... Based on the similarity of each word, a preset number of second query results are determined from the intelligent knowledge base.
8. The method as described in claim 1, characterized in that, The process of classifying the intent of the query statement and supplementing the initial query result with contextual information based on the classification results to obtain the final query result includes: The intent of the query statement is categorized to obtain the intent category to which the query statement belongs; The initial query results are supplemented with contextual information based on the intent category; Based on a preset prompt word template, a target prompt word template is generated in conjunction with the query statement. The initial query result, after supplementing with context information, is then filled into the target prompt word template to obtain the final query result. The final query result contains the filename of the file to be stored corresponding to the query statement, the chapter title in the file to be stored, and the chapter content of the file to be stored. The prompt word template includes: file requirements, chapter title requirements, and chapter content requirements.
9. A document retrieval device, characterized in that, include: The retrieval unit is used to retrieve the query statement input by the query object; The retrieval unit is used to perform a hybrid retrieval of the intelligent knowledge base based on the query statement, and obtain multiple query results and the similarity of each query result to each retrieval method; The hybrid retrieval includes multiple retrieval methods; the intelligent knowledge base includes: knowledge blocks obtained by dynamically slicing the files to be stored according to their internal hierarchical structure; An initial query unit is used to perform weighted fusion of multiple similarities of each query result based on a preset weight for each retrieval method, and determine an initial query result from the multiple query results based on the weighted fusion result. The final query unit is used to classify the intent of the query statement and supplement the initial query result with contextual information based on the classification result to obtain the final query result.
10. An electronic device, characterized in that, include: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the steps of the method according to any one of claims 1 to 8.