High-capacity text clue retrieval method and system based on enhanced retrieval, product and medium

By chunking and high-dimensional feature vector matrix processing of large-capacity text information, combined with dynamic feature library and multi-threaded scanning, the problems of low efficiency and insufficient accuracy in large-capacity electronic data retrieval are solved, and efficient and accurate clue retrieval is achieved.

CN120336494AActive Publication Date: 2025-07-18SHANGHAI XINREN INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510813019.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The prior art has problems in the search of large-capacity electronic data, such as low efficiency, insufficient accuracy, low recall and inability to capture hidden information, and large language models are prone to hallucinations, resulting in inaccurate search results.

Method used

Using an enhanced retrieval method, large-capacity text information is processed in blocks, converted into a high-dimensional feature vector matrix, and a dynamic feature library is introduced for business type classification. Through multi-threaded parallel scanning and association prompt word reply, ensuring full retrieval and accurate output.

Benefits of technology

The full scanning of large-capacity text information is realized, which improves the accuracy and recall rate of search, reduces processing time and hardware resource consumption, and ensures the relevance and accuracy of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336494A_ABST
    Figure CN120336494A_ABST
Patent Text Reader

Abstract

The invention discloses a high-capacity text clue retrieval method and system based on enhanced retrieval, a product and a medium. The method specifically comprises the steps that S100, high-capacity text information is partitioned; s200, converting the segmented text information blocks into a text high-dimensional feature vector matrix; s300, introducing a dynamic feature library in which different service type features are stored; s400, obtaining a target service feature vector matrix in the dynamic feature library based on the cue word and the service type input by the user, and obtaining an associated cue word associated with the target service feature vector matrix; s500, generating a retrieval unit, scanning the text information block, marking a scanning result meeting a hit condition of the retrieval unit as a hit text information block, and outputting the hit text information block; and S600, replying the associated prompt word in the range of the hit text information block. According to the method, accurate and efficient intelligent clue retrieval can be carried out on large-capacity texts, and accurate reply based on cue words is provided for case handling personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing, and specifically relates to a method, system, product, and medium for retrieving large-capacity text clues based on enhanced retrieval. Background Art

[0002] In the information age, criminal activities show highly digital and cross-regional characteristics, and electronic data has become the core source of evidence in case investigation. For example, in electronic evidence collection for public security case handling, chat records of instant messaging tools are used as important data analysis objects. Electronic data is characterized by a huge volume and complex structure. For example, in a single case, there may be tens of thousands or even hundreds of thousands of chat records and transaction flow data from different instant messaging tools or information platform APPs. The traditional manual screening mode has the disadvantages of low efficiency, high cost, and easy omission.

[0003] Currently, general large language models have made great progress in semantic understanding and natural language generation technologies in complex scenarios, and large language models are gradually being applied to electronic data governance, case clue induction, and evidence collection. However, there are still many problems in the current technical architecture in practical applications.

[0004] First, due to the huge volume of electronic data, natural language understanding based on large language models takes a long time, has a high feature complexity, and is limited by system resources, resulting in low efficiency and high difficulty in performing feature matching; some technologies adopt a "stop when the first hit occurs" retrieval strategy to improve retrieval efficiency, but this will cause the system to be unable to complete a full scan of all content, and there is a risk of missing retrieval fragment content. Second, the retrieval based on large language models cannot guarantee the accuracy of retrieval results. The accuracy of retrieval results depends on the matching degree between the hit fragments and the retrieval content, which has a certain randomness, and it is difficult to guarantee the accuracy of the first hit; and limited by the depth of semantic understanding, it is difficult to capture hidden information or coded language information during the scanning of electronic data, resulting in a problem of insufficient recall rate; in addition, it is also impossible to avoid the hallucinations generated by large language models, resulting in inaccurate information provided. Summary of the Invention

[0005] The purpose of this application is to overcome the deficiencies of the prior art, and provide a method, system, product, and medium for retrieving large-capacity text clues based on enhanced retrieval, which can perform accurate and efficient intelligent clue retrieval on large-capacity text, and provide accurate answers based on prompt words for case handlers.

[0006] In a first aspect, a method for retrieving large-capacity text clues based on enhanced retrieval provided by this application adopts the following technical solution: S100, divide the large-capacity text information into blocks with a certain number of characters as the limit, and divide the large-capacity text information into several consecutive text information blocks; S200. Convert each text information block into a text high-dimensional feature vector matrix through a large language processing model; S300. Introduce a dynamic feature library, which stores business type feature vector matrices differentiated based on different business types; each business type feature vector matrix is composed of several business vector matrices related to the business type, and the business vector matrix represents specific features in a specific business type; S400. Based on the prompt word and business type input by the user, extract the features of the prompt word through a large language processing model, then obtain the target business feature vector matrix in the corresponding business type in the dynamic feature library, and associate the prompt word with the target business feature vector matrix to obtain the associated prompt word; based on the prompt word input by the user, obtain several target business feature vector matrices; S500. For each target business feature vector matrix, generate a retrieval unit, set the window range of the retrieval unit and the hit condition of the text information block. Each retrieval unit independently performs a scan of the text high-dimensional feature vector matrix of the text information block based on the target business feature vector matrix. When the spatial similarity between the two meets the hit condition preset by the retrieval unit, mark the corresponding text information block as the hit text information block, output the corresponding hit text information block, and the finally output hit text information block is the sum of the hit text information blocks marked by each retrieval unit; for any one retrieval unit, create a corresponding scan thread, and each scan thread is responsible for the text information block scan task of the corresponding retrieval unit, and multiple scan threads execute the scan task in parallel; S600. Within the range of the hit text information block, perform a reply to the associated prompt word based on the large language processing model.

[0007] Through the above technical solution, after dividing the large-capacity text information into blocks, extract the high-dimensional feature vector matrix, then introduce a dynamic feature library classified according to business types, match the dynamic feature library after classifying the prompt words input by the user according to business types to obtain the target business feature vector matrix, and then scan the high-dimensional feature vector matrix of the large-capacity text information. After hitting, output the hit text information block and reply to the prompt word. The enhanced retrieval of this application based on the user's prompt word and the dynamic feature library improves the accuracy of the retrieved content; this application performs a full-scale retrieval of the large-capacity text information, and outputs and replies to the hit text information block, preventing the large language model from generating hallucinations while ensuring the retrieval range and recall rate; this application replies based on the associated prompt word, improving the relevance of the reply content to the user's actual reply target.

[0008] Through the above technical solution, multiple retrieval units are used to construct multi-parallel retrieval channels, which can achieve semantic coverage, realize feature fusion, enhance the complementarity between features, explore potential related clues, and improve the recall rate and completeness rate of retrieval; the multi-threaded parallel retrieval strategy can improve retrieval efficiency.

[0009] Preferably, for any retrieval unit, the retrieval range consists of a main text information block and several adjacent text information blocks of the main text information block.

[0010] Preferably, the adjacent text information blocks include adjacent contexts of text information published in the same instant messaging tool or information publishing platform software, and text information published in different instant messaging tools and / or information publishing platform software within the same time period.

[0011] Through the above technical solution, a dynamically sliding retrieval window formed by the main text information block and the adjacent text information block is realized, which can realize the continuous semantic capture of long texts in context, semantically associate text information across information sources, ensure the restoration of fragmented chat information, and improve retrieval coverage and accuracy.

[0012] Preferably, in S500, after marking the corresponding text information block as a hit text information block for the first time, the scanning task is continued until the scanning of the text high-dimensional feature vector matrix of all text information blocks is completed, and all marked hit text information blocks are sorted according to spatial similarity, and a set number of hit text information blocks are output from high to low.

[0013] Through the above technical solution, after executing a full search of large-capacity text information, the application prioritizes and sorts the hit text information blocks based on spatial similarity and then outputs the hit text information blocks. The output quantity can be adjusted according to user needs to ensure that the hit text information blocks that are strongly related to the user's intentions are preferentially displayed. Users can directly focus on highly relevant hit text information blocks without manually traversing all the results, thereby improving the ability to identify key evidence.

[0014] Preferably, the large-capacity text clue retrieval method based on enhanced retrieval further includes S700, updating the business type feature vector matrix in the dynamic feature library by manually supplementing or modifying the business type features based on the hit text information block.

[0015] Through the above technical solution, by manually identifying the hit text information blocks, the new clue features that change over time (such as new aliases, code words) can be updated to the dynamic feature library, so that the large language processing model retrieval strategy based on the dynamic feature library can adapt to new secret and complex clue scanning and collection services.

[0016] In a second aspect, the present application provides a large-capacity text clue retrieval system based on enhanced retrieval, including a text segmentation module, a text information processing module, a dynamic feature library module, a retrieval unit generation module, a text information scanning module, and an output module; The text segmentation module divides the large-capacity text information into several consecutive text information blocks with a certain number of characters as the limit; The text information processing module converts the text information blocks into a text high-dimensional feature vector matrix through a large language processing model; The dynamic feature library module stores the business type feature vector matrices of different business types; Based on the prompt words and business types input by the user, the retrieval unit generation module matches the target business feature vector matrix under the corresponding business type in the dynamic feature library module, then generates a retrieval unit, and associates the prompt words with the target business feature vector matrix to obtain the associated prompt words; Based on the retrieval unit, the text information scanning module scans the text high-dimensional feature vector matrix of the text information blocks based on the target business feature vector matrix. If the scanning result meets the preset hit condition of the retrieval unit, the corresponding text information block is marked as the hit text information block; The output module outputs the corresponding hit text information blocks, and within the range of the hit text information blocks, replies to the associated prompt words based on the large language processing model.

[0017] In a third aspect, the present application provides a computer program product, which includes a computer program or instruction, enabling the computer program or instruction to implement the steps in the above-mentioned large-capacity text clue retrieval method based on enhanced retrieval.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned large-capacity text clue retrieval method based on enhanced retrieval are implemented.

[0019] In summary, the present application includes at least one of the following beneficial technical effects: 1. The present application can achieve a full-scale scan of large-capacity text information, achieve full coverage of content, and avoid omission of potential clue information.

[0020] 2. By introducing a dynamic feature library, after the prompt words input by the user are subjected to feature extraction by the large language model, the present application matches the target business feature vector matrix according to the business type, and scans the text high-dimensional feature vector matrix of the text information blocks based on the target business feature vector matrix, improving the relevance between clue retrieval and user intent and business modality, and improving the retrieval accuracy.

[0021] 3. This application conducts retrieval based on multiple retrieval units. Each retrieval unit conducts retrieval based on the main text information block and several adjacent text information blocks, enhancing the depth of semantic understanding, improving the clue coverage rate of cross-platform group chats, and being able to successfully capture hidden conversation patterns, effectively improving the recall rate of the retrieval.

[0022] 4. This application sorts and outputs the spatial similarity of the hit text information blocks, reducing the difficulty for users to identify key clues and improving the judgment efficiency.

[0023] 5. This application makes full use of system resources, which can greatly shorten the processing time for large-capacity text information. At the same time, this application can reduce the consumption of hardware resources and the average calculation cost per case. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a schematic flow chart of the method for retrieving large-capacity text clues based on enhanced retrieval in the embodiment of this application; Figure 2 It is a schematic flow chart of obtaining the hit text information blocks through multiple retrieval units in the embodiment of this application; Figure 3 It is a schematic architecture diagram of the system for retrieving large-capacity text clues based on enhanced retrieval in the embodiment of this application; Figure 4 It is a schematic diagram of the structure of a computer device in the embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] This specific embodiment is only an interpretation of this application and does not limit this application. After reading this specification, those skilled in the art can make modifications without creative contributions to this embodiment as needed, but as long as it is within the scope of this application, it is protected by the patent law.

[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application. It should be noted that in the alternative embodiments of this application, for relevant data such as object information, when the embodiments in this application are applied to specific products or technologies, permission or consent from the object is required, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions. That is to say, if the embodiments of this application involve data related to an object, it needs to be obtained under the authorization and consent of the object, the authorization and consent of the relevant department, and compliance with the relevant laws, regulations, and standards of the relevant countries and regions. In the embodiments, if personal information is involved, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject needs to be obtained, and the embodiments also need to be implemented under the authorization and consent of the object.

[0027] In addition, the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, unless otherwise specified.

[0028] The embodiments of this application will be further described in detail below with reference to the accompanying drawings of the specification.

[0029] In one embodiment of this application, please refer to Figure 1 , and a large-capacity text clue retrieval method based on enhanced retrieval is provided, including the following steps: S100, for the large-capacity text information, it is segmented with a certain number of characters as the limit, and the large-capacity text information is divided into several consecutive text information blocks.

[0030] It should be noted that the character count limit for chunking large-capacity text information is preset based on the application scenarios of the targeted business types. The specific setting of the character count limit can be set by selecting a specific business type scenario or adjusted by the user's own setting. For example, the classification of business scenario types includes Scenario A for basic information verification, Scenario B for multi-source evidence chain analysis, and Scenario C for cross-domain comprehensive judgment. For Scenario A of basic information verification, the Tokens limit is approximately 1500 words. In scenarios such as ID number verification and case code matching, the input is fixed fields (such as 18-digit ID number + name), and the output is a boolean value or a preset label. Since long text generation is not required in this scenario, the relative Tokens limit is relatively low. For Scenario B of multi-source evidence chain analysis, due to the need for unstructured data processing, it is necessary to load PDF reports, email bodies, call record transcriptions, etc. simultaneously. The average single file occupies 800 - 1200 Tokens. Therefore, considering the multi-file scenario, the Tokens limit can be 5000 words. For Scenario C of cross-domain comprehensive judgment, the scenario involves the need for the integration of multi-modal data and structured data. For example, in mafia-related cases, it is necessary to synchronously analyze the fund map (structured table), surveillance video description (unstructured text), and communication network topology (graph data). After encoding, the Token occupancy exceeds the normal upper limit. Therefore, the Tokens limit needs to be customized comprehensively in combination with the specific scenario and hardware computing power.

[0031] S200. Convert each text information block into a text high - dimensional feature vector matrix through a large - language processing model. The optional types of the large - language processing model used are Qwen2.5 or DeepSeek R1. As a preference, in this embodiment, the Qwen model is adopted. The reasons are as follows: 1. Enterprise - level application maturity advantage: As the preferred solution in this embodiment, the core advantage of the Qwen model lies in its deep scenario adaptation ability verified through long - term business. In this application, since 2023, technical personnel have fully deployed Qwen series models in the intelligent judgment application model of electronic data. Through continuous scenario - based iteration, a customized ability covering the entire process of case judgment has been constructed: In a pilot work of a network - related case judgment, based on the intelligent text mining function developed by Qwen, the accuracy of the description scenario hit rate reaches as high as 82.3% (a 11.5% increase compared to the general model). This empirical data fully verifies its technical robustness in the judgment scenario and provides a reliable technical foundation for the technical solution in this application. 2. Model architecture technical advantage: Considering from the technical characteristics dimension of large - language models, the Qwen model demonstrates significantly better engineering performance than similar products: Its ultra - long context processing ability of 128k tokens (compared with the 64k upper limit of DeepSeek R1) improves the recall rate of key information in a 100,000 - word case file by 27% in the test, perfectly adapting to the ultra - long Tokens limit requirements for C - type business types (cross - domain comprehensive judgment scenarios) involved in the patent; at the same time, this model only requires 500 labeled data in the fine - tuning stage to reach a task accuracy of 90% + (a 37.5% reduction in labeling cost compared to competitors). This feature has decisive value in the processing of classified cases with scarce labeling resources.

[0032] S300. Introduce a dynamic feature library, which stores business - type feature vector matrices differentiated based on different business types. Each business - type feature vector matrix is actually composed of several business vector matrices related to the business type, and the business vector matrix represents specific features in a specific business type.

[0033] Specifically, electronic data differentiated based on different business types involves different case description scenario types, which can be classified into economic and financial scenarios, violent scenarios, network scenarios, etc. according to the large categories. Under the case description scenario type, specific business divisions are made for the involved crime types and case investigation types, and the corresponding dynamic feature library is introduced.

[0034] It should be noted that the dynamic feature library can be obtained by extracting feature vectors from an existing keyword library classified and matched based on business types, or can be obtained by training an artificial neural network with a clue training set with feature and business type labels, or can also be obtained by manually inputting in the feature library based on expert experience and extracting feature vectors. The dynamic feature library can be adjusted and expanded at any time according to actual needs and connected to corresponding other case feature libraries.

[0035] S400, based on the prompt word and business type input by the user, extracts the features of the prompt word through a large language processing model, and then obtains the target business feature vector matrix in the corresponding business type feature vector matrix under the corresponding business type in the dynamic feature library. In addition, the prompt word is associated with the target business feature vector matrix to obtain the associated prompt word, so as to enhance the business type relevance of the prompt word and improve the relevance of the reply content to the user's intention and business type when the subsequent large language processing model extracts information and replies to the prompt word.

[0036] It should be noted that for the user's input, the business type is selected in a drop-down menu through a preset existing business type, and the prompt word is input in natural language.

[0037] S500, generates a retrieval unit to scan the text high-dimensional feature vector matrix of the text information block based on the target business feature vector matrix. Then, when the spatial similarity between the two meets the hit condition preset by the retrieval unit, the corresponding text information block is marked as the hit text information block, and the corresponding hit text information block is output.

[0038] It should be noted that the parameters of the retrieval unit are specifically set by the user, and the setting content includes the window range of the retrieval unit, the hit condition of the text information block, etc. Specifically, in this embodiment, the hit condition can be the spatial similarity threshold of the text high-dimensional feature vector matrix with respect to the target business feature vector matrix, and only the hit text information blocks with spatial similarity exceeding the threshold are output.

[0039] S600, within the range of the hit text information block, replies to the associated prompt word based on the large language processing model. Strictly limiting the reply range of the large language processing model can prevent the model from referring to or creating clues outside the provided text information range, avoid the model from generating mitigation, and strictly ensure the authenticity of the clues. If the large language processing model cannot provide a reply content based on the user's prompt word, only the hit text information block is provided, and an empty set is output for the reply content. The user updates the prompt word to adjust the retrieval and reply strategies.

[0040] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0041] Further, in another embodiment, please refer to Figure 2 , based on the prompt words input by the user, several target service feature vector matrices are obtained. For each target service feature vector matrix, a retrieval unit is generated, and each retrieval unit independently performs a scan of the text high-dimensional feature vector matrix of the text information block, and the finally output hit text information block is the sum of the hit text information blocks marked by each retrieval unit.

[0042] It should be understood that since the prompt words are input by the user in natural language, after semantic understanding and feature extraction of the prompt words through a large language processing model, a plurality of different features will be generated. Generating target service feature vector matrices for different features respectively, generating different retrieval units, and performing subsequent clue scans can achieve semantic coverage, feature fusion, and improve the recall rate of retrieval. The user can uniformly set the parameters of different retrieval units, or set the parameters of each retrieval unit respectively based on the specific features represented by the retrieval units.

[0043] More specifically, for any one retrieval unit, a corresponding scan thread is created, and each scan thread is responsible for the text information block scan task of the corresponding retrieval unit, and multiple scan threads execute the scan tasks in parallel. Through the solution of this embodiment, dense vector calculation can be accelerated by GPU, different scan tasks can be assigned to different computing nodes, the data throughput can be improved, the single channel can be prevented from becoming a performance bottleneck, the utilization rate of hardware resources can be improved, and the processing time of large-capacity text information can be greatly reduced.

[0044] Furthermore, in another embodiment, for any one retrieval unit, its retrieval range consists of the main text information block and several adjacent text information blocks of the main text information block.

[0045] More specifically, the adjacent text information blocks include the adjacent context of the text information published in the same instant messaging tool or information publishing platform software, and the text information published in different instant messaging tools and / or information publishing platform software within the same time interval.

[0046] It should be noted that when setting the parameters of the retrieval unit, the user can selectively set whether to use the text information published in different instant messaging tools and / or information publishing platform software as adjacent text information blocks. If it is set not to, the adjacent text information blocks only contain the text information published in the same instant messaging tool or information publishing platform software. The parameter setting of the retrieval unit for adjacent text blocks can be the number of adjacent text blocks. For example, if it is set to 3, the text information blocks within 3 adjacent to the main text information block are the adjacent text information blocks. The parameter setting of the retrieval unit for adjacent text blocks can also be the same-source time interval. For example, if the same-source time interval is set to 2 days, all text information blocks occurring within two days adjacent to the main text information block are used as adjacent text information blocks. If it is set to be required, the adjacent text information blocks contain the text information published in different instant messaging tools and / or information publishing platform software. When setting the retrieval unit, it is necessary to additionally set the different-source time interval. For example, if the different-source time interval is set to 1 day, the text information published in different instant messaging tools and / or information publishing platform software within one day adjacent to the main text information block is used as adjacent text information blocks. The weight of the feature vector matrix of adjacent text information blocks with respect to the main text information block can be specifically set and adjusted. Through the above technical solution, an adjustable dynamic sliding window is established for clue retrieval, which can realize the context semantic understanding and clue capture of the same-source continuous long text, and can break through the barriers between different-source text information, establish the connection of different-source texts based on the time interval, conduct semantic comprehensive understanding, realize the semantic association across a single text information block, can significantly improve the hit accuracy and recall rate of clues, and can identify hidden code words and other dialogue patterns.

[0047] More specifically, after the corresponding text information block is first marked as the hit text information block, the scanning task is continued until the scanning of the text high-dimensional feature vector matrix of all text information blocks is completed. All the marked hit text information blocks are sorted according to the spatial similarity, and a set number of hit text information blocks are output from high to low. For multiple retrieval units, it is necessary to wait for the scanning tasks of all scanning threads to be completed before summarizing all the hit text information blocks. The user can set to comprehensively sort all the hit text information blocks, or sort them separately according to the hit text information blocks corresponding to different retrieval units. The specific output quantity is set by the user. For example, the user can choose to output the TOP 5 or TOP 10 hit text information blocks.

[0048] In another embodiment, the large-capacity text clue retrieval method based on enhanced retrieval further includes S700. Based on the hit text information block, the business type feature vector matrix in the dynamic feature library is updated by manually supplementing or modifying the business type features. Taking the anti-money laundering scenario as an example, criminals will adopt more and more complex and concealed money laundering means and communication codes to avoid detection. Therefore, it is necessary to dynamically adapt to new money laundering means and adjust, update, and maintain the dynamic feature library. Since the hit text information block is the text retrieved and output by the large language model, it is convenient for experts to analyze and summarize new features from it, thereby improving the real-time performance and adaptability of the system, and simplifying the analysis difficulty of experts and improving the analysis efficiency.

[0049] In another embodiment of the present application, please refer to Figure 3 , a large-capacity text clue retrieval system based on enhanced retrieval is provided, including a text segmentation module 1, a text information processing module 2, a dynamic feature library module 3, a retrieval unit generation module 4, a text information scanning module 5, and an output module 6.

[0050] The text segmentation module 1 divides the large-capacity text information into several consecutive text information blocks according to the user's setting with a certain number of characters as the limit.

[0051] The text information processing module 2 converts the text information block into a text high-dimensional feature vector matrix through a large language processing model.

[0052] The dynamic feature library module 3 stores the business type feature vector matrices of different business types.

[0053] Based on the prompt word and business type input by the user, the retrieval unit generation module 4 matches the target business feature vector matrix of the corresponding business type in the dynamic feature library module 3, then generates several retrieval units, and associates the prompt word with the target business feature vector matrix to obtain the associated prompt word.

[0054] Based on the retrieval unit, the text information scanning module 5 scans the text high-dimensional feature vector matrix of the text information block based on the target business feature vector matrix. If the scanning result meets the preset hit condition of the retrieval unit, the corresponding text information block is marked as a hit text information block.

[0055] The output module 6 outputs the corresponding hit text information block and gives a reply to the associated prompt word based on the large language processing model within the range of the hit text information block.

[0056] In other embodiments, the large-capacity text clue retrieval system based on enhanced retrieval further includes an update module 7 for updating the dynamic feature library module based on the hit text information block output by the output module.

[0057] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described large-capacity text clue retrieval system based on enhanced retrieval can refer to the corresponding process in the foregoing method embodiments and will not be elaborated herein.

[0058] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0059] In another embodiment of the present application, a computer program product is provided. The computer program product includes a computer program or instruction, causing the computer program or instruction to be able to perform the steps in the above-described large-capacity text clue retrieval method based on enhanced retrieval.

[0060] In another embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-described large-capacity text clue retrieval method based on enhanced retrieval.

[0061] In another embodiment of the present application, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store relevant data of the large-capacity text clue retrieval method based on enhanced retrieval, large-capacity text information, and dynamic feature libraries. The network interface of the computer device is used to communicate with an external terminal through a network connection to realize the input of external user prompt words, service types, and parameter adjustments, and output clue retrieval results to the user. When the computer program is executed by the processor, it implements the large-capacity text clue retrieval method based on enhanced retrieval.

[0062] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.

[0063] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware with a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as ROM or random access memory RAM, magnetic disks, or optical discs.

[0064] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for retrieving a large - volume text clue based on enhanced retrieval, characterized in that, It includes the following steps: S100, divide the large-capacity text information into chunks with a certain number of characters as the limit, and divide the large-capacity text information into several consecutive text information chunks; S200, convert each text information chunk into a text high-dimensional feature vector matrix through a large language processing model; S300, introduce a dynamic feature library, in which business type feature vector matrices differentiated based on different business types are stored; each business type feature vector matrix is composed of several business vector matrices related to the business type, and the business vector matrix represents specific features in a specific business type; S400, based on the prompt word and business type input by the user, extract the features of the prompt word through a large language processing model, then obtain the target business feature vector matrix in the corresponding business type in the dynamic feature library, and associate the prompt word with the target business feature vector matrix to obtain the associated prompt word; based on the prompt word input by the user, obtain several target business feature vector matrices; S500, for each target business feature vector matrix, generate a retrieval unit, set the window range of the retrieval unit and the hit condition of the text information chunk, and each retrieval unit independently performs a scan of the text high-dimensional feature vector matrix of the text information chunk based on the target business feature vector matrix. When the spatial similarity between the two meets the hit condition preset by the retrieval unit, mark the corresponding text information chunk as the hit text information chunk, output the corresponding hit text information chunk, and the finally output hit text information chunks are the sum of the hit text information chunks marked by each retrieval unit; for any one retrieval unit, create a corresponding scan thread, and each scan thread is responsible for the text information chunk scanning task of the corresponding retrieval unit, and multiple scan threads execute the scanning task in parallel; S600, within the range of the hit text information chunk, make a reply to the associated prompt word based on the large language processing model.

2. The method for retrieving large-capacity text clues based on enhanced retrieval according to claim 1, wherein For any one retrieval unit, its retrieval range consists of the main text information chunk and several adjacent text information chunks of the main text information chunk.

3. The method for retrieving a large-capacity text clue based on enhanced retrieval according to claim 2, wherein The adjacent text information chunks include the adjacent context of the text information published in the same instant messaging tool or information publishing platform software, and the text information published in different instant messaging tools and / or information publishing platform software within the same time interval.

4. The method for retrieving a large-capacity text clue based on enhanced retrieval according to claim 1, wherein In S500, after the corresponding text information chunk is first marked as the hit text information chunk, continue to execute the scanning task until the scanning of the text high-dimensional feature vector matrix of all text information chunks is completed, sort all the marked hit text information chunks according to the spatial similarity, and output a set number of hit text information chunks from high to low.

5. The method for retrieving a large-capacity text clue based on enhanced retrieval according to claim 1, wherein It also includes S700, based on the hit text information chunk, update the business type feature vector matrix in the dynamic feature library by manually supplementing or modifying the business type features.

6. A large-capacity text clue retrieval system based on enhanced retrieval, characterized in that, It includes a text segmentation module, a text information processing module, a dynamic feature library module, a retrieval unit generation module, a text information scanning module and an output module; The text segmentation module divides the large-capacity text information into several consecutive text information blocks with a certain number of characters as the limit; The text information processing module converts the text information blocks into a text high-dimensional feature vector matrix through a large language processing model; The dynamic feature library module stores the business type feature vector matrices of different business types; The retrieval unit generation module, based on the prompt words and business type input by the user, matches and obtains the target business feature vector matrix under the corresponding business type in the dynamic feature library module, then generates a retrieval unit, and associates the prompt words with the target business feature vector matrix to obtain the associated prompt words; The text information scanning module scans the text high-dimensional feature vector matrix of the text information blocks based on the retrieval unit using the target business feature vector matrix. If the scanning result meets the preset hit condition of the retrieval unit, the corresponding text information block is marked as the hit text information block; The output module outputs the corresponding hit text information blocks, and within the range of the hit text information blocks, replies to the associated prompt words based on the large language processing model.

7. A computer program product, characterized in that, The computer program product includes a computer program or instructions, enabling the computer program or instructions to implement the steps in the method for retrieving large-capacity text clues based on enhanced retrieval according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for retrieving large-capacity text clues based on enhanced retrieval according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Large model retrieval enhancement generation method and device

    CN117891838A

  • Multi-modal multi-scale multi-recall large language model retrieval enhancement generation method

    CN118296120A

  • Question and answer processing method and device, electronic equipment and storage medium

    CN119202151A

  • Intelligent retrieval method and system for unstructured asset content based on large model

    CN119646243A

  • Document retrieval enhancement method, device and equipment for large language model

    CN119938884A