Information retrieval method and system

By optimizing the information processing and index adjustment modules, the stability of the information retrieval system has been enhanced, the problem of incomplete retrieval caused by perfect matching in traditional systems has been solved, and more efficient multi-source information retrieval has been achieved.

CN121614604APending Publication Date: 2026-03-06BEIJING FABO HONGYE TECH DEV CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511829702.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Existing information retrieval systems rely solely on a perfect match mechanism without setting an approximate match threshold, resulting in incomplete retrieval content and insufficient retrieval stability.

Method used

By setting up information processing, retrieval sorting, conversion adjustment, expansion adjustment, and index adjustment modules, adjustments are made based on the field matching failure rate, average waiting time, and index latency of multi-source information. This increases the fault tolerance coefficient for format conversion, reduces the dynamic expansion threshold of preprocessing nodes, increases the minimum guaranteed proportion of index update resources, and optimizes data processing and index update strategies.

Benefits of technology

It improves the stability of multi-source information retrieval, reduces format conversion failure rate and data backlog, ensures data processing speed and real-time index updates, and reduces system operating pressure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121614604A_ABST
    Figure CN121614604A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information retrieval processing, in particular to an information retrieval method and system, and the system comprises an information processing module which comprises an acquisition unit for acquiring multi-source information of a plurality of data sources; the retrieval sorting module comprises a model training unit which is used for training an initial model according to the multi-source information and the lexical item sequence so as to output a retrieval model; the conversion adjusting module is used for adjusting the fault-tolerant coefficient of format conversion according to the field matching failure rate of the multi-source information; the capacity expansion adjusting module is used for adjusting the dynamic capacity expansion threshold value of the preprocessing node according to the average waiting duration of the multiple pieces of multi-source information processing; and the index adjusting module is used for adjusting the index updating resource minimum guarantee proportion according to the index delay duration of the retrieval requirement. According to the method, the retrieval stability of the multi-source information is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information retrieval and processing technology, and in particular to an information retrieval method and system. Background Technology

[0002] In today's era of widespread digital information technology and intelligent services, information retrieval methods and systems, with their core advantages of efficient information filtering, cross-source data integration, and precise demand matching, have become a key technological foundation supporting the efficient operation of the digital economy and intelligent services. In the context of drug information retrieval, their core advantages of accurate information positioning, compliant data filtering, and rapid extraction of key information have made them crucial technical support for clinical diagnosis and treatment, drug regulation, and patient consultation. As the data and application scenarios involved in the retrieval process exhibit significant characteristics of multi-source and dynamic nature, user needs are shifting from precise matching to approximate retrieval. The requirements for real-time retrieval, result recall, and fault tolerance vary significantly across different scenarios. Against this backdrop, traditional information retrieval methods and systems have gradually revealed numerous technical bottlenecks, making it difficult to meet the needs of efficient, accurate, and flexible information retrieval in complex scenarios.

[0003] Chinese Patent Publication No. CN101796493A discloses an information retrieval system, information retrieval method, and program. The system includes: a receiving unit that receives retrieval object condition information indicating the conditions for a retrieval object, source information retrieval location information indicating whether the source information for the retrieval object exists in either a file or a memory, and the source information name of the source information for the retrieval object; a management table that associates and stores the names of memory regions with memory region information indicating the memory regions; a retrieval unit that, when the source information for the retrieval object indicated by the received source information retrieval location information exists in the memory, retrieves from the management table the name of a memory region that matches the source information name of the received source information for the retrieval object; and an obtaining unit that, when the retrieval unit retrieves the name of a memory region that matches the source information name of the received source information for the retrieval object, retrieves information that matches the received retrieval object condition information from the memory region indicated by the memory region information associated with the retrieved memory region name. It is evident that the information retrieval system, information retrieval method, and program suffer from insufficient information retrieval stability due to the fact that only a completely identical matching mechanism is used when retrieving source information names, and no approximate matching threshold is set, resulting in incomplete retrieval content. Summary of the Invention

[0004] To address this issue, the present invention provides an information retrieval method and system to overcome the problem in the prior art where the retrieval of source information is unstable due to incomplete retrieval content caused by the use of only a completely identical matching mechanism and the absence of an approximate matching threshold when retrieving source information names.

[0005] To achieve the above objectives, the present invention provides an information retrieval system, comprising: The information processing module includes a collection unit for collecting multi-source information from several data sources, a preprocessing unit connected to the collection unit for sequentially performing format conversion, cleaning, word segmentation, and vocabulary normalization on the multi-source information to output a word sequence, and an index building unit connected to the preprocessing unit for building a global index based on the word sequence. The retrieval and sorting module, which is connected to the information processing module, includes a model training unit for training an initial model based on the multi-source information and the term sequence to output a retrieval model, and a retrieval unit connected to the model training unit for querying retrieval requirements to output retrieval results. A conversion adjustment module, which is connected to the information processing module, is used to adjust the fault tolerance coefficient of the format conversion according to the field matching failure rate of multi-source information; The expansion adjustment module is connected to the information processing module and the conversion adjustment module respectively, and is used to adjust the dynamic expansion threshold of the preprocessing node according to the average waiting time of several multi-source information processing. The index adjustment module, which is connected to the retrieval sorting module and the expansion adjustment module respectively, is used to adjust the minimum guaranteed ratio of index update resources according to the index delay time of retrieval needs.

[0006] Furthermore, the conversion adjustment module determines that the retrieval stability of the multi-source information does not meet the requirements when the field matching failure rate of the multi-source information is greater than the preset first failure rate.

[0007] Furthermore, in response to the field matching failure rate of the multi-source information being greater than the preset first failure rate and less than or equal to the preset second failure rate, the conversion adjustment module initially determines that the timeliness of the term sequence does not meet the requirements, and determines whether the timeliness of the term sequence meets the requirements based on the average waiting time of processing several pieces of multi-source information.

[0008] Furthermore, the conversion adjustment module increases the fault tolerance coefficient of the format conversion in response to the field matching failure rate of the multi-source information being greater than the preset second failure rate; The increase in the fault tolerance coefficient of the format conversion is determined by the difference between the field matching failure rate of multi-source information and the preset second failure rate.

[0009] Furthermore, the expansion adjustment module responds to the fact that the average waiting time for processing several multi-source information is greater than the preset first waiting time, and determines that the timeliness of the term sequence does not meet the requirements.

[0010] Furthermore, the expansion adjustment module responds to the fact that the average waiting time of processing several multi-source information is greater than the preset first waiting time and less than or equal to the preset second waiting time, and reduces the dynamic expansion threshold of the preprocessing node. The expansion adjustment module responds when the average waiting time for processing several multi-source information items exceeds the preset second waiting time, initially determining that the real-time update of the global index does not meet the requirements, and then determines whether the real-time update of the global index meets the requirements based on the index delay time of the retrieval needs.

[0011] Furthermore, the reduction in the dynamic expansion threshold of the preprocessing node is determined by the difference between the average waiting time of the processing of the multiple multi-source information and the preset first waiting time.

[0012] Furthermore, if the index delay duration in response to retrieval requests exceeds the preset delay duration, the index adjustment module determines that the real-time update of the global index does not meet the requirements and increases the minimum guaranteed percentage of index update resources.

[0013] Furthermore, the increase in the minimum guaranteed percentage of index update resources is determined by the difference between the index delay time of the retrieval demand and the preset delay time.

[0014] The present invention also provides an information retrieval method, comprising: The collected multi-source information from several data sources is sequentially processed through format conversion, cleaning, word segmentation, and vocabulary normalization to output a term sequence. A global index is then constructed using the mapping relationship between the term sequence and the multi-source information. The initial model is trained using the multi-source information and the term sequence to output a retrieval model. The retrieval model is then used to perform queries according to the retrieval requirements to output retrieval results. Obtain the field matching failure rate of multi-source information, and determine whether the retrieval stability of multi-source information meets the requirements based on the field matching failure rate of multi-source information; If the retrieval stability of multi-source information does not meet the requirements, then determine whether it is necessary to increase the fault tolerance coefficient of format conversion; If it is not necessary to increase the fault tolerance coefficient of the format conversion, then obtain the average waiting time of processing several multi-source information to determine whether the timeliness of the term sequence meets the requirements; If the timeliness of the term sequence does not meet the requirements, determine whether it is necessary to reduce the dynamic expansion threshold of the preprocessing node; If it is not necessary to reduce the dynamic expansion threshold of the preprocessing nodes, then the minimum guaranteed percentage of index update resources is determined based on the index latency duration of retrieval needs.

[0015] Compared with existing technologies, the beneficial effects of this invention are as follows: The system of this invention, by setting up an information processing module, a retrieval and sorting module, a conversion adjustment module, a capacity expansion adjustment module, and an index adjustment module, adjusts the fault tolerance coefficient of format conversion based on the field matching failure rate of multi-source information. Since the data structures and field definitions of different platforms are inconsistent when collecting data, field matching chaos and format noise can occur during merging. By increasing the fault tolerance coefficient of format conversion, more non-standard formats can be compatible, reducing conversion failures caused by format incompatibility, significantly lowering the format conversion failure rate, and reducing explicit format noise. The dynamic capacity expansion threshold of the preprocessing node is adjusted based on the average waiting time for processing several pieces of multi-source information. Since data transmission to the system needs to undergo cleaning, conversion, and fusion processes, if... Insufficient processing node resources lead to data backlog in the processing queue, resulting in untimely data updates. By reducing the dynamic expansion threshold of preprocessing nodes, the window period for supplementing computing power can be shortened, allowing processing capacity to match the data addition rate earlier. This alleviates backlog at the source, avoids processing efficiency losses due to extreme occupancy, and ensures data processing speed. The minimum guaranteed percentage of index update resources is adjusted based on the index latency of retrieval needs. Since index updates require computing resources, the system prioritizes query requests during peak retrieval periods, pausing or delaying index updates. By increasing the minimum guaranteed percentage of index update resources, basic updates can be maintained during peak periods, reducing the amount of delayed backlog index tasks and avoiding the need to process a large number of backlog tasks during off-peak periods. This reduces the subsequent operating pressure on the system and improves the retrieval stability of multi-source information.

[0016] Furthermore, the system described in this invention adjusts the fault tolerance coefficient of format conversion by setting a preset first failure rate and a preset second failure rate. Since the data structures and field definitions of different platforms are inconsistent when collecting data, field matching will be chaotic during merging, resulting in format noise. By increasing the fault tolerance coefficient of format conversion, more non-standard formats can be compatible, reducing conversion failures caused by format incompatibility, significantly reducing the format conversion failure rate, reducing explicit format noise, and further improving the retrieval stability of multi-source information.

[0017] Furthermore, the system of the present invention adjusts the dynamic expansion threshold of the preprocessing nodes by setting a preset first waiting time and a preset second waiting time. Since data needs to undergo cleaning, transformation, and fusion after being transmitted to the system, insufficient processing node resources can lead to data backlog in the processing queue, resulting in untimely data updates. By reducing the dynamic expansion threshold of the preprocessing nodes, the window period for computing power replenishment can be shortened, allowing processing capacity to match the data addition rate earlier, alleviating backlog from the source, avoiding processing efficiency loss due to extreme occupancy, ensuring data processing speed, and further improving the retrieval stability of multi-source information.

[0018] Furthermore, the system described in this invention adjusts the minimum guaranteed ratio of index update resources by setting a preset delay duration. Since index updates require computing resources, the system will prioritize query requests during peak retrieval periods and will pause or delay index updates. By increasing the minimum guaranteed ratio of index update resources, basic updates can be maintained during peak periods, reducing the amount of delayed backlog index tasks and avoiding the need to process a large number of backlog tasks during off-peak periods. This reduces the subsequent operating pressure on the system and further improves the retrieval stability of multi-source information. Attached Figure Description

[0019] Figure 1 This is a block diagram of the overall structure of the information retrieval system according to an embodiment of the present invention; Figure 2 This is a logical flowchart illustrating the process of determining the fault tolerance coefficient for format conversion in the information retrieval system according to an embodiment of the present invention. Figure 3 This is a flowchart illustrating the process of determining the dynamic expansion threshold of the preprocessing node in the information retrieval system according to an embodiment of the present invention. Figure 4 This is an overall flowchart of the information retrieval method according to an embodiment of the present invention. Detailed Implementation

[0020] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0021] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0022] Please see Figure 1 The diagram shown is an overall structural block diagram of the information retrieval system according to an embodiment of the present invention.

[0023] This invention provides an information retrieval system, comprising: The information processing module includes a collection unit for collecting multi-source information from several data sources, a preprocessing unit connected to the collection unit for sequentially performing format conversion, cleaning, word segmentation, and vocabulary normalization on the multi-source information to output a word sequence, and an index building unit connected to the preprocessing unit for building a global index based on the word sequence. The retrieval and sorting module, which is connected to the information processing module, includes a model training unit for training an initial model based on the multi-source information and the term sequence to output a retrieval model, and a retrieval unit connected to the model training unit for querying retrieval requirements to output retrieval results. A conversion adjustment module, which is connected to the information processing module, is used to adjust the fault tolerance coefficient of the format conversion according to the field matching failure rate of multi-source information; The expansion adjustment module is connected to the information processing module and the conversion adjustment module respectively, and is used to adjust the dynamic expansion threshold of the preprocessing node according to the average waiting time of several multi-source information processing. The index adjustment module, which is connected to the retrieval sorting module and the expansion adjustment module respectively, is used to adjust the minimum guaranteed ratio of index update resources according to the index delay time of retrieval needs.

[0024] Specifically, the application fields of the embodiments of the present invention can be the financial field, the medical field, and the academic literature field.

[0025] Specifically, several data sources include stock market data sources, drug information databases, and e-commerce transaction information platforms.

[0026] Specifically, multi-source information includes stock codes, adverse drug reaction information, and daily trading volume information.

[0027] Specifically, a term sequence is a sequence formed by extracting key elements from multi-source information and arranging them in a certain order.

[0028] Specifically, the key elements are the core terms that can represent the meaning of the textual information within the information context.

[0029] Specifically, the process of constructing a global index based on the term sequence involves creating and recording the mapping relationship between terms and multi-source information, and constructing an inverted index.

[0030] Specifically, an inverted index is an index structure used to store the mapping relationship between terms and multi-source information containing those terms.

[0031] Specifically, the process of training the initial model based on the multi-source information and the term sequence involves extracting features from the term sequence and multi-source information respectively, obtaining term features and information features, and dividing them into training set, validation set and test set. The initial model is then trained using the training set according to the corresponding loss function, and the parameters are optimized to minimize the loss function. The hyperparameters are adjusted using the validation set to avoid overfitting, and the recall rate and other indicators are validated using the test set, and the retrieval model is output.

[0032] Specifically, the initial model is a basic framework with basic learning capabilities and a general structure.

[0033] Specifically, the retrieval models include the Boolean retrieval model, the BM25 model, and the dual-tower model.

[0034] Specifically, search requests include stock quotes, drug names, and product names.

[0035] Specifically, the search results include real-time transaction prices, labeled retail prices of medicines, and product packaging specifications.

[0036] Specifically, the format conversion tolerance coefficient is a parameter that measures the maximum tolerance of data format deviations in the format conversion process.

[0037] Specifically, data format deviations include missing fields, non-standard formats, and encoding anomalies.

[0038] Specifically, the dynamic expansion threshold for preprocessing nodes is the maximum system resource utilization rate that triggers the automatic increase in the number of preprocessing computing nodes.

[0039] Specifically, the minimum guaranteed percentage of index update resources is the minimum resource allocation percentage that the system reserves specifically for index building and update operations.

[0040] In implementation, the system of this invention adjusts the fault tolerance coefficient of format conversion based on the field matching failure rate of multi-source information by setting up an information processing module, a retrieval and sorting module, a conversion adjustment module, a capacity expansion adjustment module, and an index adjustment module. Since the data structures and field definitions of different platforms are inconsistent when collecting data, field matching chaos and format noise can occur during merging. By increasing the fault tolerance coefficient of format conversion, more non-standard formats can be accommodated, reducing conversion failures caused by format incompatibility, significantly lowering the format conversion failure rate, and reducing explicit format noise. The dynamic capacity expansion threshold of the preprocessing node is adjusted based on the average waiting time for processing several pieces of multi-source information. Since data transmission to the system requires cleaning, conversion, and fusion, if the processing node resources are insufficient... Insufficient processing capacity leads to data backlog in the processing queue, resulting in untimely data updates. By reducing the dynamic expansion threshold of preprocessing nodes, the window of opportunity for computing power replenishment can be shortened, allowing processing capacity to match the data addition rate earlier. This alleviates backlog at the source, avoids processing efficiency losses due to extreme occupancy, and ensures data processing speed. The minimum guaranteed percentage of index update resources is adjusted according to the index latency of retrieval needs. Since index updates require computing resources, the system will prioritize query requests during peak retrieval periods and may pause or delay index updates. By increasing the minimum guaranteed percentage of index update resources, basic updates can be maintained during peak periods, reducing the amount of delayed backlog index tasks and avoiding the need to process a large number of backlog tasks during off-peak periods. This reduces the subsequent operating pressure on the system and improves the retrieval stability of multi-source information.

[0041] Please continue reading Figure 2 The diagram shown is a logical flowchart of the process for determining the fault tolerance coefficient of the format conversion in the information retrieval system of this embodiment of the invention.

[0042] Specifically, the conversion adjustment module determines that the retrieval stability of the multi-source information meets the requirements when the field matching failure rate of the multi-source information is less than or equal to a preset first failure rate. The conversion adjustment module determines that the retrieval stability of the multi-source information does not meet the requirements when the field matching failure rate of the multi-source information is greater than the preset first failure rate.

[0043] Specifically, the conversion adjustment module responds to the fact that the field matching failure rate of the multi-source information is greater than the preset first failure rate and less than or equal to the preset second failure rate, initially determines that the timeliness of the term sequence does not meet the requirements, and determines whether the timeliness of the term sequence meets the requirements based on the average waiting time of processing several pieces of multi-source information.

[0044] It is understandable that the preset first failure rate is lower than the preset second failure rate. The three intervals divided by the preset first failure rate and the preset second failure rate correspond to three different scenarios: The first interval is when the field matching failure rate of multi-source information is less than or equal to the preset first failure rate, which corresponds to the situation where the retrieval stability of multi-source information meets the requirements. The second interval is when the field matching failure rate of multi-source information is greater than the preset first failure rate and less than or equal to the preset second failure rate. The corresponding situation is: after the data is transmitted to the system, it needs to be cleaned, transformed, and merged. If the processing node resources are insufficient, the data will accumulate in the processing queue, resulting in untimely data updates. The third interval is where the field matching failure rate of multi-source information is greater than the preset second failure rate. The corresponding situation is that when data is collected from different platforms, the data structures and field definitions of each platform are inconsistent, which will cause field matching chaos and form format noise during merging.

[0045] Understandably, in information retrieval systems, using preset first and second failure rates to characterize the retrieval stability of multi-source information is based on the core logic of transforming multi-source information retrieval stability into a quantifiable failure rate range. The preset first failure rate serves as a dividing line for determining whether the stability of multi-source information retrieval is adequate, its core function being to determine whether data processing and field matching meet the basic requirements of the information retrieval system. The preset second failure rate serves as a dividing line for determining the severity of problems affecting retrieval stability, its core function being to distinguish the types of problems where the system's retrieval stability does not meet requirements. The preset first and second failure rates can be set according to actual operating conditions. The setting of the preset first and second failure rates aims to ensure the stability and usability of multi-source information retrieval. Optionally, the preset first and second failure rates are determined through a limited number of trials by evaluating the retrieval effect of different matching failure rates on multi-source information. The determined preset first and second failure rates should be neither too low nor excessively interfere with the multi-source information retrieval process. For example, the preset first failure rate is generally selected in the range of [0.9%, 1.1%], and the preset second failure rate is generally selected in the range of [4.9%, 6.1%].

[0046] Preferably, the first failure rate is 1.0% in a preferred embodiment, and the second failure rate is 5.0% in a preferred embodiment.

[0047] Specifically, the field matching failure rate for multi-source information is the ratio of the number of fields that do not match between several data sources to the total number of information fields.

[0048] Specifically, the conversion adjustment module increases the fault tolerance coefficient of the format conversion in response to the field matching failure rate of the multi-source information being greater than the preset second failure rate; The increase in the fault tolerance coefficient of the format conversion is determined by the difference between the field matching failure rate of multi-source information and the preset second failure rate.

[0049] Specifically, when the difference between the field matching failure rate of multi-source information and the preset second failure rate is within 1%, the fault tolerance coefficient of the format conversion increases to 1.1 times the original value. When the difference between the field matching failure rate of multi-source information and the preset second failure rate exceeds 1%, the fault tolerance coefficient of the format conversion increases by 0.02 for every 0.5% increase beyond the original 1.1 times. For example, when the difference between the field matching failure rate of multi-source information and the preset second failure rate is 2%, the current fault tolerance coefficient of the format conversion is 0.4, and the increased fault tolerance coefficient of the format conversion is 0.4×1.1+0.02×2=0.48.

[0050] In practice, the system described in this invention adjusts the fault tolerance coefficient of format conversion by setting a preset first failure rate and a preset second failure rate. Since the data structures and field definitions of different platforms are inconsistent when collecting data, field matching will be chaotic and form format noise during merging. By increasing the fault tolerance coefficient of format conversion, more non-standard formats can be compatible, reducing conversion failures caused by format incompatibility, significantly reducing the format conversion failure rate, reducing explicit format noise, and further improving the retrieval stability of multi-source information.

[0051] Please continue reading Figure 3 The diagram shown is a logical flowchart of the process of determining the dynamic expansion threshold of the preprocessing node in the information retrieval system of this embodiment of the invention.

[0052] Specifically, the expansion adjustment module responds to the average waiting time of processing several multi-source information items being less than or equal to a preset first waiting time, thus determining that the timeliness of the term sequence meets the requirements; The expansion adjustment module responds to the fact that the average waiting time for processing several multi-source information is greater than the preset first waiting time, and determines that the timeliness of the term sequence does not meet the requirements.

[0053] Specifically, the expansion adjustment module responds to the fact that the average waiting time of processing several multi-source information is greater than the preset first waiting time and less than or equal to the preset second waiting time, and reduces the dynamic expansion threshold of the preprocessing node. The expansion adjustment module responds when the average waiting time for processing several multi-source information items exceeds the preset second waiting time, initially determining that the real-time update of the global index does not meet the requirements, and then determines whether the real-time update of the global index meets the requirements based on the index delay time of the retrieval needs.

[0054] It is understandable that the preset first waiting time is shorter than the preset second waiting time, and the three intervals divided by the preset first and second waiting times correspond to three different scenarios: The first interval is where the average waiting time for processing several multi-source information items is less than or equal to the preset first waiting time, which corresponds to the situation where the timeliness of the term sequence meets the requirements. The second interval is the average waiting time for processing several multi-source information that is greater than the preset first waiting time and less than or equal to the preset second waiting time. The corresponding situation is: after the data is transmitted to the system, it needs to be cleaned, transformed, and merged. If the processing node resources are insufficient, the data will accumulate in the processing queue, resulting in untimely data updates. The third interval is when the average waiting time for processing several multi-source information items is greater than the preset second waiting time. The corresponding situation is: since index updates require computing resources, during peak retrieval periods, the system will prioritize query requests and will pause or delay index updates.

[0055] Understandably, in information retrieval systems, the use of preset first and second waiting times to characterize the timeliness of term sequences is based on the core logic of transforming the timeliness of term sequences into a quantifiable average waiting time range. The preset first waiting time serves as the dividing line for determining whether the timeliness of a term sequence is acceptable; its core function is to determine whether the data processing and response time meets the basic timeliness requirements of information retrieval. The preset second waiting time serves as the dividing line for determining the severity of timeliness issues; its core function is to distinguish how to adjust and alleviate situations where the system's timeliness fails to meet requirements. The preset first and second waiting times can be set according to actual operating conditions. The setting of the preset first and second waiting times aims to ensure the stability and practicality of multi-source information retrieval. Optionally, the preset first and second waiting times are determined through a limited number of experiments by evaluating the retrieval effect of different waiting times on multi-source information. The determined preset first and second waiting times should be neither too small nor excessively interfere with the multi-source information retrieval process. For example, the preset first waiting time is generally selected in the range of [400ms, 600ms], and the preset second waiting time is generally selected in the range of [900ms, 1100ms].

[0056] Preferably, the first waiting time is 500ms in a preferred embodiment, and the second waiting time is 1000ms in a preferred embodiment.

[0057] Specifically, the average waiting time for processing several pieces of multi-source information is the ratio of the total waiting time in the processing of several pieces of multi-source information to the total number of pieces of multi-source information.

[0058] Specifically, the reduction in the dynamic expansion threshold of the preprocessing node is determined by the difference between the average waiting time of the processing of the multiple multi-source information and the preset first waiting time.

[0059] Specifically, when the difference between the average waiting time of several multi-source information processing sessions and the preset first waiting time is within 100ms, the dynamic expansion threshold of the preprocessing node is reduced to 0.9 times the original value. When the difference between the average waiting time of several multi-source information processing sessions and the preset first waiting time exceeds 100ms, the dynamic expansion threshold of the preprocessing node is reduced by 2% for every 50ms exceeding the original value, in addition to being reduced to 0.9 times the original value. For example, when the difference between the average waiting time of several multi-source information processing sessions and the preset first waiting time is 200ms, the current dynamic expansion threshold of the preprocessing node is 70%, and the reduced dynamic expansion threshold of the preprocessing node is 70%×0.9-2%×2=59%.

[0060] In implementation, the system of the present invention adjusts the dynamic expansion threshold of the preprocessing nodes by setting a preset first waiting time and a preset second waiting time. Since data needs to be cleaned, transformed, and fused after being transmitted to the system, insufficient processing node resources can cause data to accumulate in the processing queue, resulting in untimely data updates. By reducing the dynamic expansion threshold of the preprocessing nodes, the window period for replenishing computing power can be shortened, allowing processing capacity to match the data addition rate earlier, alleviating the backlog from the source, avoiding processing efficiency loss due to extreme occupancy, ensuring data processing speed, and further improving the retrieval stability of multi-source information.

[0061] Specifically, the index adjustment module responds to the retrieval request with an index delay time that is less than or equal to a preset delay time, thus determining that the real-time update of the global index meets the requirements. The index adjustment module responds to the index delay duration of the retrieval request being greater than the delay duration, determines that the real-time update of the global index does not meet the requirements, and increases the minimum guaranteed proportion of index update resources.

[0062] Specifically, the increase in the minimum guaranteed percentage of index update resources is determined by the difference between the index delay time of the retrieval requirement and the preset delay time.

[0063] It is understandable that the two intervals divided by the preset delay duration correspond to two different scenarios: The first interval is when the index delay time for retrieval needs is less than or equal to the preset delay time, which corresponds to the situation where the real-time update of the global index meets the requirements. The second interval is when the index delay time for retrieval needs is greater than the preset delay time. The corresponding situation is: since index updates require computing resources, the system will prioritize query requests during peak retrieval periods and will pause or delay index updates.

[0064] Understandably, in information retrieval systems, using a preset delay duration to characterize the real-time performance of the global index update essentially transforms the abstract concept of real-time performance into a quantifiable index delay duration range. The preset delay duration serves as a critical dividing line for determining whether the real-time performance of the global index update is adequate, its core function being to clarify whether the index update status meets the basic real-time requirements of information retrieval. The preset delay duration can be set according to actual operating conditions. The setting of the preset delay duration aims to ensure the stability and usability of multi-source information retrieval. Optionally, the preset delay duration is determined through a limited number of trials by evaluating the retrieval effect of different delay durations on multi-source information. The determined preset delay duration should be neither too small nor cause excessive interference to the multi-source information retrieval process. For example, the preset delay duration is generally selected within the range of [40ms, 60ms].

[0065] Preferably, the preset delay duration is 50ms.

[0066] Specifically, the indexing delay time for a retrieval request is the difference between the actual time taken from receiving the retrieval request to the system generating retrieval results and returning them to the user, and the theoretical time taken.

[0067] Specifically, the increase in the minimum guaranteed percentage of index update resources is determined by the difference between the index delay time of the retrieval requirement and the preset delay time.

[0068] Specifically, when the difference between the index delay time of the retrieval request and the preset delay time is within 10ms, the minimum guaranteed ratio of index update resources increases to 1.1 times the original value. When the difference between the index delay time of the retrieval request and the preset delay time exceeds 10ms, in addition to increasing to 1.1 times the original value, the minimum guaranteed ratio of index update resources increases by 2% for every 5ms exceeding the original value. For example, when the difference between the index delay time of the retrieval request and the preset delay time is 20ms, the current minimum guaranteed ratio of index update resources is 10%, and the increased minimum guaranteed ratio of index update resources is 10×1.1+2×2=15%.

[0069] In practice, the system described in this invention adjusts the minimum guaranteed ratio of index update resources by setting a preset delay duration. Since index updates require computing resources, the system will prioritize query requests during peak retrieval periods and will pause or delay index updates. By increasing the minimum guaranteed ratio of index update resources, basic updates can be maintained during peak periods, reducing the amount of delayed backlog index tasks and avoiding the need to process a large number of backlog tasks during off-peak periods. This reduces the subsequent operating pressure on the system and further improves the retrieval stability of multi-source information.

[0070] Please continue reading Figure 4 The diagram shown is an overall flowchart of the information retrieval method according to an embodiment of the present invention.

[0071] An information retrieval method, comprising: Step S1: The multi-source information collected from several data sources is sequentially processed by format conversion, cleaning, word segmentation and vocabulary normalization to output a word sequence. A global index is constructed using the mapping relationship between the word sequence and the multi-source information. Step S2: Use the multi-source information and the term sequence to train the initial model to output a retrieval model, and use the retrieval model to perform a query according to the retrieval requirements to output retrieval results; Step S3: Obtain the field matching failure rate of multi-source information, and determine whether the retrieval stability of multi-source information meets the requirements based on the field matching failure rate of multi-source information; Step S4: If the retrieval stability of multi-source information does not meet the requirements, determine whether it is necessary to increase the fault tolerance coefficient of format conversion. Step S5: If it is not necessary to increase the fault tolerance coefficient of the format conversion, then obtain the average waiting time of processing several multi-source information to determine whether the timeliness of the term sequence meets the requirements. Step S6: If the timeliness of the term sequence does not meet the requirements, determine whether it is necessary to reduce the dynamic expansion threshold of the preprocessing node. Step S7: If it is not necessary to reduce the dynamic expansion threshold of the preprocessing node, then determine the minimum guaranteed proportion of index update resources based on the index latency duration of the retrieval requirements.

[0072] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. An information retrieval system, characterized by The information processing module comprises a collection unit configured to collect multi-source information of a plurality of data sources, a preprocessing unit connected to the collection unit and configured to sequentially perform format conversion, cleaning, word segmentation and vocabulary normalization processing on the multi-source information to output a word sequence, and an index construction unit connected to the preprocessing unit and configured to construct a global index according to the word sequence. The retrieval and sorting module is connected to the information processing module and comprises a model training unit configured to train an initial model according to the multi-source information and the word sequence to output a retrieval model, and a retrieval unit connected to the model training unit and configured to query a retrieval requirement to output a retrieval result. The conversion adjustment module is connected to the information processing module and configured to adjust a fault tolerance coefficient of format conversion according to a field matching failure rate of the multi-source information. The capacity expansion adjustment module is connected to the information processing module and the conversion adjustment module and configured to adjust a dynamic expansion threshold of a preprocessing node according to an average waiting time length of a plurality of multi-source information processing. The index adjustment module is connected to the retrieval and sorting module and the capacity expansion adjustment module and configured to adjust a minimum guarantee proportion of index update resources according to an index delay time length of a retrieval requirement. The conversion adjustment module determines that retrieval stability of the multi-source information does not meet requirements in response to the field matching failure rate of the multi-source information being greater than a preset first failure rate.

2. The information retrieval system of claim 1, wherein, The conversion adjustment module preliminarily determines that time effectiveness of the word sequence does not meet requirements in response to the field matching failure rate of the multi-source information being greater than the preset first failure rate and less than or equal to a preset second failure rate, and determines whether the time effectiveness of the word sequence meets requirements according to an average waiting time length of a plurality of multi-source information processing.

3. The information retrieval system of claim 2, wherein, The conversion adjustment module increases the fault tolerance coefficient of format conversion in response to the field matching failure rate of the multi-source information being greater than the preset second failure rate.

4. The information retrieval system of claim 3, wherein, The increase range of the fault tolerance coefficient of format conversion is determined by a difference between the field matching failure rate of the multi-source information and the preset second failure rate. The capacity expansion adjustment module determines that time effectiveness of the word sequence does not meet requirements in response to the average waiting time length of a plurality of multi-source information processing being greater than a preset first waiting time length.

5. The information retrieval system of claim 4, wherein, The capacity expansion adjustment module decreases the dynamic expansion threshold of the preprocessing node in response to the average waiting time length of a plurality of multi-source information processing being greater than the preset first waiting time length and less than or equal to a preset second waiting time length.

6. The information retrieval system of claim 5, wherein, The capacity expansion adjustment module preliminarily determines that update real-time performance of the global index does not meet requirements in response to the average waiting time length of a plurality of multi-source information processing being greater than the preset second waiting time length, and determines whether the update real-time performance of the global index meets requirements according to an index delay time length of a retrieval requirement. The decrease range of the dynamic expansion threshold of the preprocessing node is determined by a difference between the average waiting time length of the plurality of multi-source information processing and the preset first waiting time length.

7. The information retrieval system of claim 6, wherein, ​ 8. The information retrieval system of claim 7, wherein, The index adjustment module determines that the update real-time performance of the global index does not meet the requirement and increases the minimum guarantee proportion of index update resources in response to the index delay duration of the retrieval requirement being greater than the preset delay duration.

9. The information retrieval system of claim 8, wherein, The increase range of the minimum guarantee proportion of the index update resources is determined by a difference between the index delay duration of the retrieval requirement and the preset delay duration.

10. A search method applied to the information search system as claimed in any one of claims 1 to 9, characterized by, The method comprises the following steps: performing format conversion, cleaning, word segmentation and vocabulary normalization processing on a plurality of data sources to output a word sequence, and using a mapping relationship between the word sequence and the plurality of data sources to create a global index; training an initial model using the plurality of data sources and the word sequence to output a retrieval model, and using the retrieval model to perform a query according to a retrieval requirement to output a retrieval result; obtaining a field matching failure rate of the plurality of data sources, and determining whether the retrieval stability of the plurality of data sources meets the requirement according to the field matching failure rate of the plurality of data sources; if the retrieval stability of the plurality of data sources does not meet the requirement, determining whether the fault tolerance coefficient of format conversion needs to be increased; if the fault tolerance coefficient of format conversion does not need to be increased, obtaining an average waiting time of a plurality of pieces of information processing to determine whether the timeliness of the word sequence meets the requirement; if the timeliness of the word sequence does not meet the requirement, determining whether the dynamic expansion threshold of the preprocessing node needs to be reduced; if the dynamic expansion threshold of the preprocessing node does not need to be reduced, determining the minimum guarantee proportion of index update resources based on the index delay duration of the retrieval requirement.

Citation Information

Patent Citations

  • Information search system, information search method, and program

    CN101796493A

  • A picture storage method in a two-dimensional code

    CN109902242A

  • Automatic data processing method and system, computer equipment and readable storage medium

    CN111046035A

  • Resource management method and device, equipment and storage medium

    CN117608823A

  • Quick retrieval method and system for multi-hop relationship of data warehouse based on graph embedded index

    CN120723757A