Information retrieval method and device, equipment and storage medium

By preset data representation and organization processing of the original data, score distribution is generated, and data storage and routing distribution is carried out, the problem of low information retrieval efficiency based on large language models is solved, and more efficient and quality information retrieval effects are achieved.

CN120045748APending Publication Date: 2025-05-27XIAN SECLOVER INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510043026.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When searching information based on large language models, the search efficiency is low, resulting in complex relationship networks, affecting the search effect and efficiency.

Method used

The original data is processed through preset data representation and organizational policy, and score distribution is generated on the preset measurement dimension, and data storage and routing are distributed based on this. Finally, the target data is retrieved through the preset search strategy.

Benefits of technology

The quality and efficiency of information retrieval are improved, and the process of information retrieval is simplified by combining different data storage methods to quickly locate and query at different storage locations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045748A_ABST
    Figure CN120045748A_ABST
Patent Text Reader

Abstract

The invention discloses an information retrieval method and device, equipment and a storage medium, and the method comprises the steps: processing each piece of original data through a preset data representation and organization strategy, generating the score distribution of each piece of original data on a preset measurement dimension, and obtaining the score distribution of each piece of original data on the basis of the score distribution of the original data on the preset measurement dimension; and performing routing distribution on the original data according to a preset data storage mode through a preset data storage and routing strategy, storing the original data obtained by routing distribution to a preset position, and performing retrieval processing from the preset position through a preset retrieval strategy to obtain target data corresponding to the target retrieval intention. According to the scheme, multiple data storage modes and data retrieval strategies are adopted, so that the information retrieval quality and efficiency are improved; in addition, quick positioning query is carried out at different storage positions by combining different data storage modes, so that the retrieval efficiency of information retrieval is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and particularly to an information retrieval method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of large language models (LLMs) and related technologies, information retrieval is no longer limited to using browsers and other methods, but is implemented based on the retrieval and generation method of large language models. This method has greatly improved the retrieval effect and efficiency.

[0003] Currently, when performing information retrieval based on large language models, the final retrieval effect is closely related to the capabilities of the large language models. Usually, information retrieval is implemented in vector form. Specifically, data is first vectorized and stored in a vector database. After retrieving in the vector database through similarity calculation, the original data is then retrieved from the original database, resulting in a relatively complex relationship network.

[0004] Therefore, the above information retrieval method has the problem of low retrieval efficiency. Summary of the Invention

[0005] The present application aims to at least solve the technical problems existing in the prior art. To this end, a first aspect of the present application proposes an information retrieval method, which includes:

[0006] Processing each piece of original data through a preset data representation and organization strategy to generate a score distribution of each piece of original data on a preset measurement dimension; wherein, the preset measurement dimension is used to determine the data representation and organization method of the original data, and the data representation and organization method includes any one of relationship representation, vector representation, graph representation, and hybrid representation;

[0007] Based on the score distribution of the original data on the preset measurement dimension, routing and distributing the original data according to a preset data storage and routing strategy in a preset data storage manner, and storing the routed and distributed original data at a preset location; wherein, the preset data storage manner corresponds to the data representation and organization method, and the preset data storage manner includes any one of relationship storage, vector storage, graph storage, and hybrid storage;

[0008] Performing a retrieval process from the preset location through a preset retrieval strategy to obtain target data corresponding to a target retrieval intention.

[0009] In a possible implementation manner, processing each piece of original data through a preset data representation and organization strategy to generate a score distribution of each piece of original data on a preset measurement dimension includes:

[0010] Perform data standardization and formatting processing on each piece of original data through preset data representation and organizational strategies to generate a processing result;

[0011] Perform data splitting processing on the processing result according to the preset data splitting strategy to generate multiple data units;

[0012] Perform data fusion processing on the multiple data units according to the preset data fusion strategy to generate fusion data; wherein, the preset data fusion strategy corresponds to the preset data splitting strategy, and the semantic correlation degree of the fusion data is higher than that of the original data;

[0013] Perform data representation and organizational processing on the fusion data to generate the score distribution of each piece of fusion data on the preset measurement dimension.

[0014] In a possible implementation manner, performing data representation and organizational processing on the fusion data to generate the score distribution of each piece of fusion data on the preset measurement dimension includes:

[0015] Obtain the preset measurement dimension of the data representation and organizational method; wherein, the preset measurement dimension includes any one of a data business characteristic measurement dimension, a splitting strategy measurement dimension, and a retrieval form measurement dimension;

[0016] Based on the preset measurement dimension, determine the target data representation and organizational method corresponding to the fusion data;

[0017] Calculate the score value of the fusion data on the preset measurement dimension;

[0018] Based on the preset measurement dimension weight and score value corresponding to each preset measurement dimension, calculate the score distribution of each piece of fusion data on the preset measurement dimension.

[0019] In a possible implementation manner, based on the preset measurement dimension, determining the target data representation and organizational method corresponding to the fusion data includes:

[0020] Obtain the dimension index value corresponding to the preset measurement dimension;

[0021] Based on the dimension index value and the preset measurement standard, determine the data representation and organizational method corresponding to the preset measurement standard;

[0022] Determine the data representation and organizational method corresponding to the preset measurement standard as the target data representation and organizational method.

[0023] In a possible implementation manner, based on the score distribution of the original data on the preset measurement dimension, route and distribute the original data according to the preset data storage and routing strategy in the preset data storage manner, including:

[0024] Determine the target dimension index value closest to the score distribution based on the score distribution of the original data on the preset measurement dimension and the index values of each dimension;

[0025] Based on the target dimension index value, determine the target data representation and organization method corresponding to the original data;

[0026] Based on the target data representation and organization method, determine the corresponding preset data storage method, and route and distribute the original data through the preset data storage method.

[0027] In a possible implementation manner, perform a retrieval process from a preset location through a preset retrieval strategy to obtain target data corresponding to the target retrieval intention, including:

[0028] Perform a matching process from a preset location through a preset retrieval strategy, and obtain the target retrieval intention in the case of successful matching;

[0029] Determine the preset data storage method corresponding to the target retrieval intention, and perform data retrieval based on the preset data storage method to obtain the target data.

[0030] In a possible implementation manner, the method further includes:

[0031] Make a judgment through a preset retrieval ability evaluation strategy to obtain a judgment result; wherein, the preset retrieval ability evaluation strategy is determined based on multiple evaluation indicators, and the evaluation indicators at least include quality indicators, efficiency indicators, usage frequency and effect indicators of the preset retrieval strategy.

[0032] A second aspect of the present application proposes an information retrieval device, and the device includes:

[0033] A generation module, configured to process each piece of original data through a preset data representation and organization strategy to generate a score distribution of each piece of original data on a preset measurement dimension; wherein, the preset measurement dimension is used to determine the data representation and organization method of the original data, and the data representation and organization method includes any one of relationship representation, vector representation, graph representation, and hybrid representation;

[0034] A storage module, configured to route and distribute the original data according to a preset data storage method based on the score distribution of the original data on the preset measurement dimension through a preset data storage and routing strategy, and store the routed and distributed original data at a preset location; wherein, the preset data storage method corresponds to the data representation and organization method, and the preset data storage method includes any one of relationship storage, vector storage, graph storage, and hybrid storage;

[0035] A retrieval module, configured to perform a retrieval process from a preset location through a preset retrieval strategy to obtain target data corresponding to the target retrieval intention.

[0036] In a possible implementation, the above-mentioned generation module is specifically configured to:

[0037] Perform data standardization and formatting processing on each piece of original data through preset data representation and organization policies to generate a processing result;

[0038] Perform data splitting processing on the processing result according to a preset data splitting strategy to generate multiple data units;

[0039] Perform data fusion processing on the multiple data units according to a preset data fusion strategy to generate fusion data; wherein, the preset data fusion strategy corresponds to the preset data splitting strategy, and the semantic correlation degree of the fusion data is higher than that of the original data;

[0040] Perform data representation and organization processing on the fusion data to generate the score distribution of each piece of fusion data on a preset measurement dimension.

[0041] In a possible implementation, the above-mentioned generation module is further configured to:

[0042] Obtain a preset measurement dimension of the data representation and organization method; wherein, the preset measurement dimension includes any one of a data business feature measurement dimension, a splitting strategy measurement dimension, and a retrieval form measurement dimension;

[0043] Based on the preset measurement dimension, determine the target data representation and organization method corresponding to the fusion data;

[0044] Calculate the score value of the fusion data on the preset measurement dimension;

[0045] Based on the preset measurement dimension weight and score value corresponding to each preset measurement dimension, calculate the score distribution of each piece of fusion data on the preset measurement dimension.

[0046] In a possible implementation, the above-mentioned generation module is further configured to:

[0047] Obtain the dimension index value corresponding to the preset measurement dimension;

[0048] Based on the dimension index value and the preset measurement standard, determine the data representation and organization method corresponding to the preset measurement standard;

[0049] Determine the data representation and organization method corresponding to the preset measurement standard as the target data representation and organization method.

[0050] In a possible implementation, the above-mentioned storage module is specifically configured to:

[0051] Based on the score distribution of the original data on the preset measurement dimension and each dimension index value, determine the target dimension index value closest to the score distribution;

[0052] Determine the target data representation and organization method corresponding to the original data based on the target dimension index value;

[0053] Determine the corresponding preset data storage method based on the target data representation and organization method, and route and distribute the original data through the preset data storage method.

[0054] In a possible implementation manner, the above-mentioned retrieval module is specifically used for:

[0055] Perform a matching process from a preset location through a preset retrieval strategy, and obtain the target retrieval intention in the case of successful matching;

[0056] Determine the preset data storage method corresponding to the target retrieval intention, and perform data retrieval based on the preset data storage method to obtain the target data.

[0057] In a possible implementation manner, the above-mentioned information retrieval device is further used for:

[0058] Make a judgment through a preset retrieval ability evaluation strategy to obtain a judgment result; wherein, the preset retrieval ability evaluation strategy is determined based on multiple evaluation indexes, and the evaluation indexes at least include a quality index, an efficiency index, the usage frequency of the preset retrieval strategy, and an effect index.

[0059] A third aspect of this application proposes an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the information retrieval method as described in the first aspect.

[0060] A fourth aspect of this application proposes a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the information retrieval method as described in the first aspect.

[0061] The embodiments of this application have the following beneficial effects:

[0062] The information retrieval method provided by the embodiments of the present application includes: processing each piece of original data through a preset data representation and organization strategy to generate a score distribution of each piece of original data in a preset measurement dimension; based on the score distribution of the original data in the preset measurement dimension, routing and distributing the original data according to a preset data storage method through a preset data storage and routing strategy, and storing the routed and distributed original data at a preset location; and performing a retrieval process from the preset location through a preset retrieval strategy to obtain target data corresponding to a target retrieval intention. This solution improves the quality and efficiency of information retrieval by adopting various data storage methods and data retrieval strategies; in addition, by combining different data storage methods for quick positioning and querying at different storage locations, the retrieval efficiency of information retrieval is improved. Description of the Drawings

[0063] Figure 1 It is a block diagram of a computer device provided by the embodiments of the present application;

[0064] Figure 2 It is a flowchart of the steps of an information retrieval method provided by the embodiments of the present application;

[0065] Figure 3 It is a flowchart of the steps of generating a score distribution provided by the embodiments of the present application;

[0066] Figure 4 It is another flowchart of the steps of generating a score distribution provided by the embodiments of the present application;

[0067] Figure 5 It is a flowchart of the steps of determining a target data representation and organization method provided by the embodiments of the present application;

[0068] Figure 6 It is a flowchart of the steps of routing and distributing the original data provided by the embodiments of the present application;

[0069] Figure 7 It is a flowchart of the steps of obtaining target data provided by the embodiments of the present application;

[0070] Figure 8 It is a block diagram of the structure of an information retrieval device provided by the embodiments of the present application. Detailed Embodiments

[0071] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0072] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise specified, the meaning of "a plurality" is two or more. Additionally, the use of "based on" or "according to" implies openness and inclusiveness, because a process, step, calculation, or other action "based on" or "according to" one or more of the stated conditions or values may, in practice, be based on additional conditions or values beyond those stated.

[0073] The information retrieval method provided in this application can be applied to a computer device (electronic device). The computer device can be a server or a terminal. Among them, the server can be a single server or a server cluster composed of multiple servers. The embodiments of this application do not make specific limitations in this regard. The terminal can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers, and portable wearable devices.

[0074] Taking the computer device being a server as an example, Figure 1 A block diagram of a server is shown, as Figure 1 shown, the server may include a processor and a memory connected by a system bus. Among them, the processor of the server is used to provide computing and control capabilities. The memory of the server includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it realizes an information retrieval method.

[0075] Those skilled in the art can understand that Figure 1 the structure shown in

[0076] is only a block diagram of a part of the structure related to the solution of this application and does not constitute a limitation on the server to which the solution of this application is applied. Optionally, the server may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.

[0077] Figure 2 is a flowchart of the steps of an information retrieval method provided in an embodiment of this application. As Figure 2 shown, the method includes the following steps:

[0078] Step 202: Process each piece of original data through a preset data representation and organization strategy to generate the score distribution of each piece of original data on a preset measurement dimension.

[0079] Among them, the preset measurement dimension is used to determine the data representation and organization method of the original data. The data representation and organization method includes any one of relationship representation, vector representation, graph representation, and hybrid representation, and may also include other types of data representation and organization methods. The embodiments of the present application do not make specific limitations in this regard. In some optional embodiments, as Figure 3 shown, Figure 3 is a step flowchart for generating a score distribution provided by an embodiment of the present application, including:

[0080] Step 302: Perform data standardization and formatting processing on each piece of original data through a preset data representation and organization strategy to generate a processing result.

[0081] Step 304: Perform data splitting processing on the processing result according to a preset data splitting strategy to generate multiple data units.

[0082] Step 306: Perform data fusion processing on the multiple data units according to a preset data fusion strategy to generate fusion data.

[0083] Step 308: Perform data representation and organization processing on the fusion data to generate the score distribution of each piece of fusion data on a preset measurement dimension.

[0084] Among them, after obtaining the original data, data standardization and formatting processing can be first performed on each piece of original data through a preset data representation and organization strategy to generate a processing result. The specific process of specific standardization and formatting can refer to the prior art and will not be elaborated here.

[0085] Then, data splitting processing can be performed on the processing result according to a preset data splitting strategy to generate multiple data units. The data unit is the smallest unit. The preset data splitting strategy can include various types. Optionally, it can include strategies such as splitting by word segmentation, splitting by single-sentence semantics, and splitting according to part of speech.

[0086] Furthermore, corresponding preset data fusion strategies are respectively formulated for different preset data splitting strategies, so that multiple data units can be subjected to data fusion processing according to the preset data fusion strategy to generate fusion data. Among them, the preset data fusion strategy corresponds to the preset data splitting strategy, and the semantic correlation degree of the fusion data is higher than that of the original data. Since the storage efficiency and compression rate of the split data units are high, but the semantic and spatial retrieval effects are poor, it is necessary to formulate a preset data fusion strategy for different preset data splitting strategies, and the accuracy of information retrieval through the fusion data is higher.

[0087] Exemplarily, for the case where the preset data splitting strategy is splitting according to semantics, a preset data fusion strategy for fusing strongly associated semantic statements can be adopted for data fusion processing. For example, after splitting according to semantics, there may be two statements: Statement 1: Electric vehicles have much less environmental pollution than gasoline vehicles, and environmental pollution is a topic of widespread concern in the international community; Statement 2: Many countries and organizations advocate vigorously developing new energy vehicles. If the smallest unit is directly used for subsequent information retrieval, problems such as information loss and low recall rate may occur. Therefore, a preset data fusion strategy for fusing strongly associated semantic statements can be adopted for data fusion processing, so that the accuracy of information retrieval through fused data is higher.

[0088] Finally, data representation and organization processing can be performed on the fused data to generate the score distribution of each piece of fused data on the preset measurement dimension. In some optional embodiments, as Figure 4 shown, Figure 4 Another step flowchart for generating a score distribution provided by an embodiment of the present application includes:

[0089] Step 402, obtain the preset measurement dimension of the data representation and organization method.

[0090] Step 404, based on the preset measurement dimension, determine the target data representation and organization method corresponding to the fused data.

[0091] Step 406, calculate the score value of the fused data on the preset measurement dimension.

[0092] Step 408, based on the preset measurement dimension weights and score values corresponding to each preset measurement dimension, calculate the score distribution of each piece of fused data on the preset measurement dimension.

[0093] Among them, the preset measurement dimension includes any one of the data service feature measurement dimension, the splitting strategy measurement dimension, and the retrieval form measurement dimension, and can also include other types of preset measurement dimensions. The data service feature measurement dimension can include multiple types, for example, data volume, data timeliness, data diversity, data credibility, normativity, compliance, accessibility, sensitivity, security level, etc. The splitting strategy measurement dimension can include various types of the above-mentioned preset data splitting strategies and preset data fusion strategies. The retrieval form measurement dimension can include relational retrieval, semantic retrieval, aggregation retrieval, statistical retrieval, etc.

[0094] Based on the preset measurement dimension again, determine the target data representation and organization method corresponding to the fused data. In some optional embodiments, as Figure 5 shown, Figure 5 A step flowchart for determining the target data representation and organization method provided by an embodiment of the present application includes:

[0095] Step 502: Obtain the dimension index values corresponding to the preset measurement dimensions.

[0096] Step 504: Based on the dimension index values and the preset measurement criteria, determine the data representation and organization method corresponding to the preset measurement criteria.

[0097] Step 506: Determine the data representation and organization method corresponding to the preset measurement criteria as the target data representation and organization method.

[0098] Among them, the dimension index values corresponding to the preset measurement dimensions are the specific values corresponding to each preset measurement dimension. Thus, based on the dimension index values and the preset measurement criteria, the data representation and organization method corresponding to the preset measurement criteria can be determined. In other words, it can be judged whether the dimension index values fall within the index range of a certain preset measurement criteria, so that the data representation and organization method corresponding to the preset measurement criteria can be determined, and the data representation and organization method corresponding to the preset measurement criteria is determined as the target data representation and organization method.

[0099] Next, the score value of the fusion data on the preset measurement dimension can be calculated. Optionally, traditional expert scoring, rule matching, semantic automatic evaluation by large language models, etc. can be used to calculate the score value.

[0100] Then, based on the preset measurement dimension weights and score values corresponding to each preset measurement dimension, the score value distribution of each piece of fusion data on the preset measurement dimension is calculated. Among them, the preset measurement dimension weights can be determined by experts for different business scenarios, and the product of the preset measurement dimension weights and score values corresponding to each preset measurement dimension can be used to obtain the score values on each preset measurement dimension one by one, and the multiple score values obtained form the score value distribution of each piece of fusion data on the preset measurement dimension.

[0101] Step 204: Based on the score value distribution of the original data on the preset measurement dimension, route and distribute the original data according to the preset data storage and routing strategy in the preset data storage method, and store the original data obtained by route and distribution at the preset location.

[0102] Among them, the preset data storage method corresponds to the data representation and organization method. The preset data storage method includes any one of relational storage, vector storage, graph storage, and hybrid storage, and can also include other types of data storage methods. In some optional embodiments, as Figure 6 shown Figure 6 is a flowchart of the steps for routing and distributing the original data provided by the embodiment of the present application, including:

[0103] Step 602: Determine the target dimension index value that is closest to the score distribution based on the score distribution of the original data on the preset measurement dimension and the index values of each dimension.

[0104] Step 604: Determine the target data representation and organization method corresponding to the original data based on the target dimension index value.

[0105] Step 606: Determine the corresponding preset data storage method based on the target data representation and organization method, and route and distribute the original data through the preset data storage method.

[0106] Among them, a similarity algorithm can be used to calculate the distance value between the score distribution on the preset measurement dimension and the index values of each dimension, so as to determine the target dimension index value that is closest to the score distribution according to the distance value. Other methods can also be used to determine the target dimension index value, and the embodiments of the present application do not make specific limitations on this.

[0107] Then, based on the target dimension index value, determine the target data representation and organization method corresponding to the original data. Finally, determine the corresponding preset data storage method based on the target data representation and organization method, route and distribute the original data through the preset data storage method, and store the routed and distributed original data at a preset location.

[0108] Among them, the preset location can be a database. Exemplarily, the original data with a storage method of relational storage can be stored in a relational database, the original data with a storage method of vector storage can be stored in a vector database, the original data with a storage method of graph storage can be stored in a graph database, and the original data with a storage method of hybrid storage can be stored in a hybrid database.

[0109] Step 206: Perform a retrieval process from the preset location through a preset retrieval strategy to obtain target data corresponding to the target retrieval intention.

[0110] Among them, in some optional embodiments, as Figure 7 shown, Figure 7 is a flowchart of the steps for obtaining target data provided by the embodiments of the present application, including:

[0111] Step 702: Perform a matching process from the preset location through a preset retrieval strategy, and obtain the target retrieval intention in the case of successful matching.

[0112] Step 704: Determine the preset data storage method corresponding to the target retrieval intention, and perform data retrieval based on the preset data storage method to obtain the target data.

[0113] Among them, before performing matching processing from a preset position through a preset retrieval strategy, it is necessary to formulate a caching strategy first. Specifically, the hot information criteria can be set first, that is, the hot information range, expiration time, etc. are determined, and the hot information is determined for different retrieval scenarios. Then the hot information is collected, and different hot information matching methods are flexibly selected, such as various algorithms like semantic similarity. Finally, for a certain piece of information in information retrieval, the hot information matching method is used to match with the hot information. If the threshold is reached, the matching is successful; otherwise, the matching is unsuccessful. Among them, the threshold needs to be set in advance and is generally related to the accuracy and efficiency requirements of the information retrieval task.

[0114] In addition, a buffering strategy also needs to be formulated. Specifically, the buffering data range and mechanism can be set, that is, the buffering data range can be set for the retrieval task, the original data set, performance requirements, etc., and its expiration time, update mechanism, etc. are set. Then the buffering information is collected, and different buffering information matching methods are flexibly selected, such as various algorithms like semantic similarity. Finally, for a certain piece of information in information retrieval, the buffering information matching method is used to match with the buffering information. If the threshold is reached, the matching is successful; otherwise, the matching is unsuccessful. Among them, the threshold needs to be set in advance and is generally related to the accuracy and efficiency requirements of the information retrieval task.

[0115] Thus, when performing matching processing from a preset position through a preset retrieval strategy, caching matching and buffering matching can be carried out in sequence. If both match successfully, the target retrieval intention is obtained in the case of successful matching. By setting optimization mechanisms such as caching strategies and buffering strategies, the same or similar problems can obtain the target data without recalculation or with little calculation through the caching strategy, effectively saving system resources.

[0116] Then, determine the preset data storage method corresponding to the target retrieval intention, and perform data retrieval based on the preset data storage method to obtain the target data. Specifically, when obtaining the target retrieval intention, methods such as large language models and rule matching can be used for identification. The target retrieval intention can also include various types such as association relation retrieval, semantic retrieval, aggregation retrieval, and statistical retrieval. Finally, data retrieval can be performed based on the preset data storage method to obtain the target data.

[0117] In some optional embodiments, it can also be judged through a preset retrieval ability evaluation strategy to obtain a judgment result. Among them, the preset retrieval ability evaluation strategy is determined based on multiple evaluation indicators, and the evaluation indicators at least include quality indicators, efficiency indicators, the usage frequency and effect indicators of the preset retrieval strategy. Various hyperparameters can be improved in real time according to the judgment result of the preset retrieval ability evaluation strategy to optimize the effect of information retrieval.

[0118] The present application provides an information retrieval method, which includes: processing each piece of original data through a preset data representation and organization strategy to generate a score distribution of each piece of original data on a preset measurement dimension; based on the score distribution of the original data on the preset measurement dimension, routing and distributing the original data according to a preset data storage and routing strategy in a preset data storage manner, and storing the routed and distributed original data at a preset location; performing a retrieval process from the preset location through a preset retrieval strategy to obtain target data corresponding to a target retrieval intention. This solution improves the quality and efficiency of information retrieval by adopting various data storage methods and data retrieval strategies; in addition, by combining different data storage methods for quick positioning and querying at different storage locations, the retrieval efficiency of information retrieval is improved.

[0119] Figure 8 It is a structural block diagram of an information retrieval device provided by an embodiment of the present application.

[0120] As Figure 8 shown, the information retrieval device 800 includes:

[0121] A generation module 802, configured to process each piece of original data through a preset data representation and organization strategy to generate a score distribution of each piece of original data on a preset measurement dimension; wherein, the preset measurement dimension is used to determine the data representation and organization method of the original data, and the data representation and organization method includes any one of relationship representation, vector representation, graph representation, and hybrid representation.

[0122] A storage module 804, configured to route and distribute the original data according to a preset data storage and routing strategy in a preset data storage manner based on the score distribution of the original data on the preset measurement dimension, and store the routed and distributed original data at a preset location; wherein, the preset data storage manner corresponds to the data representation and organization method, and the preset data storage manner includes any one of relationship storage, vector storage, graph storage, and hybrid storage.

[0123] A retrieval module 806, configured to perform a retrieval process from the preset location through a preset retrieval strategy to obtain target data corresponding to a target retrieval intention.

[0124] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein. Each module in the above information retrieval device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in a computer device in the form of software, so as to facilitate the processor to call and execute the operations of the above modules.

[0125] In one embodiment of the present application, a computer device is provided. The computer device includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0126] Process each piece of original data through a preset data representation and organization strategy to generate a score distribution of each piece of original data on a preset measurement dimension; wherein, the preset measurement dimension is used to determine the data representation and organization method of the original data, and the data representation and organization method includes any one of relationship representation, vector representation, graph representation, and hybrid representation;

[0127] Based on the score distribution of the original data on the preset measurement dimension, route and distribute the original data according to a preset data storage and routing strategy in a preset data storage method, and store the original data obtained by the route and distribution at a preset location; wherein, the preset data storage method corresponds to the data representation and organization method, and the preset data storage method includes any one of relationship storage, vector storage, graph storage, and hybrid storage;

[0128] Perform a retrieval process from the preset location through a preset retrieval strategy to obtain target data corresponding to a target retrieval intention.

[0129] In one embodiment of the present application, when the processor executes the computer program, the following steps are also implemented:

[0130] Perform data standardization and formatting processing on each piece of original data through a preset data representation and organization strategy to generate a processing result;

[0131] Perform data splitting processing on the processing result according to a preset data splitting strategy to generate multiple data units;

[0132] Perform data fusion processing on the multiple data units according to a preset data fusion strategy to generate fusion data; wherein, the preset data fusion strategy corresponds to the preset data splitting strategy, and the semantic correlation degree of the fusion data is higher than that of the original data;

[0133] Perform data representation and organization processing on the fusion data to generate a score distribution of each piece of fusion data on a preset measurement dimension.

[0134] In one embodiment of the present application, when the processor executes the computer program, the following steps are also implemented:

[0135] Obtain the preset measurement dimension of the data representation and organization method; wherein, the preset measurement dimension includes any one of a data service feature measurement dimension, a splitting strategy measurement dimension, and a retrieval form measurement dimension;

[0136] Based on the preset measurement dimension, determine the target data representation and organization method corresponding to the fusion data;

[0137] Calculate the score value of the fusion data on a preset measurement dimension;

[0138] Based on the preset measurement dimension weights and score values corresponding to each preset measurement dimension, calculate the score distribution of each piece of fusion data on the preset measurement dimension.

[0139] In an embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:

[0140] Obtain the dimension index value corresponding to the preset measurement dimension;

[0141] Based on the dimension index value and the preset measurement standard, determine the data representation and organization method corresponding to the preset measurement standard;

[0142] Determine the data representation and organization method corresponding to the preset measurement standard as the target data representation and organization method.

[0143] In an embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:

[0144] Based on the score distribution of the original data on the preset measurement dimension and each dimension index value, determine the target dimension index value closest to the score distribution;

[0145] Based on the target dimension index value, determine the target data representation and organization method corresponding to the original data;

[0146] Based on the target data representation and organization method, determine the corresponding preset data storage method, and route and distribute the original data through the preset data storage method.

[0147] In an embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:

[0148] Perform matching processing from a preset location through a preset retrieval strategy, and obtain the target retrieval intention in the case of successful matching;

[0149] Determine the preset data storage method corresponding to the target retrieval intention, and perform data retrieval based on the preset data storage method to obtain the target data.

[0150] In an embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:

[0151] Make a judgment through a preset retrieval ability evaluation strategy to obtain a judgment result; wherein, the preset retrieval ability evaluation strategy is determined based on multiple evaluation indicators, and the evaluation indicators at least include quality indicators, efficiency indicators, the usage frequency of the preset retrieval strategy, and effect indicators.

[0152] The computer device provided by the embodiment of the present application has a similar implementation principle and technical effect to the above method embodiment, which will not be elaborated here.

[0153] In an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0154] Process each piece of original data through a preset data representation and organization strategy to generate a score distribution of each piece of original data on a preset measurement dimension; wherein, the preset measurement dimension is used to determine the data representation and organization method of the original data, and the data representation and organization method includes any one of relationship representation, vector representation, graph representation, and hybrid representation;

[0155] Based on the score distribution of the original data on the preset measurement dimension, route and distribute the original data according to the preset data storage and routing strategy in the preset data storage manner, and store the original data obtained by the route distribution at a preset position; wherein, the preset data storage manner corresponds to the data representation and organization method, and the preset data storage manner includes any one of relationship storage, vector storage, graph storage, and hybrid storage;

[0156] Perform a retrieval process from the preset position through a preset retrieval strategy to obtain target data corresponding to the target retrieval intention.

[0157] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are also implemented:

[0158] Perform data standardization and formatting processing on each piece of original data through a preset data representation and organization strategy to generate a processing result;

[0159] Perform data splitting processing on the processing result according to a preset data splitting strategy to generate multiple data units;

[0160] Perform data fusion processing on the multiple data units according to a preset data fusion strategy to generate fusion data; wherein, the preset data fusion strategy corresponds to the preset data splitting strategy, and the semantic correlation degree of the fusion data is higher than that of the original data;

[0161] Perform data representation and organization processing on the fusion data to generate a score distribution of each piece of fusion data on a preset measurement dimension.

[0162] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are also implemented:

[0163] Obtain the preset measurement dimension of the data representation and organization method; wherein, the preset measurement dimension includes any one of a data service feature measurement dimension, a splitting strategy measurement dimension, and a retrieval form measurement dimension;

[0164] Based on a preset measurement dimension, determine the target data representation and organization method corresponding to the fusion data;

[0165] Calculate the score value of the fusion data on the preset measurement dimension;

[0166] Based on the preset measurement dimension weights and score values corresponding to each preset measurement dimension, calculate the score value distribution of each piece of fusion data on the preset measurement dimension.

[0167] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:

[0168] Obtain the dimension index value corresponding to the preset measurement dimension;

[0169] Based on the dimension index value and the preset measurement standard, determine the data representation and organization method corresponding to the preset measurement standard;

[0170] Determine the data representation and organization method corresponding to the preset measurement standard as the target data representation and organization method.

[0171] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:

[0172] Based on the score value distribution of the original data on the preset measurement dimension and each dimension index value, determine the target dimension index value closest to the score value distribution;

[0173] Based on the target dimension index value, determine the target data representation and organization method corresponding to the original data;

[0174] Based on the target data representation and organization method, determine the corresponding preset data storage method, and route and distribute the original data through the preset data storage method.

[0175] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:

[0176] Perform a matching process from a preset location through a preset retrieval strategy, and obtain the target retrieval intention in the case of successful matching;

[0177] Determine the preset data storage method corresponding to the target retrieval intention, and perform data retrieval based on the preset data storage method to obtain the target data.

[0178] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:

[0179] Judgment is made through a preset retrieval ability evaluation strategy to obtain a judgment result; among them, the preset retrieval ability evaluation strategy is determined based on multiple evaluation indicators, and the evaluation indicators at least include quality indicators, efficiency indicators, the usage frequency of the preset retrieval strategy, and effect indicators.

[0180] The computer-readable storage medium provided in this embodiment has the same implementation principle and technical effects as the above method embodiment, and will not be elaborated here.

[0181] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0182] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0183] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An information retrieval method, characterized in that: The method comprises: Processing each piece of raw data through a preset data representation and organization strategy to generate a score distribution of each piece of raw data on a preset measurement dimension; wherein the preset measurement dimension is used to determine a data representation and organization method of the raw data, and the data representation and organization method includes any one of a relational representation, a vector representation, a graph representation, and a hybrid representation; Based on the score distribution of the original data on the preset measurement dimension, the original data is routed and distributed according to the preset data storage method through the preset data storage and routing strategy, and the original data obtained by routing distribution is stored in a preset location; wherein the preset data storage method corresponds to the data representation and organization method, and the preset data storage method includes any one of relational storage, vector storage, graph storage, and hybrid storage; The preset search strategy is used to perform search processing from the preset position to obtain target data corresponding to the target search intention.

2. The method according to claim 1, characterized in that The processing of each piece of original data by using a preset data representation and organization strategy to generate a score distribution of each piece of original data on a preset measurement dimension includes: Perform data standardization and formatting on each piece of raw data through preset data representation and organization strategies to generate processing results; Performing data splitting processing on the processing result according to a preset data splitting strategy to generate multiple data units; The plurality of data units are subjected to data fusion processing according to a preset data fusion strategy to generate fused data; wherein the preset data fusion strategy corresponds to the preset data splitting strategy, and the semantic relevance of the fused data is higher than the semantic relevance of the original data; The fused data is represented and organized to generate a score distribution of each piece of the fused data on a preset measurement dimension.

3. The method according to claim 2, characterized in that The performing data representation and organization processing on the fused data to generate a score distribution of each piece of the fused data on a preset measurement dimension includes: Obtaining preset measurement dimensions of data representation and organization; wherein the preset measurement dimensions include any one of data business feature measurement dimensions, splitting strategy measurement dimensions, and retrieval form measurement dimensions; Based on the preset measurement dimension, determining a target data representation and organization method corresponding to the fused data; Calculating the score value of the fused data on the preset measurement dimension; Based on the preset measurement dimension weights and the score values ​​corresponding to each of the preset measurement dimensions, the score distribution of each of the fused data on the preset measurement dimensions is calculated.

4. The method according to claim 3, characterized in that The determining, based on the preset measurement dimension, a representation and organization method of target data corresponding to the fused data includes: Obtaining the dimension indicator value corresponding to the preset measurement dimension; Based on the dimension indicator value and the preset measurement standard, determine the data representation and organization method corresponding to the preset measurement standard; The data representation and organization method corresponding to the preset measurement standard is determined as the target data representation and organization method.

5. The method according to claim 4, characterized in that The routing and distributing of the original data based on the score distribution of the original data on the preset measurement dimension according to the preset data storage and routing strategy in accordance with the preset data storage method includes: Based on the score distribution of the original data on the preset measurement dimension and the indicator value of each dimension, determine the target dimension indicator value that is closest to the score distribution; Based on the target dimension indicator value, determine the target data representation and organization method corresponding to the original data; A corresponding preset data storage method is determined based on the target data representation and organization method, and the original data is routed and distributed through the preset data storage method.

6. The method according to any one of claims 1 to 4, characterized in that: The retrieval process is performed from the preset position by using the preset retrieval strategy to obtain target data corresponding to the target retrieval intention, including: Perform matching processing from the preset position through a preset search strategy, and obtain the target search intent if the match is successful; Determine a preset data storage method corresponding to the target retrieval intention, and perform data retrieval based on the preset data storage method to obtain the target data.

7. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The judgment result is obtained by judging through a preset search capability judgment strategy; wherein the preset search capability judgment strategy is determined based on multiple judgment indicators, and the judgment indicators at least include quality indicators, efficiency indicators, and the frequency of use and effect indicators of the preset search strategy.

8. An information retrieval device, characterized in that: The device comprises: A generation module, used to process each piece of raw data through a preset data representation and organization strategy, and generate a score distribution of each piece of raw data on a preset measurement dimension; wherein the preset measurement dimension is used to determine the data representation and organization method of the raw data, and the data representation and organization method includes any one of a relational representation, a vector representation, a graph representation, and a hybrid representation; A storage module, configured to distribute the raw data by routing according to a preset data storage method through a preset data storage and routing strategy based on the score distribution of the raw data on a preset measurement dimension, and store the raw data obtained by routing distribution in a preset location; wherein the preset data storage method corresponds to the data representation and organization method, and the preset data storage method includes any one of relational storage, vector storage, graph storage, and hybrid storage; The retrieval module is used to perform retrieval processing from the preset position through a preset retrieval strategy to obtain target data corresponding to the target retrieval intention.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the information retrieval method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the information retrieval method as described in any one of claims 1-7.