Data processing method and device, storage medium and electronic equipment
By combining the similarity and sorting position of the search data with the target description text, the data processing of the large language model LLM is optimized using the reciprocal ranking fusion algorithm, which solves the problem of data inaccuracy in the prior art and achieves higher quality generation results.
Patent Information
- Application Number
- CN202510369145.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, in the large language model LLM, when generating RAGs through search enhancement, the similarity and sorting position of the search data and the target description text cannot be effectively combined, resulting in the data provided to the intelligent model being inaccurate enough, affecting the quality of the generated results.
By combining the data similarity and sorting position between the search data and the target description text, the reciprocal ranking fusion algorithm is used to score and sort the search data to provide more accurate target data to the intelligent model.
It improves the accuracy of the output content of the intelligent model, improves user satisfaction and the integration effect of multi-data source retrieval.
Smart Images

Figure CN120256752A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a data processing method, apparatus, storage medium, and electronic device. Background Art
[0002] In the process of using the large language model LLM (Large Language Model), through retrieval-augmented generation RAG (Retrieval-augmented Generation), multiple retrieval results are retrieved from multiple data sources respectively based on the text input by the user, and then provided to the LLM after integration. The LLM generates text or images for the user based on these retrieval results. Summary of the Invention
[0003] In view of this, this application provides a data processing method, apparatus, storage medium, and electronic device, as follows:
[0004] A data processing method includes:
[0005] Obtaining a retrieval result for a target description text, where the retrieval result at least includes retrieval subsets corresponding to multiple data sources, and the retrieval subset includes retrieval data retrieved from the corresponding data source;
[0006] Obtaining target data according to the data similarity between the retrieval data and the target description text and the sorting position of the retrieval data in the corresponding retrieval subset;
[0007] Providing the target data to an intelligent model, and the target data enables the intelligent model to output target content based on the target description text.
[0008] In the above method, preferably, the method further includes:
[0009] Extracting the data similarity between each retrieval data in each retrieval subset and the target description text from the retrieval result;
[0010] Wherein, the retrieval data in the retrieval subset is obtained based on the data similarity, and the sorting position between the retrieval data in the retrieval subset is determined based on the data similarity.
[0011] In the above method, preferably, obtaining target data according to the data similarity between the retrieval data and the target description text and the sorting position of the retrieval data in the corresponding retrieval subset includes:
[0012] Obtain the sorting score corresponding to the retrieved data according to the sorting position of the retrieved data in the retrieved subset to which it belongs and the data similarity between the retrieved data and the target description text;
[0013] Obtain the target data according to the sorting score.
[0014] Preferably, for the above method, obtaining the target data according to the sorting score includes:
[0015] Sort all the retrieved data according to the sorting score corresponding to each retrieved data to obtain a sorting result;
[0016] Obtain the target data according to the sorting result.
[0017] Preferably, for the above method, obtaining the sorting score corresponding to the retrieved data according to the sorting position of the retrieved data in the retrieved subset to which it belongs and the data similarity between the retrieved data and the target description text includes:
[0018] Obtain the initial score corresponding to the retrieved data according to the sorting position of the retrieved data in the retrieved subset to which it belongs;
[0019] Adjust the initial score based on the data similarity between the retrieved data and the target description text to obtain the sorting score corresponding to the retrieved data.
[0020] Preferably, for the above method, obtaining the sorting score corresponding to the retrieved data according to the sorting position of the retrieved data in the retrieved subset to which it belongs and the data similarity between the retrieved data and the target description text includes:
[0021] Process the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the retrieved subset to which it belongs based on the inverse ranking fusion algorithm to obtain the sorting score corresponding to the retrieved data.
[0022] Preferably, in the inverse ranking fusion algorithm, use the data similarity between the retrieved data and the target description text as the numerator, and use the sum of the smoothing constant and the sorting position of the retrieved data in the retrieved subset to which it belongs as the denominator.
[0023] A data processing device includes:
[0024] A retrieval obtaining unit, configured to obtain a retrieval result for a target description text, where the retrieval result at least includes retrieved subsets corresponding to multiple data sources, and the retrieved subset includes retrieved data retrieved from the data source to which it belongs;
[0025] A target acquisition unit, configured to acquire target data according to the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the retrieved subset to which it belongs;
[0026] A data providing unit, configured to provide the target data to an intelligent model, and the target data enables the intelligent model to output target content based on the target description text.
[0027] A storage medium, configured to store a computer program, and when the computer program is executed by a processor, the following is implemented:
[0028] Obtain a retrieval result for a target description text, where the retrieval result at least includes retrieved subsets corresponding to multiple data sources, and the retrieved subset includes retrieved data retrieved from the data source to which it belongs;
[0029] According to the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the retrieved subset to which it belongs, obtain target data;
[0030] Provide the target data to an intelligent model, and the target data enables the intelligent model to output target content based on the target description text.
[0031] An electronic device, comprising:
[0032] A memory, configured to store a computer program and data generated by running the computer program;
[0033] A processor, configured to execute the computer program to implement: obtain a retrieval result for a target description text, where the retrieval result at least includes retrieved subsets corresponding to multiple data sources, and the retrieved subset includes retrieved data retrieved from the data source to which it belongs; according to the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the retrieved subset to which it belongs, obtain target data; provide the target data to an intelligent model, and the target data enables the intelligent model to output target content based on the target description text.
[0034] A storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the data processing method described in any one of the above is implemented.
[0035] A computer program product, comprising computer program / instructions, and when the computer program / instructions are executed by a processor, the data processing method described in any one of the above is implemented. Description of the Drawings
[0036] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0037] Figure 1 It is a flowchart of the implementation of a data processing method provided by an embodiment of the present application;
[0038] Figure 2 It is an example diagram of the retrieved data of RAG;
[0039] Figure 3 It is a partial flowchart of a data processing method provided by an embodiment of the present application;
[0040] Figure 4 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0041] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0043] Reference Figure 1 As shown, it is a flowchart of the implementation of a data processing method provided by an embodiment of the present application. This method can be applied to an electronic device capable of data processing, and the electronic device can be a computer, a server, etc. The technical solution in this embodiment is mainly used to improve the accuracy of the target data provided to the intelligent model.
[0044] Specifically, the method in this embodiment may include the following steps:
[0045] Step 101: Obtain a retrieval result for the target description text.
[0046] Among them, the retrieval result includes at least retrieval subsets corresponding to multiple data sources. The retrieval subset includes the retrieved data from the corresponding data source.
[0047] For example, as Figure 2As shown in , in this embodiment, a search tool based on retrieval-augmented generation (RAG) can be used to search for text data matching a target description text such as "Where is the best place to go in one day?" in multiple data sources. Multiple search data can be retrieved from each data source, such as "A scenic spot" and "B scenic spot", etc. The search data belonging to the same data source constitute a search subset, and the search subsets corresponding to each data source constitute the search results. In each search subset, each search data has a ranking position. The higher the ranking position, the more the search data matches the target description text relative to other search data in the search subset.
[0048] It should be noted that in the search sub-sets corresponding to different data sources, there may be one or more search data that are the same. That is, a search data may be retrieved from different data sources. Based on this, a search data may appear in different search sub-sets, and the ranking positions of the search data in different search sub-sets may be the same or different. Figure 2 As shown in , "Scenic Area A" appears in the search sub-sets retrieved from the three data sources, but the ranking position of "Scenic Area A" in each search sub-set is different.
[0049] Step 102: Obtain target data according to the data similarity between the search data and the target description text and the ranking position of the search data in the search sub-set to which it belongs.
[0050] Specifically, the target data may be data obtained by integrating the search data in all search subsets based on the ranking position and data similarity. The target data is different from the data set composed of the search data.
[0051] Step 103: Provide the target data to the intelligent model, and the target data enables the intelligent model to output the target content based on the target description text.
[0052] Among them, the intelligent model can be a large language model LLM (Large Language Model). LLM can process the target data, such as text generation, text modification, etc., to output target content that matches the target description text. The target content can be text content or image content, such as the LLM output "You can give priority to scenic spot A. If you have the physical strength, consider deep climbing in scenic spot B."
[0053] As can be seen from the above technical solution, in a data processing method provided by an embodiment of the present application, when retrieving text for a target description text from a data source, instead of simply providing target data for an intelligent model based on the sorting result of the retrieved data in the retrieval result, target data is provided for the intelligent model by combining the data similarity between the retrieved data and the target description text. That is to say, according to the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the retrieved subset to which it belongs, the target data provided to the intelligent model is obtained. Furthermore, the intelligent model outputs target content based on the target description text according to the target data. Thus, the present application can provide more accurate target data for the intelligent model, thereby improving the accuracy of the target content output by the intelligent model.
[0054] In one implementation, the data similarity in this embodiment can be obtained in the following manner:
[0055] From the retrieval result, extract the data similarity between each retrieved data in each retrieved subset and the target description text.
[0056] Among them, the retrieved data in the retrieved subset is obtained based on the data similarity, and the sorting position among the retrieved data in the retrieved subset is determined based on the data similarity. That is to say, in the retrieval result obtained from the data source, each retrieved data corresponds to the data similarity with the target description text, and the data similarity determines the sorting position of the retrieved data in the retrieved subset corresponding to the corresponding data source. Based on this, the data similarity between each retrieved data and the target description text can be extracted from the retrieval result in this embodiment.
[0057] For example, in the retrieval result output by a retrieval tool based on RAG, each retrieved data (i.e., the retrieved text data) carries the data similarity between it and the target description text. Based on this, the data similarity can be extracted from the retrieval result in this embodiment.
[0058] In one implementation, the target data in step 102 can be obtained in the following manner, as Figure 3 shown in:
[0059] Step 301: According to the sorting position of the retrieved data in the retrieved subset to which it belongs and the data similarity between the retrieved data and the target description text, obtain the sorting score corresponding to the retrieved data.
[0060] Among them, the sorting score can represent the degree of matching of the retrieved data with the target description text. Specifically, in this embodiment, the data similarity can be used to process the sorting position of the retrieved data in each retrieved subset to obtain the sorting score.
[0061] In one implementation, in step 301, the initial score corresponding to the retrieved data can be obtained first according to the sorting position of the retrieved data in the retrieved subset to which it belongs, and then the initial score can be adjusted (such as increased or decreased) based on the data similarity between the retrieved data and the target description text to obtain the sorting score corresponding to the retrieved data.
[0062] Specifically, if the data similarity between the retrieved data and the target description text is relatively high, then the initial score is increased; if the data similarity between the retrieved data and the target description text is relatively low, then the initial score is decreased to obtain the sorting score.
[0063] For example, in step 301, based on the Reciprocal Rank Fusion (RRF) algorithm, the sorting position of the retrieved data in the retrieved subset to which it belongs can be scored to obtain the initial score corresponding to the retrieved data, as shown in formula (1). Then, the data similarity between the retrieved data and the target description text is used to multiply (i.e., weight) the initial score to obtain the sorting score corresponding to the retrieved data:
[0064]
[0065] where score is the initial score corresponding to the retrieved data, n is the number of retrieved subsets, c is the smoothing constant, and r i is the sorting position of the retrieved data in the i-th retrieved subset.
[0066] It can be seen that in this embodiment, after integrating and sorting the sorting positions of the retrieved data in the retrieved subsets of multiple data sources based on the Reciprocal Rank Fusion algorithm, the data similarity between the retrieved data and the target description text can be used to adjust the sorted sorting positions, which can improve the accuracy of the integrated sorting of the retrieved data, that is, the accuracy of the sorting score.
[0067] In another implementation, in step 301, the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the retrieved subset to which it belongs can be processed based on the Reciprocal Rank Fusion algorithm to obtain the sorting score corresponding to the retrieved data.
[0068] Among them, in this embodiment, the data similarity can be used to improve the Reciprocal Rank Fusion algorithm, and then the improved Reciprocal Rank Fusion algorithm is used to process the data similarity and the sorting position to obtain the sorting score corresponding to the retrieved data.
[0069] In the specific implementation, in the improved Reciprocal Rank Fusion algorithm, the data similarity between the retrieved data and the target description text is used as the numerator, and the sum of the smoothing constant and the sorting position of the retrieved data in the retrieved subset to which it belongs is used as the denominator.
[0070] For example, in step 301, the sorting score corresponding to the retrieved data can be obtained through an improved reciprocal rank fusion algorithm such as formula (2):
[0071]
[0072] where score is the sorting score corresponding to the retrieved data, sim i is the data similarity between the retrieved data in the i-th retrieval subset and the target description text, S is the number of retrieval subsets, c is a smoothing constant, and r i is the sorting position of the retrieved data in the i-th retrieval subset.
[0073] Step 302: Obtain the target data according to the sorting score.
[0074] Specifically, in this embodiment, the target data can be extracted from the retrieved data according to the sorting score or the retrieved data can be integrated to obtain the target data.
[0075] In one implementation, in step 302, all the retrieved data can be sorted according to the sorting score corresponding to each retrieved data to obtain a sorting result, and then the target data can be obtained according to the sorting result.
[0076] For example, in this embodiment, the retrieved data is sorted according to the sorting score from largest to smallest, and then the target data is obtained according to the top N retrieved data, where N is a positive integer greater than or equal to 1.
[0077] For another example, in this embodiment, the top N with the highest sorting scores are used as the target data from the retrieved data, or data generation or adjustment is performed on the top N with the highest sorting scores to obtain the target data.
[0078] Reference Figure 4 , is a schematic structural diagram of a data processing device provided by an embodiment of the present application. The device can be configured in an electronic device capable of data processing, and the electronic device can be a computer or a server, etc. The technical solution in this embodiment is mainly used to improve the accuracy of the target data provided to the intelligent model.
[0079] Specifically, the device in this embodiment may include the following units:
[0080] A retrieval and acquisition unit 401, configured to obtain a retrieval result for a target description text, where the retrieval result at least includes retrieval subsets corresponding to multiple data sources, and the retrieval subsets include retrieved data retrieved from the corresponding data sources;
[0081] A target acquisition unit 402, configured to obtain target data according to the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the retrieved subset to which it belongs;
[0082] A data providing unit 403, configured to provide the target data to an intelligent model, and the target data enables the intelligent model to output target content based on the target description text.
[0083] As can be seen from the above technical solution, in a data processing device provided in an embodiment of the present application, when retrieving text for a target description text from a data source, the target data is not simply provided to the intelligent model based on the sorting result of the retrieved data in the retrieval result, but the target data is provided to the intelligent model in combination with the data similarity between the retrieved data and the target description text. That is to say, according to the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the retrieved subset to which it belongs, the target data provided to the intelligent model is obtained. Furthermore, the intelligent model outputs target content based on the target data and the target description text. Therefore, the present application can provide more accurate target data for the intelligent model, thereby improving the accuracy of the target content output by the intelligent model.
[0084] In one implementation, the target acquisition unit 402 is further configured to: extract the data similarity between each retrieved data and the target description text in each retrieved subset from the retrieval result; wherein, the retrieved data in the retrieved subset is obtained based on the data similarity, and the sorting position between the retrieved data in the retrieved subset is determined based on the data similarity.
[0085] In one implementation, the target acquisition unit 402 is specifically configured to: obtain a sorting score corresponding to the retrieved data according to the sorting position of the retrieved data in the retrieved subset to which it belongs and the data similarity between the retrieved data and the target description text; and obtain target data according to the sorting score.
[0086] Wherein, when the target acquisition unit 402 obtains target data according to the sorting score, it is specifically configured to: sort all the retrieved data according to the sorting score corresponding to each retrieved data to obtain a sorting result; and obtain target data according to the sorting result.
[0087] In one implementation, when the target obtaining unit 402 obtains the sorting score corresponding to the retrieved data according to the sorting position of the retrieved data in the retrieved subset to which it belongs and the data similarity between the retrieved data and the target description text, it is specifically configured to: obtain the initial score corresponding to the retrieved data according to the sorting position of the retrieved data in the retrieved subset to which it belongs; and adjust the initial score based on the data similarity between the retrieved data and the target description text to obtain the sorting score corresponding to the retrieved data.
[0088] In one implementation, when the target obtaining unit 402 obtains the sorting score corresponding to the retrieved data according to the sorting position of the retrieved data in the retrieved subset to which it belongs and the data similarity between the retrieved data and the target description text, it is specifically configured to: process the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the retrieved subset to which it belongs based on the inverse ranking fusion algorithm to obtain the sorting score corresponding to the retrieved data.
[0089] Wherein, in the inverse ranking fusion algorithm, the data similarity between the retrieved data and the target description text is used as the numerator, and the sum of the smoothing constant and the sorting position of the retrieved data in the retrieved subset to which it belongs is used as the denominator.
[0090] It should be noted that the specific implementation manners of the units in this embodiment may refer to the corresponding contents in the foregoing, and will not be elaborated herein.
[0091] The embodiment of the present application further provides a storage medium for storing a computer program, and when the computer program is executed by a processor, it implements the data processing method described in the foregoing embodiments.
[0092] Refer to Figure 5 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device may include the following structures:
[0093] A memory 501 for storing a computer program and the data generated by the running of the computer program;
[0094] A processor 502 for executing the computer program to implement: obtaining a retrieval result for a target description text, where the retrieval result at least includes retrieved subsets corresponding to multiple data sources, and the retrieved subset includes retrieved data retrieved from the data source to which it belongs; obtaining target data according to the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the retrieved subset to which it belongs; and providing the target data to an intelligent model, where the target data enables the intelligent model to output target content based on the target description text.
[0095] As can be seen from the above technical solution, in an electronic device provided in an embodiment of the present application, when retrieving text for a target description text from a data source, instead of simply providing target data for an intelligent model based on the sorting result of the retrieved data in the retrieval result, target data is provided for the intelligent model by combining the data similarity between the retrieved data and the target description text. That is to say, according to the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the retrieved subset to which it belongs, the target data provided to the intelligent model is obtained, and then the intelligent model outputs target content based on the target data based on the target description text. Thus, the present application can provide more accurate target data for the intelligent model, thereby improving the accuracy of the target content output by the intelligent model.
[0096] Taking the RAG-based retrieval tool as an example, the technical solution of the present application will be illustrated as follows:
[0097] During the retrieval process of RAG, if there are multiple data sources, such as documents, dbschemas, pictures, etc. In the RAG retrieval process, it is necessary to integrate the retrieved data queried for each data source to obtain the corresponding target data and return it to the large language model. In this process, the RRF algorithm is generally used for multi-data source fusion. However, this algorithm does not consider the relevance between the retrieved data and the user's question (i.e., the target description text), and only performs a comprehensive sorting according to the sorting position (i.e., the fixed position) of the retrieved data in the retrieved subset corresponding to the data source to which it belongs.
[0098] In response to the above problems, the present application proposes an optimization solution. This solution is based on the RAG retrieval process and makes corresponding improvements to the RRF algorithm, which can consider the sorting position of the retrieved data in different data sources and the relevance between the retrieved data and the user's question, and improve the satisfaction of users when using RAG with multiple data sources.
[0099] For example, as shown in formula (1), it is the original RRF algorithm. The sorting score of the retrieved data is calculated based on the sorting position of the retrieved data in the subsets retrieved from each data source, and then sorted. Based on this, the present application makes improvements to the RRF algorithm. As shown in formula (2), the sorting score of the retrieved data is calculated by combining the data similarity between the retrieved data and the target description text and the sorting position of the retrieved data in the subsets retrieved from each data source, and then sorted.
[0100] After adopting the technical solution of the present application, the following advantages are obtained, as shown in Table 1:
[0101] Table 1 Advantages
[0102]
[0103] In the description of the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0104] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0105] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0106] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data processing method, comprising: Obtaining a retrieval result for a target description text, the retrieval result at least including retrieval subsets corresponding to multiple data sources, and the retrieval subset including retrieval data retrieved from the corresponding data source; Obtaining target data according to the data similarity between the retrieval data and the target description text and the sorting position of the retrieval data in the corresponding retrieval subset; Providing the target data to an intelligent model, and the target data enables the intelligent model to output target content based on the target description text.
2. The method according to claim 1, the method further comprising: Extracting, from the retrieval result, the data similarity between each piece of the retrieval data and the target description text in each retrieval subset; Wherein, the retrieval data in the retrieval subset is obtained based on the data similarity, and the sorting position between the retrieval data in the retrieval subset is determined based on the data similarity.
3. The method according to claim 1 or 2, obtaining target data according to the data similarity between the retrieval data and the target description text and the sorting position of the retrieval data in the corresponding retrieval subset, comprising: Obtaining a sorting score corresponding to the retrieval data according to the sorting position of the retrieval data in the corresponding retrieval subset and the data similarity between the retrieval data and the target description text; Obtaining target data according to the sorting score.
4. The method according to claim 3, obtaining target data according to the sorting score, comprising: Sorting all the retrieval data according to the sorting score corresponding to each piece of the retrieval data to obtain a sorting result; Obtaining target data according to the sorting result.
5. The method according to claim 3, obtaining a sorting score corresponding to the retrieval data according to the sorting position of the retrieval data in the corresponding retrieval subset and the data similarity between the retrieval data and the target description text, comprising: Obtaining an initial score corresponding to the retrieval data according to the sorting position of the retrieval data in the corresponding retrieval subset; Adjusting the initial score based on the data similarity between the retrieval data and the target description text to obtain the sorting score corresponding to the retrieval data.
6. The method according to claim 3, obtaining a sorting score corresponding to the retrieval data according to the sorting position of the retrieval data in the corresponding retrieval subset and the data similarity between the retrieval data and the target description text, comprising: Processing the data similarity between the retrieval data and the target description text and the sorting position of the retrieval data in the corresponding retrieval subset based on a reciprocal rank fusion algorithm to obtain the sorting score corresponding to the retrieval data.
7. The method according to claim 6, in the reciprocal rank fusion algorithm, using the data similarity between the retrieval data and the target description text as the numerator and the sum of a smoothing constant and the sorting position of the retrieval data in the corresponding retrieval subset as the denominator.
8. A data processing device, comprising: A retrieval and acquisition unit, configured to obtain a retrieval result for a target description text, where the retrieval result at least includes retrieval subsets corresponding to multiple data sources, and the retrieval subset includes retrieval data retrieved from the corresponding data source; A target acquisition unit, configured to obtain target data according to the data similarity between the retrieval data and the target description text and the sorting position of the retrieval data in the corresponding retrieval subset; A data providing unit, configured to provide the target data to an intelligent model, and the target data enables the intelligent model to output target content based on the target description text.
9. A storage medium, configured to store a computer program, and when the computer program is executed by a processor, the following functions are implemented: Obtaining a retrieval result for a target description text, where the retrieval result at least includes retrieval subsets corresponding to multiple data sources, and the retrieval subset includes retrieval data retrieved from the corresponding data source; Obtaining target data according to the data similarity between the retrieval data and the target description text and the sorting position of the retrieval data in the corresponding retrieval subset; Providing the target data to an intelligent model, and the target data enables the intelligent model to output target content based on the target description text.
10. An electronic device, comprising: A memory, configured to store a computer program and data generated by the running of the computer program; A processor, configured to execute the computer program to implement: obtaining a retrieval result for a target description text, where the retrieval result at least includes retrieval subsets corresponding to multiple data sources, and the retrieval subset includes retrieval data retrieved from the corresponding data source; obtaining target data according to the data similarity between the retrieval data and the target description text and the sorting position of the retrieval data in the corresponding retrieval subset; providing the target data to an intelligent model, and the target data enables the intelligent model to output target content based on the target description text.
Citation Information
Cited By
Multi-source data search method and device based on artificial intelligence model
CN121542375A