Data retrieval optimization method and system

By obtaining the information entropy and semantic similarity of the wide knowledge table to merge data fields, and using the dual-depth Q-network algorithm to dynamically adjust the weights, the problem of low data retrieval efficiency under fixed templates is solved, and efficient and intelligent data retrieval and integration are achieved.

CN120632066APending Publication Date: 2025-09-12CHINA TELECOM CORP LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510727258.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When processing multi-source heterogeneous data, the existing data governance and retrieval systems use fixed templates, resulting in low retrieval efficiency and low accuracy of retrieval results. They are unable to effectively identify and integrate field differences between different data sources, leading to data redundancy and consistency issues.

Method used

By acquiring the knowledge table of the target domain, determining the information entropy and semantic similarity of data fields, merging similar or redundant fields, and dynamically adjusting field weights using the weight allocation decision model trained by the dual deep Q-network algorithm, a target retrieval template is constructed.

Benefits of technology

It has improved the pertinence and response speed of data retrieval, significantly improved retrieval efficiency and result quality, and realized the intelligence and efficiency of data governance and retrieval systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632066A_ABST
    Figure CN120632066A_ABST
Patent Text Reader

Abstract

The invention discloses a data retrieval optimization method and system. The method comprises the steps that a knowledge wide table of a target domain is obtained, the knowledge wide table comprises a plurality of knowledge databases in the target domain, and each knowledge database comprises a plurality of structured knowledge data composed of at least one data field; determining the information entropy of each data field in the knowledge wide table and the semantic similarity between every two data fields, and merging the plurality of data fields in the knowledge wide table according to the information entropy and the semantic similarity to obtain a plurality of standard data fields; analyzing the field information of each standard data field by using a weight distribution decision model to obtain a field weight of each standard data field; and constructing a target retrieval template of the knowledge wide table according to the respective field weight of each standard data field. According to the method and the device, the technical problems of low retrieval efficiency and low retrieval result accuracy caused by cross-source retrieval of data by adopting a fixed retrieval template are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data retrieval optimization method and system. Background Art

[0002] Existing data governance and retrieval systems often rely on fixed templates and preset retrieval rules when processing multi-source, heterogeneous data. This model exposes significant limitations when dealing with widely distributed and diversely formatted data sources such as propaganda, cultural tourism, and cultural law enforcement. Fixed templates were originally designed to unify data structures and facilitate data integration and retrieval. However, when data formats, field naming, and descriptions differ significantly across data sources, the universality of these templates is severely challenged. For example, a "scenic spot introduction" in tourist guide data may be called an "exhibit description" in cultural and museum data, and may be labeled "on-site conditions" in cultural law enforcement records. Although they essentially describe the same type of subject information, fixed templates cannot accurately identify and integrate these fields, leading to data redundancy and consistency issues.

[0003] Furthermore, fixed-template retrieval mechanisms typically rely on keyword matching or simple field comparison. This approach ignores the differences in information entropy between different fields, meaning that each field contributes differently to the search results. For example, the "location" field often distinguishes specific cultural heritage sites or attractions more effectively than the "ticket price" field. However, existing technologies treat all fields equally, which not only reduces retrieval efficiency but also may produce irrelevant or low-quality search results, increasing user search costs.

[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0005] The embodiments of the present application provide a data retrieval optimization method and system to at least solve the technical problem that cross-source retrieval of data using a fixed retrieval template results in low retrieval efficiency and low accuracy of retrieval results.

[0006] According to one aspect of an embodiment of the present application, a data retrieval optimization method is provided, including: obtaining a knowledge wide table of a target domain, wherein the knowledge wide table includes multiple knowledge databases within the target domain, and each knowledge database includes multiple structured knowledge data consisting of at least one data field; determining the information entropy of each data field in the knowledge wide table and the semantic similarity between every two data fields, and merging the multiple data fields in the knowledge wide table based on the information entropy and the semantic similarity to obtain multiple standard data fields; using a pre-trained weight allocation decision model to analyze the field information of each standard data field to obtain a field weight of each standard data field, wherein the weight allocation decision model is trained based on a dual-depth Q network algorithm; and constructing a target retrieval template for the knowledge wide table based on the field weights of each standard data field.

[0007] Optionally, obtaining a wide knowledge table of the target domain includes: obtaining multiple knowledge databases within the target domain; for each knowledge database, preprocessing multiple structured knowledge data in the knowledge database, and using a subject recognition model to determine the business subjects in each preprocessed structured data, wherein the preprocessing includes at least one of the following: data cleaning, format conversion, and standardization; aggregating multiple structured data describing the same business entity in multiple knowledge databases to obtain a wide knowledge table, wherein each row in the wide knowledge table represents a business entity, and each column represents an attribute information of a business entity.

[0008] Optionally, multiple structured data describing the same business entity in multiple knowledge databases are aggregated to obtain a knowledge wide table, including: for each business entity, determining the first data field describing the business entity in each knowledge database, and calculating the first semantic similarity between each first data field; based on the first semantic similarity and the entity type of the business entity, determining at least one second data field describing the same attribute of the business entity from multiple first data fields, and standardizing each second data field describing the same attribute of the business entity to obtain a unique identification field describing the same attribute of the business entity; based on the unique identification field, aggregating multiple structured data describing the same business entity in multiple knowledge databases to obtain a knowledge wide table.

[0009] Optionally, the field information includes at least: a field value, wherein the information entropy of each data field in the knowledge wide table and the semantic similarity between each two data fields are determined, and multiple data fields in the knowledge wide table are merged based on the information entropy and semantic similarity to obtain multiple standard data fields, including: for each data field in the knowledge wide table, the information entropy of the data field is calculated based on the field information of the data field, wherein, the more field values ​​of the data field appearing in the knowledge wide table and the more dispersed the occurrence frequency of each field value, the greater the information entropy of the data field; the fewer field values ​​of the data field appearing in the knowledge wide table and / or the more concentrated the occurrence frequency of each field value, the smaller the information entropy of the data field; determining at least one candidate data field whose information entropy is lower than a preset first threshold value from multiple data fields in the knowledge wide table, and calculating the semantic similarity between the field information of each candidate data field; merging at least two candidate data fields whose semantic similarity is higher than a preset second threshold value to obtain multiple standard data fields.

[0010] Optionally, the training process of the weight allocation decision model includes: constructing an online Q network and a target network for solving the weight allocation strategy, and initializing the network weight parameters of the online Q network and the target network; setting an experience pool, determining the capacity of the experience pool and initializing the priority of each sample in the experience pool; determining a first preset number of iteration cycles, and iteratively solving the weight allocation strategy and network weight parameters through the following steps: in each iteration cycle, initializing the state, and determining a second preset number of calculation cycles, wherein the state includes at least: the semantic similarity between each standard data field, the information entropy of each standard data field and the retrieval usage frequency; in each calculation cycle, inputting the current state into the online Q network, calculating the predicted Q values ​​corresponding to different weight allocation strategies, and selecting the target weight corresponding to the maximum Q value based on the greedy strategy. allocation strategy, execute the target weight allocation strategy and calculate the reward, obtain the new state, take the current state, target weight allocation strategy, reward and new state as a sample, and store the sample in the experience pool; sample multiple samples based on the priority of each sample in the experience pool, input the neural network containing a bidirectional long short-term memory network, calculate the probability of each sample being sampled, the mean square error loss function and the loss function weight, determine the target Q value based on the calculation results and update the network weight parameters of the online Q network, calculate the error of all samples in the experience pool and update the priority of all samples; update the network weight parameters of the target network based on the network weight parameters of the online Q network after a third preset number of calculation cycles, where the third preset number is less than the second preset number; after the iteration is completed, the obtained target network is used as the weight allocation decision model.

[0011] Optionally, the current state is input into the online Q network to calculate the predicted Q values ​​corresponding to different weight allocation strategies, including: determining multiple weight allocation strategies; for each weight allocation strategy, determining the retrieval time overhead, retrieval resource overhead, retrieval accuracy and user satisfaction corresponding to the weight allocation strategy, and determining the total retrieval overhead of the weight allocation strategy based on the retrieval time overhead, retrieval resource overhead, retrieval accuracy and user satisfaction; determining the predicted Q values ​​corresponding to various weight allocation strategies based on the total retrieval overheads of various weight allocation strategies, wherein the weight allocation strategy with a larger total retrieval overhead has a smaller predicted Q value.

[0012] Optionally, a target retrieval template for the wide knowledge table is constructed based on the field weights of each standard data field, including: arranging and combining each standard data field in order from large to small according to the field weights to obtain the target retrieval template for the wide knowledge table.

[0013] According to another aspect of an embodiment of the present application, a data retrieval optimization system is also provided, including: an acquisition module for acquiring a knowledge wide table of a target domain, wherein the knowledge wide table includes multiple knowledge databases within the target domain, and each knowledge database includes multiple structured knowledge data consisting of at least one data field; a field processing module for determining the information entropy of each data field in the knowledge wide table and the semantic similarity between each two data fields, and merging the multiple data fields in the knowledge wide table based on the information entropy and semantic similarity to obtain multiple standard data fields; a weight allocation module for analyzing the field information of each standard data field using a pre-trained weight allocation decision model to obtain the field weight of each standard data field, wherein the weight allocation decision model is trained based on a dual-depth Q network algorithm; a construction module for constructing a target retrieval template for the knowledge wide table based on the field weights of each standard data field.

[0014] According to another aspect of an embodiment of the present application, a computer program product is further provided, the computer program product comprising: a computer program, wherein when the computer program is executed by a processor, the above-mentioned data retrieval optimization method is implemented.

[0015] According to another aspect of an embodiment of the present application, an electronic device is provided, which includes: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned data retrieval optimization method through the computer program.

[0016] In an embodiment of the present application, the information entropy of each data field in the knowledge wide table of the aggregating multi-source knowledge database and the semantic similarity between each two data fields are judged, similar fields or redundant fields are merged to form a unified standard data field; then, the weight distribution decision model obtained by training the double-depth Q network algorithm is used to analyze the field information of each standard data field to obtain a dynamically adjusted field weight. This weight distribution mechanism can ensure that in the data retrieval process, high information entropy fields are matched first, thereby improving the relevance of the retrieval results and the system response speed; finally, a target retrieval template of the knowledge wide table is constructed based on the field weights of each standard data field. The template not only contains the standardized fields, but also embeds dynamic weight information, so that in subsequent retrieval operations, the system can perform intelligent matching based on the field weights, significantly improving retrieval efficiency and result quality. This significantly improves the pertinence and response speed of data retrieval, and provides a more intelligent and efficient solution for data governance and retrieval systems. This solves the technical problem of using a fixed retrieval template to perform cross-source retrieval of data, resulting in low retrieval efficiency and low accuracy of retrieval results. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 is a flow chart of an optional data retrieval optimization method according to an embodiment of the present application;

[0019] Figure 2 is a structural diagram of an optional data retrieval optimization system according to an embodiment of the present application;

[0020] Figure 3 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0022] It should be noted that the terms "first", "second", etc. in the specification, claims, and drawings of the present application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.

[0023] In order to better understand the embodiments of the present application, some nouns or terms that appear in the description of the embodiments of the present application are first translated and explained as follows:

[0024] The Double Deep Q Network (DDQN) algorithm combines deep learning and Q-learning, using neural networks to approximate the Q function to solve complex, high-dimensional state problems. DDQN uses two independent value functions to select and evaluate actions, reducing bias in estimates and improving algorithm performance. Specifically, when updating the value function, DDQN uses one network to select the optimal action and another to evaluate the value of that action. This approach reduces bias caused by using the same value function for both selection and evaluation.

[0025] Example 1

[0026] According to an embodiment of the present application, a data retrieval optimization method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0027] Figure 1 is a flow chart of a data retrieval optimization method provided according to an embodiment of the present application, such as Figure 1 As shown, the method includes the following steps:

[0028] Step S102: Obtain a wide knowledge table of the target domain, wherein the wide knowledge table includes multiple knowledge databases in the target domain, and each knowledge database includes multiple structured knowledge data consisting of at least one data field.

[0029] Step S104 , determining the information entropy of each data field in the knowledge wide table and the semantic similarity between every two data fields, and merging multiple data fields in the knowledge wide table according to the information entropy and semantic similarity to obtain multiple standard data fields.

[0030] Step S106 : Analyze the field information of each standard data field using the pre-trained weight allocation decision model to obtain the field weight of each standard data field.

[0031] Step S108: constructing a target search template for the wide knowledge table based on the field weights of the respective standard data fields.

[0032] Based on the scheme defined by the above steps S102 to S108, it can be known that in the embodiment of the present application, the information entropy of each data field in the knowledge wide table of the aggregating multi-source knowledge database and the semantic similarity between each two data fields are judged, and similar fields or redundant fields are merged to form a unified standard data field; then, the weight allocation decision model obtained by training the dual-depth Q network algorithm is used to analyze the field information of each standard data field to obtain dynamically adjusted field weights. This weight allocation mechanism can ensure that high information entropy fields are matched first during the data retrieval process, thereby improving the relevance of the retrieval results and the system response speed; finally, a target retrieval template for the knowledge wide table is constructed based on the field weights of each standard data field. The template not only contains the standardized fields, but also embeds dynamic weight information, so that in subsequent retrieval operations, the system can perform intelligent matching based on the field weights, significantly improving retrieval efficiency and result quality.

[0033] Through the above scheme, the embodiment of the present application overcomes the limitations of the existing technology of static fixed templates, low retrieval efficiency and inability to intelligently process cross-source data integration, and realizes automatic aggregation of data fields and dynamic optimization of weights, thereby greatly improving the accuracy of data retrieval and system operation efficiency while ensuring data quality.

[0034] The following describes each step of the data retrieval optimization method in conjunction with a specific implementation process.

[0035] As an optional implementation, in the technical solution provided in step S102 above, the system may obtain the broad knowledge table of the target domain according to the following method, including:

[0036] Step S1021: Acquire multiple knowledge databases in the target domain.

[0037] For example, if the target domain is cultural promotion and related fields, such as the cultural tourism industry (including data on tourist attraction introductions, cultural heritage protection, and travel route planning), the cultural museum industry (including information on collections, exhibitions, and research materials from cultural institutions such as museums, art galleries, and libraries), cultural law enforcement (including case records, legal documents, and enforcement documents in scenarios such as cultural market supervision, copyright protection, and cultural heritage enforcement), and academic research (such as data on cultural studies, history, and art), then the multiple knowledge databases within the target domain could include museum databases, law enforcement record databases, and tourist guide databases.

[0038] Step S1022 : For each knowledge database, pre-process the plurality of structured knowledge data in the knowledge database, and use the subject identification model to determine the business subject in each pre-processed structured data.

[0039] Specifically, a series of preprocessing measures, such as data cleaning, format conversion, and standardization, are applied to the structured data within each knowledge database to ensure data quality and consistency. Data cleaning involves removing noise, filling missing values, and correcting erroneous data. Format conversion and standardization convert data into a unified format for subsequent processing and analysis.

[0040] Next, the subject identification model is used to identify the business entities within each preprocessed structured data set. This subject identification model is based on a deep learning model trained on a large-scale corpus to capture rich contextual information and long-range dependencies, thereby more accurately identifying and classifying entities. Furthermore, the business entities mentioned here can be specific entity categories within the target domain. For example, business entities in the field of cultural promotion could include artists, cultural relics, and cultural activities.

[0041] It should be noted that domain experts or professional teams will annotate a portion of the collected data with business entities to form a high-quality "benchmark annotation dataset," which can be used as a corpus for training the subject recognition model. The subject recognition results obtained from the model analysis are regularly compared with the corresponding data in the "benchmark annotation dataset," and statistical indicators (such as accuracy, recall, F1 value, etc.) are used to evaluate the performance of the current model. If the model performance degrades, online incremental learning is continued using data with large deviations as feedback data to optimize model parameters and improve the accuracy of the processing results. At the same time, a shadow model (i.e., an old model with good historical performance) is run to monitor the real-time performance of the current model. If the key performance indicators of the current model are detected to be lower than the shadow model or a predetermined threshold, a rollback mechanism is automatically triggered to switch to the historically stable model to avoid negative impacts caused by sudden performance declines.

[0042] In step S1023, multiple structured data describing the same business entity in multiple knowledge databases are aggregated to generate a wide knowledge table. Each row in the wide knowledge table represents a business entity, and each column represents an attribute of a business entity (such as name, era, material, location, etc.). Because this wide knowledge table provides a complete set of attributes for each business entity while eliminating data redundancy, it can be used for subsequent data retrieval, analysis, and management.

[0043] It should be noted that if the initial data to be stored is unstructured data or semi-structured data, before storing the data in the corresponding knowledge database, the unstructured data or semi-structured data can be converted into structured data in advance to facilitate systematic management.

[0044] Specifically, if the initial data is text, it can be broken down into sentences and phrases using word segmentation and syntactic analysis tools. If the initial data is image data, optical character recognition algorithms can be used to identify the text within the image. Next, named entity recognition algorithms are used to identify entities within the results of the previous step, and relationship extraction techniques are applied to determine the relationships between entities. For example, the identified entities could be "Bronze Tripod," "Late Shang Dynasty," "National Museum," and so on. The extracted relationships between entities could be the "creation period" relationship between "Bronze Tripod" and "Late Shang Dynasty," and the "collection ownership" relationship between "National Museum" and "Bronze Tripod." The identified entities are then standardized and mapped to a predefined data schema. For example, "Late Shang Dynasty" could be uniformly labeled "Late Shang Dynasty," and a structured field "creation period" could be created to store this information. Finally, the results from the previous step are summarized and stored in a knowledge database.

[0045] Optionally, in the technical solution provided in step S1023 above, the system may perform data aggregation according to the following steps to obtain a wide knowledge table, including:

[0046] The first step is to determine the first data field describing the business entity in each knowledge database and calculate the first semantic similarity between each first data field; based on the first semantic similarity and the entity type of the business entity, determine at least one second data field describing the same attribute of the business entity from multiple first data fields, and standardize each second data field describing the same attribute of the business entity to obtain a unique identification field describing the same attribute of the business entity.

[0047] Step 2: Based on the unique identification field, multiple structured data describing the same business entity in multiple knowledge databases are aggregated to obtain a knowledge wide table.

[0048] That is, for each identified business entity, all data fields describing that entity are automatically scanned and identified from various knowledge databases. By calculating semantic similarity and entity type matching between fields, it is determined which fields describe the same attribute of the business entity. For example, for the business entity "Bronze Tripod," "Casting Period" and "Year of Production" might be identified as temporal attributes describing "Bronze Tripod." Next, the fields describing the same attribute of the business entity are standardized, such as unifying the field name, data type, and format, to ensure that subsequent data aggregation can proceed smoothly. Finally, based on the generated unique identification field, all structured data describing the same business entity from different knowledge databases are aggregated to obtain a knowledge wide table, where each row represents a business entity and each column contains various attribute information of that entity.

[0049] For example, Table 1 is a schematic diagram of a knowledge database, as shown in Table 1 below.

[0050] Table 1

[0051] Cultural Relic Number name Material Creation period Favorite Location Database Source 001 Bronze tripod bronze Late Shang Dynasty Museum A Database 1 001 Bronze tripod / / Museum B Database 2 002 painted pottery jars clay Neolithic Museum B Database 2

[0052] It is not difficult to see from Table 1 that "Cultural Relic No. 001" appears twice in different knowledge databases. Therefore, through attribute aggregation, we can obtain the wide table shown in Table 2 below.

[0053] Table 2

[0054] Cultural Relic Number name Material Creation period Favorite Location 001 Bronze tripod bronze Late Shang Dynasty Museum A, Museum B 002 painted pottery jars clay Neolithic Museum B

[0055] As an optional implementation, in the technical solution provided in step S104 above, the system may merge the data fields according to the following method to obtain multiple standard data fields, including:

[0056] Step S1041 : For each data field in the knowledge wide table, the information entropy of the data field is calculated based on the field information of the data field.

[0057] The information entropy of a data field is calculated based on the types and frequency of field values ​​that appear in that data field. The more diverse the field values ​​corresponding to a data field are and the more dispersed the frequency of occurrence of different field values, the higher the information entropy of the data field. Conversely, the fewer field values ​​corresponding to a data field are or the frequency of occurrence of different field values ​​is concentrated, the lower the information entropy of the data field. Therefore, the information entropy of data fields can ensure that the system can identify which fields provide rich and valuable information and which fields may be redundant or lack distinctiveness.

[0058] Step S1042 : determining at least one candidate data field whose information entropy is lower than a preset first threshold value from the multiple data fields in the knowledge wide table, and calculating the semantic similarity between the field information of each candidate data field.

[0059] Among them, the characteristics of these candidate data fields are that they provide less value information, or the frequency of the same field value in the target database is too high, that is, these candidate data fields can be considered redundant or lacking in distinctiveness. Therefore, the embodiments of the present application calculate the semantic similarity between the field information of each candidate data field to determine whether the two fields contain duplicate or similar information when describing the same business entity or concept.

[0060] Step S1043 : Merge at least two candidate data fields whose semantic similarity is higher than a preset second threshold value to obtain a plurality of standard data fields.

[0061] The merged multiple standard data fields also contain information of the original multiple data fields, thereby reducing redundancy and improving the retrieval efficiency and data consistency of the target database.

[0062] It should be noted that in the process of merging at least two candidate data fields whose semantic similarity is higher than a preset second threshold value, metadata such as field names and data types can also be standardized to ensure that the merged fields can be correctly identified and used in subsequent data retrieval and processing.

[0063] Furthermore, the system calls the weight allocation decision model obtained by training the dual deep Q network algorithm to analyze the field information (such as field name, field description, and information entropy) of each standard data field to obtain the field weight of each standard data field. The field weight reflects the importance and information content of the data field in practical applications. Therefore, the larger the field weight, the more critical the role of the data field in data retrieval and processing.

[0064] Specifically, the training process of the above weight allocation decision model includes:

[0065] Construct an online Q network and a target network for solving the weight allocation strategy, and initialize the network weight parameters of the online Q network and the target network;

[0066] Set up the experience pool, determine the capacity of the experience pool and initialize the priority of each sample in the experience pool;

[0067] Determine a first preset number of iteration cycles, and iteratively solve the weight allocation strategy and network weight parameters through the following steps:

[0068] In each iteration cycle, a state is initialized, and a second preset number of calculation cycles is determined, wherein the state includes at least: information entropy of each standard data field;

[0069] In each calculation cycle, the current state is input into the online Q network, the predicted Q values ​​corresponding to different weight allocation strategies are calculated, and the target weight allocation strategy corresponding to the maximum Q value is selected based on the greedy strategy. The target weight allocation strategy is executed and the reward is calculated to obtain the new state. The current state, target unloading strategy, reward and new state are taken as a sample and stored in the experience pool. Based on the priority of each sample in the experience pool, multiple samples are sampled and input into a neural network containing a bidirectional long short-term memory network. The probability of each sample being sampled, the mean square error loss function and the loss function weight are calculated. Based on the calculation results, the target Q value is determined and the network weight parameters of the online Q network are updated. The errors of all samples in the experience pool are calculated and the priorities of all samples are updated.

[0070] updating the network weight parameters of the target network according to the network weight parameters of the online Q network after a third preset number of calculation cycles, wherein the third preset number is less than the second preset number;

[0071] After the iteration is completed, the obtained target network is used as the weight allocation decision model.

[0072] Specifically, during the model training process, the size of the Q value reflects the expected value of the future cumulative reward after taking a certain action. Therefore, when calculating the predicted Q value corresponding to different weight distribution strategies, the embodiment of the present application can be performed as follows:

[0073] Step 1: Determine multiple weight distribution strategies.

[0074] Step 2: For each weight allocation strategy, determine the retrieval time cost, retrieval resource cost, retrieval accuracy and user satisfaction corresponding to the weight allocation strategy, and determine the total retrieval cost corresponding to the weight allocation strategy based on the retrieval time cost, retrieval resource cost, retrieval accuracy and user satisfaction.

[0075] Retrieval time overhead is the average time required to complete a search operation, which includes data preprocessing, index lookup, data matching, and sorting. Therefore, retrieval time overhead can be estimated by measuring the response time of a series of search operations and then calculating the average or percentile. Retrieval resource overhead is the computing resources required to complete a search operation, including CPU usage, memory usage, and disk I / O. Therefore, computing resource overhead can be measured by monitoring the system's resource consumption during search execution. Retrieval accuracy indirectly affects total retrieval cost because low-accuracy searches require users to perform additional operations to find the correct information, which indirectly increases total retrieval cost. Therefore, searches with low accuracy can be considered to incur additional costs, for example, by converting the percentage of low accuracy into a virtual loss of time or computing resources. User satisfaction is the user's satisfaction with the search results, which also affects total retrieval cost. Dissatisfied users may need to search again or seek information through other channels, which invisibly increases the system's burden. User satisfaction can generally be quantified through user behavior data such as questionnaires, click-through rates, and dwell time.

[0076] Then, by taking a weighted sum of retrieval time, resource cost, accuracy, and user satisfaction, the total retrieval cost of the weighted allocation strategy is obtained. The weighting coefficients for retrieval time, resource cost, accuracy, and user satisfaction can be customized based on actual business needs. For example, in some cases, the accuracy and completeness of retrieval results may be more important than retrieval speed; in this case, retrieval accuracy should be given a higher weight. In high-traffic scenarios, user waiting time and computational efficiency may take precedence; in this case, retrieval time and user satisfaction should be given higher weights.

[0077] Step 3: Determine the predicted Q-values ​​for each weighted allocation strategy based on the corresponding total search overhead. The smaller the total search overhead, the larger the predicted Q-value for the corresponding weighted allocation strategy. This is because a smaller overhead means more efficient resource utilization or a better user experience, resulting in a larger cumulative reward. Conversely, a larger total search overhead means a smaller predicted Q-value for the corresponding weighted allocation strategy. This is because a larger overhead means significant resource waste, long user wait times, or unsatisfactory search results, resulting in a smaller cumulative reward.

[0078] Therefore, through the above model training method, we can continuously learn and adjust the weights to find an optimal balance between different weight allocation strategies, that is, to find a weight allocation strategy that can bring high cumulative rewards while maintaining low overhead, and ultimately achieve an efficient, low-cost and highly satisfying retrieval experience.

[0079] As an optional implementation, in the technical solution provided in the above step S108, the system can arrange and combine the various standard data fields in order from large to small field weights to obtain a target retrieval template of the target database.

[0080] In other words, all standard data fields are sorted from largest to smallest by weight, and the target search template is constructed based on this sorted sequence of standard data fields. Therefore, within the target search template, the standard data field with the highest field weight is ranked first, followed by the standard data field with the next highest field weight, and so on. This optimization not only considers the weight of individual data fields but also the interactions and combined effects between data fields, adjusting the relative positions of data fields to achieve more efficient and accurate data retrieval.

[0081] Once the target search template is constructed, the system can dynamically adjust the field weights within it based on real-time feedback and data changes, ensuring that the template consistently reflects the current data value and characteristics of the wide knowledge table. The adjusted target search template is then applied to the data search process, guiding the system to search based on the importance of data fields when retrieving the wide knowledge table, significantly improving search efficiency and accuracy.

[0082] In summary, in the above-mentioned data retrieval optimization method, considering the diversity of business entity identifiers in different databases, it is difficult to effectively associate them, which affects the efficiency of data integration and analysis. For this reason, it is proposed to align cross-source data through business entity pairs, solve the problem of cross-source data association, and improve the consistency and linkability of data. In addition, considering that existing data retrieval is mainly through fixed retrieval templates and manual rules, it responds slowly to changes in data formats and has poor template adaptability. For this reason, it is proposed to generate retrieval templates by automatically aggregating fields and intelligently assigning appropriate field weights to the aggregated fields, which greatly reduces the need for manual intervention and improves retrieval accuracy and response speed. Therefore, the embodiment of the present application can solve the core problems of the prior art such as rigid retrieval templates, inconsistent subject identifiers, low retrieval efficiency, and excessive consumption of retrieval resources, realize the intelligence, automation and security of data governance, and significantly improve the efficiency, quality and security of data integration.

[0083] Example 2

[0084] According to an embodiment of the present application, a data retrieval optimization system for implementing the data retrieval optimization method in embodiment 1 is also provided. Figure 2 As shown, the data retrieval optimization system at least includes: an acquisition module 22, a field processing module 24, a weight allocation module 26 and a construction module 28, wherein:

[0085] An acquisition module 22 is configured to acquire a wide knowledge table of a target domain, wherein the wide knowledge table includes a plurality of knowledge databases within the target domain, and each knowledge database includes a plurality of structured knowledge data consisting of at least one data field;

[0086] A field processing module 24 is used to determine the information entropy of each data field in the knowledge wide table and the semantic similarity between each two data fields, and merge multiple data fields in the knowledge wide table based on the information entropy and semantic similarity to obtain multiple standard data fields;

[0087] A weight allocation module 26 is configured to analyze the field information of each standard data field using a pre-trained weight allocation decision model to obtain a field weight for each standard data field, wherein the weight allocation decision model is trained based on a dual deep Q network algorithm;

[0088] The construction module 28 is used to construct a target search template of the knowledge wide table according to the field weights of each standard data field.

[0089] The following describes the functions of each module of the data retrieval optimization system in conjunction with the specific implementation process.

[0090] As an optional implementation, the acquisition module 22 may acquire the target domain knowledge table in the following manner, including:

[0091] Step S11: Acquire multiple knowledge databases in the target domain.

[0092] For example, if the target domain is cultural promotion and related fields, such as the cultural tourism industry (including data on tourist attraction introductions, cultural heritage protection, and travel route planning), the cultural museum industry (including information on collections, exhibitions, and research materials from cultural institutions such as museums, art galleries, and libraries), cultural law enforcement (including case records, legal documents, and enforcement documents in scenarios such as cultural market supervision, copyright protection, and cultural heritage enforcement), and academic research (such as data on cultural studies, history, and art), then the multiple knowledge databases within the target domain could include museum databases, law enforcement record databases, and tourist guide databases.

[0093] Step S12: for each knowledge database, pre-process the plurality of structured knowledge data in the knowledge database, and use the subject identification model to determine the business subject in each pre-processed structured data.

[0094] Specifically, a series of preprocessing measures, such as data cleaning, format conversion, and standardization, are applied to the structured data within each knowledge database to ensure data quality and consistency. Data cleaning involves removing noise, filling missing values, and correcting erroneous data. Format conversion and standardization convert data into a unified format for subsequent processing and analysis.

[0095] Next, the subject identification model is used to identify the business entities within each preprocessed structured data set. This subject identification model is based on a deep learning model trained on a large-scale corpus to capture rich contextual information and long-range dependencies, thereby more accurately identifying and classifying entities. Furthermore, the business entities mentioned here can be specific entity categories within the target domain. For example, business entities in the field of cultural promotion could include artists, cultural relics, and cultural activities.

[0096] In step S13, multiple structured data describing the same business entity in multiple knowledge databases are aggregated to form a wide knowledge table. Each row in the wide knowledge table represents a business entity, and each column represents an attribute of a business entity (such as name, era, material, location, etc.). Because this wide knowledge table provides a complete set of attributes for each business entity while eliminating data redundancy, it can be used for subsequent data retrieval, analysis, and management.

[0097] Optionally, in the technical solution provided in step S13 above, the acquisition module 22 may perform data aggregation according to the following steps to obtain a wide knowledge table, including:

[0098] The first step is to determine the first data field describing the business entity in each knowledge database and calculate the first semantic similarity between each first data field; based on the first semantic similarity and the entity type of the business entity, determine at least one second data field describing the same attribute of the business entity from multiple first data fields, and standardize each second data field describing the same attribute of the business entity to obtain a unique identification field describing the same attribute of the business entity.

[0099] Step 2: Based on the unique identification field, multiple structured data describing the same business entity in multiple knowledge databases are aggregated to obtain a knowledge wide table.

[0100] That is, for each identified business entity, all data fields describing that entity are automatically scanned and identified from various knowledge databases. By calculating semantic similarity and entity type matching between fields, it is determined which fields describe the same attribute of the business entity. For example, for the business entity "Bronze Tripod," "Casting Period" and "Year of Production" might be identified as temporal attributes describing "Bronze Tripod." Next, the fields describing the same attribute of the business entity are standardized, such as unifying the field name, data type, and format, to ensure that subsequent data aggregation can proceed smoothly. Finally, based on the generated unique identification field, all structured data describing the same business entity from different knowledge databases are aggregated to obtain a knowledge wide table, where each row represents a business entity and each column contains various attribute information of that entity.

[0101] As an optional implementation, the field processing module 24 may merge the data fields according to the following method to obtain multiple standard data fields, including:

[0102] Step S21 : for each data field in the knowledge wide table, the information entropy of the data field is calculated based on the field information of the data field.

[0103] The information entropy of a data field is calculated based on the types and frequency of field values ​​that appear in that data field. The more diverse the field values ​​corresponding to a data field are and the more dispersed the frequency of occurrence of different field values, the higher the information entropy of the data field. Conversely, the fewer field values ​​corresponding to a data field are or the frequency of occurrence of different field values ​​is concentrated, the lower the information entropy of the data field. Therefore, the information entropy of data fields can ensure that the system can identify which fields provide rich and valuable information and which fields may be redundant or lack distinctiveness.

[0104] Step S22: determining at least one candidate data field whose information entropy is lower than a preset first threshold value from the multiple data fields in the knowledge wide table, and calculating the semantic similarity between the field information of each candidate data field.

[0105] Among them, the characteristics of these candidate data fields are that they provide less value information, or the frequency of the same field value in the target database is too high, that is, these candidate data fields can be considered redundant or lacking in distinctiveness. Therefore, the embodiments of the present application calculate the semantic similarity between the field information of each candidate data field to determine whether the two fields contain duplicate or similar information when describing the same business entity or concept.

[0106] Step S23 : merging at least two candidate data fields whose semantic similarity is higher than a preset second threshold value to obtain a plurality of standard data fields.

[0107] The merged multiple standard data fields also contain information of the original multiple data fields, thereby reducing redundancy and improving the retrieval efficiency and data consistency of the target database.

[0108] Furthermore, the weight allocation module 26 calls the weight allocation decision model obtained by training the dual-depth Q network algorithm to analyze the field information of each standard data field to obtain the field weight of each standard data field, wherein the field weight reflects the importance and information content of the data field in practical applications. Therefore, the larger the field weight, the more critical the role of the data field in data retrieval and processing.

[0109] Specifically, the training process of the above weight allocation decision model includes:

[0110] Construct an online Q network and a target network for solving the weight allocation strategy, and initialize the network weight parameters of the online Q network and the target network;

[0111] Set up the experience pool, determine the capacity of the experience pool and initialize the priority of each sample in the experience pool;

[0112] Determine a first preset number of iteration cycles, and iteratively solve the weight allocation strategy and network weight parameters through the following steps:

[0113] In each iteration cycle, a state is initialized, and a second preset number of calculation cycles is determined, wherein the state includes at least: information entropy of each standard data field;

[0114] In each calculation cycle, the current state is input into the online Q network, the predicted Q values ​​corresponding to different weight allocation strategies are calculated, and the target weight allocation strategy corresponding to the maximum Q value is selected based on the greedy strategy. The target weight allocation strategy is executed and the reward is calculated to obtain the new state. The current state, target unloading strategy, reward and new state are taken as a sample and stored in the experience pool. Based on the priority of each sample in the experience pool, multiple samples are sampled and input into a neural network containing a bidirectional long short-term memory network. The probability of each sample being sampled, the mean square error loss function and the loss function weight are calculated. Based on the calculation results, the target Q value is determined and the network weight parameters of the online Q network are updated. The errors of all samples in the experience pool are calculated and the priorities of all samples are updated.

[0115] updating the network weight parameters of the target network according to the network weight parameters of the online Q network after a third preset number of calculation cycles, wherein the third preset number is less than the second preset number;

[0116] After the iteration is completed, the obtained target network is used as the weight allocation decision model.

[0117] Specifically, during the model training process, the size of the Q value reflects the expected value of the future cumulative reward after taking a certain action. Therefore, when calculating the predicted Q value corresponding to different weight distribution strategies, the embodiment of the present application can be performed as follows:

[0118] Step 1: Determine multiple weight distribution strategies.

[0119] Step 2: For each weight allocation strategy, determine the retrieval time cost, retrieval resource cost, retrieval accuracy and user satisfaction corresponding to the weight allocation strategy, and determine the total retrieval cost corresponding to the weight allocation strategy based on the retrieval time cost, retrieval resource cost, retrieval accuracy and user satisfaction.

[0120] Retrieval time overhead is the average time required to complete a search operation, which includes data preprocessing, index lookup, data matching, and sorting. Therefore, retrieval time overhead can be estimated by measuring the response time of a series of search operations and then calculating the average or percentile. Retrieval resource overhead is the computing resources required to complete a search operation, including CPU usage, memory usage, and disk I / O. Therefore, computing resource overhead can be measured by monitoring the system's resource consumption during search execution. Retrieval accuracy indirectly affects total retrieval cost because low-accuracy searches require users to perform additional operations to find the correct information, which indirectly increases total retrieval cost. Therefore, searches with low accuracy can be considered to incur additional costs, for example, by converting the percentage of low accuracy into a virtual loss of time or computing resources. User satisfaction is the user's satisfaction with the search results, which also affects total retrieval cost. Dissatisfied users may need to search again or seek information through other channels, which invisibly increases the system's burden. User satisfaction can generally be quantified through user behavior data such as questionnaires, click-through rates, and dwell time.

[0121] Then, by taking a weighted sum of retrieval time, resource cost, accuracy, and user satisfaction, the total retrieval cost of the weighted allocation strategy is obtained. The weighting coefficients for retrieval time, resource cost, accuracy, and user satisfaction can be customized based on actual business needs. For example, in some cases, the accuracy and completeness of retrieval results may be more important than retrieval speed; in this case, retrieval accuracy should be given a higher weight. In high-traffic scenarios, user waiting time and computational efficiency may take precedence; in this case, retrieval time and user satisfaction should be given higher weights.

[0122] Step 3: Determine the predicted Q-values ​​for each weighted allocation strategy based on the corresponding total search overhead. The smaller the total search overhead, the larger the predicted Q-value for the corresponding weighted allocation strategy. This is because a smaller overhead means more efficient resource utilization or a better user experience, resulting in a larger cumulative reward. Conversely, a larger total search overhead means a smaller predicted Q-value for the corresponding weighted allocation strategy. This is because a larger overhead means significant resource waste, long user wait times, or unsatisfactory search results, resulting in a smaller cumulative reward.

[0123] As an optional implementation, the construction module 28 may arrange and combine the standard data fields in order according to the field weight from large to small to obtain a target search template of the target database.

[0124] In other words, all standard data fields are sorted from largest to smallest by weight, and the target search template is constructed based on this sorted sequence of standard data fields. Therefore, within the target search template, the standard data field with the highest field weight is ranked first, followed by the standard data field with the next highest field weight, and so on. This optimization not only considers the weight of individual data fields but also the interactions and combined effects between data fields, adjusting the relative positions of data fields to achieve more efficient and accurate data retrieval.

[0125] Once the target search template is constructed, the system can dynamically adjust the field weights within it based on real-time feedback and data changes, ensuring that the template consistently reflects the current data value and characteristics of the wide knowledge table. The adjusted target search template is then applied to the data search process, guiding the system to search based on the importance of data fields when retrieving the wide knowledge table, significantly improving search efficiency and accuracy.

[0126] It should be noted that each module in the data retrieval optimization system in the embodiment of the present application corresponds one-to-one to each implementation step of the data retrieval optimization method in Example 1. Since a detailed description has been given in Example 1, some details not reflected in this embodiment can be referred to Example 1 and will not be elaborated here.

[0127] Example 3

[0128] According to an embodiment of the present application, a computer program product is further provided. The computer program product includes a computer program, wherein when the computer program is executed by a processor, the data retrieval optimization method in Example 1 is implemented.

[0129] According to an embodiment of the present application, a non-volatile storage medium is also provided, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the data retrieval optimization method in Example 1 by running the computer program.

[0130] According to an embodiment of the present application, a processor is further provided, which is used to run a computer program, wherein the data retrieval optimization method in Example 1 is executed when the computer program is running.

[0131] According to an embodiment of the present application, an electronic device is also provided, which includes: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the data retrieval optimization method in Example 1 through the computer program.

[0132] Specifically, the computer program executes the following steps when it is running: obtaining a knowledge wide table of the target domain, wherein the knowledge wide table includes multiple knowledge databases in the target domain, and each knowledge database includes multiple structured knowledge data consisting of at least one data field; determining the information entropy of each data field in the knowledge wide table and the semantic similarity between each two data fields, and merging the multiple data fields in the knowledge wide table based on the information entropy and the semantic similarity to obtain multiple standard data fields; using a pre-trained weight allocation decision model to analyze the field information of each standard data field to obtain the field weight of each standard data field, wherein the weight allocation decision model is trained based on a dual-depth Q network algorithm; and constructing a target retrieval template for the knowledge wide table based on the field weights of each standard data field.

[0133] As an optional implementation, the electronic device may be in the form of a mobile terminal, an electronic device or a similar computing system. Figure 3 FIG1 shows a hardware structure block diagram of an electronic device for implementing a data retrieval optimization method. Figure 3 As shown, the electronic device 30 may include one or more (illustrated as 302a, 302b, ..., 302n in the figure) processors 302 (the processor 302 may include but is not limited to a processing system such as a microprocessor MCU or a programmable logic device FPGA), a memory 304 for storing data, and a transmission system 306 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 3 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 3More or fewer components than shown, or with Figure 3 Different configurations shown.

[0134] It should be noted that the one or more processors 302 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the electronic device 30. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0135] The memory 304 can be used to store software programs and modules of application software, such as the program instructions / data storage system corresponding to the data retrieval optimization method in the embodiment of the present application. The processor 302 executes various functional applications and data processing by running the software programs and modules stored in the memory 304, that is, implementing the vulnerability detection method of the above-mentioned application. The memory 304 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage systems, flash memory, or other non-volatile solid-state memory. In some instances, the memory 304 may further include a memory remotely located relative to the processor 302, and these remote memories may be connected to the electronic device 30 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0136] The transmission system 306 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the electronic device 30. In one embodiment, the transmission system 306 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission system 306 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0137] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the electronic device 30 .

[0138] The serial numbers of the above embodiments are for description only and do not represent the advantages or disadvantages of the embodiments.

[0139] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0140] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the system embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0141] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0142] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0143] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program code.

[0144] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A data retrieval optimization method, characterized in that: include: Acquire a wide knowledge table of a target domain, wherein the wide knowledge table includes a plurality of knowledge databases in the target domain, and each of the knowledge databases includes a plurality of structured knowledge data consisting of at least one data field; Determining the information entropy of each data field in the wide knowledge table and the semantic similarity between each two data fields, and merging multiple data fields in the wide knowledge table according to the information entropy and the semantic similarity to obtain multiple standard data fields; Analyzing the field information of each of the standard data fields using a pre-trained weight allocation decision model to obtain a field weight for each of the standard data fields, wherein the weight allocation decision model is trained based on a dual deep Q network algorithm; A target search template of the wide knowledge table is constructed according to the field weights of the respective standard data fields.

2. The method according to claim 1, characterized in that Acquire a broad knowledge base in the target domain, including: Acquiring multiple knowledge databases in the target domain; For each of the knowledge databases, preprocessing is performed on a plurality of structured knowledge data in the knowledge database, and a subject identification model is used to determine a business subject in each of the preprocessed structured data, wherein the preprocessing includes at least one of the following: data cleaning, format conversion, and standardization processing; The multiple structured data describing the same business entity in the multiple knowledge databases are aggregated to obtain the knowledge wide table, wherein each row in the knowledge wide table represents a business entity, and each column represents an attribute information of a business entity.

3. The method according to claim 2, characterized in that Aggregating multiple structured data describing the same business entity in the multiple knowledge databases to obtain the knowledge wide table includes: For each of the business entities, determining first data fields describing the business entity in each of the knowledge databases, and calculating first semantic similarities between the first data fields; determining at least one second data field describing the same attribute of the business entity from a plurality of the first data fields based on the first semantic similarities and the entity type of the business entity, and performing standardization on the second data fields describing the same attribute of the business entity to obtain a unique identification field describing the same attribute of the business entity; Based on the unique identification field, a plurality of structured data describing the same business entity in the plurality of knowledge databases are aggregated to obtain the knowledge wide table.

4. The method according to claim 1, wherein The field information includes at least: a field value, wherein the information entropy of each data field in the knowledge wide table and the semantic similarity between each two data fields are determined, and multiple data fields in the knowledge wide table are merged according to the information entropy and the semantic similarity to obtain multiple standard data fields, including: For each data field in the wide knowledge table, the information entropy of the data field is calculated based on the field information of the data field, wherein the more field values ​​the data field appears in the wide knowledge table and the more dispersed the occurrence frequencies of the various field values ​​are, the greater the information entropy of the data field is; the fewer field values ​​the data field appears in the wide knowledge table and / or the more concentrated the occurrence frequencies of the various field values ​​are, the smaller the information entropy of the data field is; Determine at least one candidate data field whose information entropy is lower than a preset first threshold value from a plurality of data fields in the knowledge wide table, and calculate the semantic similarity between the field information of each candidate data field; At least two of the candidate data fields whose semantic similarity is higher than a preset second threshold are merged to obtain the multiple standard data fields.

5. The method according to claim 1, characterized in that The training process of the weight distribution decision model includes: Constructing an online Q network and a target network for solving the weight allocation strategy, and initializing network weight parameters of the online Q network and the target network; Set up the experience pool, determine the capacity of the experience pool and initialize the priority of each sample in the experience pool; Determine a first preset number of iteration cycles, and iteratively solve the weight allocation strategy and network weight parameters through the following steps: In each iteration cycle, a state is initialized and a second preset number of calculation cycles is determined, wherein the state includes at least: semantic similarity between each of the standard data fields, information entropy of each of the standard data fields, and retrieval usage frequency; In each calculation cycle, the current state is input into the online Q network, the predicted Q values ​​corresponding to different weight allocation strategies are calculated, and the target weight allocation strategy corresponding to the maximum Q value is selected based on the greedy strategy, the target weight allocation strategy is executed and the reward is calculated to obtain a new state, the current state, the target weight allocation strategy, the reward and the new state are used as a sample, and the sample is stored in the experience pool; based on the priority of each sample in the experience pool, multiple samples are sampled and input into a neural network including a bidirectional long short-term memory network, the probability of each sample being sampled, the mean square error loss function and the loss function weight are calculated, the target Q value is determined based on the calculation results and the network weight parameters of the online Q network are updated, the errors of all samples in the experience pool are calculated and the priorities of all samples are updated; updating the network weight parameters of the target network according to the network weight parameters of the online Q network after a third preset number of calculation cycles, wherein the third preset number is less than the second preset number; After the iteration is completed, the obtained target network is used as the weight allocation decision model.

6. The method according to claim 5, characterized in that Input the current state into the online Q network and calculate the predicted Q values ​​corresponding to different weight allocation strategies, including: Determine multiple weight allocation strategies; For each weight allocation strategy, determine the retrieval time cost, retrieval resource cost, retrieval accuracy and user satisfaction corresponding to the weight allocation strategy, and determine the total retrieval cost of the weight allocation strategy based on the retrieval time cost, the retrieval resource cost, the retrieval accuracy and the user satisfaction; The predicted Q values ​​corresponding to the various weight allocation strategies are determined according to the total search costs of the various weight allocation strategies, wherein the weight allocation strategy with a larger total search cost has a smaller predicted Q value.

7. The method according to claim 1, characterized in that Constructing a target search template for the wide knowledge table based on the field weights of the respective standard data fields includes: According to the order of the field weights from large to small, the standard data fields are arranged and combined in sequence to obtain the target retrieval template of the knowledge wide table.

8. A data retrieval optimization system, characterized in that: include: an acquisition module, configured to acquire a wide knowledge table of a target domain, wherein the wide knowledge table includes a plurality of knowledge databases within the target domain, and each of the knowledge databases includes a plurality of structured knowledge data consisting of at least one data field; a field processing module, configured to determine the information entropy of each data field in the wide knowledge table and the semantic similarity between each two data fields, and merge the multiple data fields in the wide knowledge table according to the information entropy and the semantic similarity to obtain multiple standard data fields; A weight allocation module, configured to analyze the field information of each of the standard data fields using a pre-trained weight allocation decision model to obtain a field weight for each of the standard data fields, wherein the weight allocation decision model is trained based on a dual deep Q network algorithm; A construction module is used to construct a target retrieval template of the knowledge wide table according to the field weights of each of the standard data fields.

9. A computer program product, characterized in that include: A computer program, wherein when the computer program is executed by a processor, the data retrieval optimization method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: A memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the data retrieval optimization method according to any one of claims 1 to 7 through the computer program.