Data search tool identification method, device, apparatus and storage medium

By acquiring and analyzing the search record item values ​​and abnormal status values ​​of data search tools, and calculating the corresponding probabilities, the problem of tool identification with low search efficiency in financial databases is solved, and the identification efficiency is improved.

CN116821147BActive Publication Date: 2025-12-26INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310789612.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-12-26
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

Existing financial database data search tools become less efficient when business changes and data volume increases, resulting in a significant waste of manpower and resources for manual screening. Furthermore, existing technologies struggle to efficiently identify inefficient data search tools.

Method used

By obtaining the search record item values ​​of the data search tool, calculating the abnormal state value, and obtaining the first sub-probability and the second sub-probability based on the abnormal state value, the target data search tool is identified using the probability relationship.

Benefits of technology

It enables efficient identification and retrieval of inefficient data search tools, reduces manual intervention, and improves the identification efficiency of data search tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116821147B_ABST
    Figure CN116821147B_ABST
Patent Text Reader

Abstract

The application relates to a data search tool identification method and device and relates to the technical field of big data. The method comprises the following steps: acquiring search records corresponding to data search tools to be identified in a financial database, and acquiring project values corresponding to a plurality of search record projects contained in the search records; acquiring abnormal state values corresponding to the search record projects respectively according to the project values; acquiring first sub-probabilities and second sub-probabilities corresponding to the search record projects respectively based on the abnormal state values; acquiring a first probability that the data search tool is a target data search tool according to the first sub-probabilities, and acquiring a second probability that the data search tool is not the target data search tool according to the second sub-probabilities; and identifying whether the data search tool is the target data search tool according to the size relationship between the first probability and the second probability. The method can efficiently identify whether the data search tool is a target data search tool with low search efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data, and in particular, relates to a data search tool identification method and device, a computer device, a storage medium, and a computer program product. BACKGROUND

[0002] With the development of the technical field of big data, a financial database data search technology has appeared. The technology searches for target data from a financial database through a data search tool corresponding to the financial database.

[0003] In the above technical solution, as financial business changes and the way in which business data volume increases changes, the search efficiency of some data search tools may decrease, and thus inefficient data search tools need to be manually checked. However, manual identification of inefficient data search tools consumes a large amount of manpower and resources, and the identification efficiency of inefficient data search tools is low. SUMMARY

[0004] Therefore, it is necessary to provide a data search tool identification method, device, computer device, computer readable storage medium, and computer program product that can efficiently identify whether a data search tool is a target data search tool with low search efficiency.

[0005] In a first aspect, the present application provides a data search tool identification method. The method comprises:

[0006] obtaining search records corresponding to data search tools to be identified of a financial database, and obtaining item values corresponding to a plurality of search record items included in the search records, respectively;

[0007] obtaining abnormal state values corresponding to the search record items, respectively, according to the item values; the abnormal state values represent whether the search record items are in an abnormal state;

[0008] obtaining first sub-probabilities and second sub-probabilities corresponding to the search record items, respectively, based on the abnormal state values; the first sub-probabilities represent probabilities of the search record items being in the abnormal state values when the data search tools are target data search tools, the second sub-probabilities represent probabilities of the search record items being in the abnormal state values when the data search tools are not the target data search tools, and the target data search tools represent data search tools with search efficiency lower than a preset search efficiency value;

[0009] According to each of the first sub-probabilities, a first probability that the data search tool is the target data search tool is obtained, and according to each of the second sub-probabilities, a second probability that the data search tool is not the target data search tool is obtained;

[0010] According to a size relationship between the first probability and the second probability, whether the data search tool is the target data search tool is identified.

[0011] In one of the embodiments, before the obtaining, further comprising: obtaining historical search records respectively corresponding to a plurality of historical data search tools of the financial database, and historical item values respectively corresponding to a plurality of search record items included in each of the historical search records, and obtaining historical abnormal state values respectively corresponding to the search record items based on the historical item values; obtaining a first number of the historical data search tools being the target data search tool and a second number of the historical data search tools not being the target data search tool based on each of the historical search records; and obtaining the first sub-probability and the second sub-probability respectively corresponding to each of the search record items according to the first number, the second number, and the historical abnormal state values respectively corresponding to the search record items.

[0012] In one of the embodiments, the obtaining the first sub-probability and the second sub-probability respectively corresponding to each of the search record items according to the first number, the second number, and the historical abnormal state values respectively corresponding to the search record items comprises: obtaining a first number of abnormal state values of a current historical abnormal state value of a current search record item corresponding to the historical data search tool in a case that the historical data search tool is the target data search tool, and obtaining a second number of abnormal state values of the current historical abnormal state value in a case that the historical data search tool is not the target data search tool; taking a ratio of the first number of abnormal state values to the first number as the first sub-probability corresponding to the current search record item; and taking a ratio of the second number of abnormal state values to the second number as the second sub-probability corresponding to the current search record item.

[0013] In one of the embodiments, the obtaining the first probability that the data search tool is the target data search tool according to each of the first sub-probabilities and the second probability that the data search tool is not the target data search tool according to each of the second sub-probabilities comprises: obtaining the first probability that the data search tool is the target data search tool based on each of the first sub-probabilities and a first ratio of the first number to the total number of the historical data search tools; and obtaining the second probability that the data search tool is not the target data search tool based on each of the second sub-probabilities and a second ratio of the second number to the total number of the historical data search tools.

[0014] In one of the embodiments, the obtaining the historical abnormal state value corresponding to each of the search record items based on each of the historical item values comprises: obtaining the historical item value corresponding to the current search record item for each of the historical data search tools; dividing the historical item values to obtain two historical item value sets corresponding to the current search record item; obtaining an average value of the historical item values in each of the historical item value sets; and obtaining a set abnormal state value corresponding to each of the historical item value sets based on the average value, and taking the set abnormal state value as the historical abnormal state value corresponding to each of the historical item values in each of the historical item value sets.

[0015] In one of the embodiments, the dividing the historical item values to obtain two historical item value sets corresponding to the current search record item comprises: randomly selecting two historical item values from the historical item values as initial clustering centers; obtaining distance information of each of the historical item values to the initial clustering centers, and dividing each of the historical item values into two initial historical item value sets based on the distance information; obtaining clustering centers corresponding to each of the initial historical item value sets based on the historical item values in each of the initial historical item value sets; and taking the clustering centers as the initial clustering centers, repeating the step of obtaining the distance information of each of the historical item values to the initial clustering centers until the initial clustering centers and the clustering centers are the same, and taking the initial historical item value sets as the historical item value sets corresponding to the current search record item.

[0016] In one of the embodiments, after the obtaining the set abnormal state value corresponding to each of the historical item value sets based on the average value, the method further comprises: obtaining two clustering centers respectively corresponding to the two historical item value sets; dividing the two clustering centers into a target clustering center and a non-target clustering center based on the set abnormal state value; and the target clustering center is the clustering center of the historical item value set with the abnormal state value.

[0017] In one of the embodiments, the acquiring of the abnormal state value corresponding to each of the search record items according to the item value comprises: acquiring a first distance between the item value corresponding to a current search record item and the target cluster center, and a second distance between the item value corresponding to the current search record item and the non-target cluster center; if the first distance is greater than the second distance, the abnormal state value corresponding to the current search record item represents that the current search record item is in an abnormal state; if the first distance is less than or equal to the second distance, the abnormal state value corresponding to the current search record item represents that the current search record item is in a normal state.

[0018] In one of the embodiments, the current search record item is missing one or more historical item values; and the obtaining of the first sub-probability and the second sub-probability corresponding to each of the search record items according to the first quantity, the second quantity and the historical abnormal state value corresponding to each of the search record items comprises: acquiring an initial quantity of the historical abnormal state value corresponding to the current search record item, and adding a preset quantity to the initial quantity to obtain a target quantity corresponding to the initial quantity; and obtaining the first sub-probability and the second sub-probability corresponding to the current search record item based on the first quantity, the second quantity and the target quantity.

[0019] In one of the embodiments, the identifying of whether the data search tool is the target data search tool according to the size relationship between the first probability and the second probability comprises: if the first probability is greater than the second probability, the data search tool is the target data search tool; and if the first probability is less than or equal to the second probability, the data search tool is not the target data search tool.

[0020] In one of the embodiments, after the identifying of whether the data search tool is the target data search tool according to the size relationship between the first probability and the second probability, the method further comprises: in a case where it is identified that the data search tool is the target data search tool, acquiring target data within a search range of the data search tool in the financial database; and constructing a database index for the data search tool to search the financial database based on the target data.

[0021] In a second aspect, the present application further provides a data search tool identification device. The device comprises:

[0022] an item value acquisition module, configured to acquire search records corresponding to a data search tool to be identified in a financial database, and acquire item values corresponding to a plurality of search record items included in the search records;

[0023] an abnormal state value obtaining module, configured to obtain, according to the item value, an abnormal state value corresponding to each of the search record items; the abnormal state value represents whether each of the search record items is in an abnormal state;

[0024] a sub-probability obtaining module, configured to obtain, based on the abnormal state value, a first sub-probability and a second sub-probability corresponding to each of the search record items, which are obtained in advance; the first sub-probability represents a probability of each of the search record items being in the abnormal state value when the data search tool is a target data search tool, the second sub-probability represents a probability of each of the search record items being in the abnormal state value when the data search tool is not the target data search tool; the target data search tool represents a data search tool with a search efficiency lower than a preset search efficiency value;

[0025] a target probability obtaining module, configured to obtain, according to each of the first sub-probabilities, a first probability of the data search tool being the target data search tool, and obtain, according to each of the second sub-probabilities, a second probability of the data search tool not being the target data search tool;

[0026] a data search tool identifying module, configured to identify, according to a size relationship between the first probability and the second probability, whether the data search tool is the target data search tool.

[0027] In a third aspect, the present application further provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0028] obtaining search records corresponding to a data search tool to be identified for a financial database, and obtaining item values corresponding to a plurality of search record items included in the search records;

[0029] obtaining, according to the item values, abnormal state values corresponding to each of the search record items; the abnormal state values represent whether each of the search record items is in an abnormal state;

[0030] obtaining, based on the abnormal state values, a first sub-probability and a second sub-probability corresponding to each of the search record items, which are obtained in advance; the first sub-probability represents a probability of each of the search record items being in the abnormal state value when the data search tool is a target data search tool, the second sub-probability represents a probability of each of the search record items being in the abnormal state value when the data search tool is not the target data search tool; the target data search tool represents a data search tool with a search efficiency lower than a preset search efficiency value;

[0031] According to each of the first sub-probabilities, a first probability that the data searching tool is the target data searching tool is obtained, and according to each of the second sub-probabilities, a second probability that the data searching tool is not the target data searching tool is obtained;

[0032] According to a size relationship between the first probability and the second probability, whether the data searching tool is the target data searching tool is identified.

[0033] In a fourth aspect, the present application further provides a computer readable storage medium. The computer readable storage medium has a computer program stored thereon, and the computer program is executed by a processor to implement the following steps:

[0034] A searching record corresponding to a data searching tool to be identified for a financial database is obtained, and a plurality of searching record items included in the searching record are obtained, and each searching record item corresponds to an item value;

[0035] According to the item value, an abnormal state value corresponding to each searching record item is obtained; the abnormal state value represents whether each searching record item is in an abnormal state;

[0036] Based on the abnormal state value, a first sub-probability and a second sub-probability corresponding to each searching record item are obtained in advance; the first sub-probability represents a probability that each searching record item is in the abnormal state value when the data searching tool is a target data searching tool, the second sub-probability represents a probability that each searching record item is in the abnormal state value when the data searching tool is not the target data searching tool; the target data searching tool represents a data searching tool with a searching efficiency lower than a preset searching efficiency value;

[0037] According to each of the first sub-probabilities, a first probability that the data searching tool is the target data searching tool is obtained, and according to each of the second sub-probabilities, a second probability that the data searching tool is not the target data searching tool is obtained;

[0038] According to a size relationship between the first probability and the second probability, whether the data searching tool is the target data searching tool is identified.

[0039] In a fifth aspect, the present application further provides a computer program product. The computer program product comprises a computer program, and the computer program is executed by a processor to implement the following steps:

[0040] A searching record corresponding to a data searching tool to be identified for a financial database is obtained, and a plurality of searching record items included in the searching record are obtained, and each searching record item corresponds to an item value;

[0041] According to the project value, an abnormal state value corresponding to each of the search record projects is obtained, and the abnormal state value represents whether each of the search record projects is in an abnormal state;

[0042] Based on the abnormal state value, a first sub-probability and a second sub-probability corresponding to each of the search record projects are obtained, the first sub-probability represents a probability of each of the search record projects being in the abnormal state value when the data search tool is a target data search tool, the second sub-probability represents a probability of each of the search record projects being in the abnormal state value when the data search tool is not the target data search tool, and the target data search tool represents a data search tool with a search efficiency lower than a preset search efficiency value;

[0043] According to each of the first sub-probabilities, a first probability that the data search tool is the target data search tool is obtained, and according to each of the second sub-probabilities, a second probability that the data search tool is not the target data search tool is obtained;

[0044] According to a size relationship between the first probability and the second probability, whether the data search tool is the target data search tool is identified.

[0045] The data search tool identification method, device, computer device, storage medium, and computer program product identify a data search tool by obtaining a search record corresponding to the data search tool to be identified for a financial database, obtaining a plurality of search record items included in the search record, and obtaining a project value corresponding to each search record item. According to the project value, an abnormal state value corresponding to each search record item is obtained. The abnormal state value represents whether each search record item is in an abnormal state. Based on the abnormal state value, a first sub-probability and a second sub-probability corresponding to each search record item are obtained in advance. The first sub-probability represents a probability of each search record item being in an abnormal state value when the data search tool is a target data search tool. The second sub-probability represents a probability of each search record item being in an abnormal state value when the data search tool is not a target data search tool. The target data search tool represents a data search tool with a search efficiency lower than a preset search efficiency value. According to each first sub-probability, a first probability that the data search tool is a target data search tool is obtained, and according to each second sub-probability, a second probability that the data search tool is not a target data search tool is obtained. According to the size relationship between the first probability and the second probability, whether the data search tool is a target data search tool is identified. The present application obtains the project value corresponding to each search record item included in the search record corresponding to the data search tool, then obtains the abnormal state value corresponding to each search record item according to the project value, and then obtains the first sub-probability that the data search tool is a target data search tool and the second sub-probability that the data search tool is not a target data search tool based on the abnormal state value. Further, the first probability that the data search tool is a target data search tool is obtained according to the first sub-probability, and the second probability that the data search tool is not a target data search tool is obtained according to the second sub-probability. Finally, whether the data search tool is a target data search tool is identified according to the size relationship between the first probability and the second probability. The present method can efficiently identify whether the data search tool is a target data search tool with low search efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 A flowchart of a data search tool identification method in an embodiment;

[0047] Figure 2 A flowchart of obtaining a first sub-probability and a second sub-probability in an embodiment;

[0048] Figure 3 A flowchart of obtaining a first sub-probability and a second sub-probability in another embodiment;

[0049] Figure 4 A flowchart of obtaining a first probability and a second probability in an embodiment;

[0050] Figure 5A structural block diagram of a data search tool identification device in an embodiment;

[0051] Figure 6 An internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0052] For the purpose, technical solutions and advantages of the present application to be more clearly and obviously understood, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0053] It should be noted that the terms "first" and "second" involved in the embodiments of the present application are only to distinguish similar objects, and do not represent a specific order of the objects. Understandably, "first" and "second" can be interchanged in a specific order or sequence as allowed. It should be understood that the objects distinguished by "first" and "second" can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.

[0054] In an embodiment, as shown in Figure 1 A data search tool identification method is provided, and the embodiment is exemplified by applying the method to a terminal. It should be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and can be implemented through the interaction of the terminal and the server. In the embodiment, the method includes the following steps:

[0055] In step S101, a search record corresponding to a data search tool to be identified for a financial database is obtained, and a plurality of search record items contained in the search record are obtained, and each search record item corresponds to an item value.

[0056] The financial database is a database corresponding to a financial service, and the data search tool to be identified is a data search tool to be identified for search efficiency. The data search tool can be used to search for financial service data of the financial database. The search record refers to a search record when the data search tool searches for the financial database data. The data search tool includes a plurality of search record items. For example, the search record item can be an index screening rate, an index field number, a proportion of the index field number to a total field number, a single data search tool execution time, a time consumption per thousand rows of records, an average execution time, a single execution time, an access frequency, a daily average execution frequency, an average time consumption, an average CPU time consumption (seconds), an average IO time consumption (seconds), an average application time consumption (seconds), an average cluster time consumption (seconds), an average logical read, and a statement scan hit ratio. Finally, the item value is a value of the search record item corresponding to the data search tool. For example, the search record item corresponding to the data search tool is a time consumption per thousand rows of records, and the current item value is a time consumption per thousand rows of records of the data search tool, which is 3 seconds.

[0057] Specifically, the search record corresponding to the data search tool to be identified is obtained from the financial database, and then the item values corresponding to the plurality of search record items are obtained from the search record.

[0058] In step S102, according to the item values, the abnormal state values corresponding to the search record items are obtained. The abnormal state value represents whether the search record item is in an abnormal state.

[0059] The abnormal state value is a value representing whether the item value is abnormal, that is, the abnormal state value represents whether the search record item is in an abnormal state. For example, the abnormal state value can be 1 and 0, 1 indicating that the item value is abnormal, and 0 indicating that the item value is not abnormal. The current item value is a time consumption per thousand rows of records of the data search tool, which is 3 seconds. At this time, the time consumption per thousand rows of records is abnormal, and the abnormal state value corresponding to the time consumption per thousand rows of records of 3 seconds is 1.

[0060] Specifically, according to the corresponding relationship between the obtained item values and the abnormal state values, the abnormal state values corresponding to the search record items can be obtained based on the item values corresponding to the search record items.

[0061] In step S103, based on the abnormal state values, the first sub-probability and the second sub-probability corresponding to the search record items are obtained. The first sub-probability represents the probability of the search record item being in an abnormal state value when the data search tool is a target data search tool. The second sub-probability represents the probability of the search record item being in an abnormal state value when the data search tool is not a target data search tool. The target data search tool represents a data search tool with a search efficiency lower than a preset search efficiency value.

[0062] The first sub-probability is a conditional probability of each search record item being an abnormal state value when the data search tool is the target data search tool, and the second sub-probability is a conditional probability of each search record item being an abnormal state value when the data search tool is not the target data search tool.

[0063] Specifically, according to a correspondence relationship between the abnormal state value and the first sub-probability and the second sub-probability obtained in advance, the first sub-probability and the second sub-probability corresponding to each search record item are obtained based on the abnormal state value.

[0064] In step S104, the first probability that the data search tool is the target data search tool is obtained according to the first sub-probability, and the second probability that the data search tool is not the target data search tool is obtained according to the second sub-probability.

[0065] The first probability is a posterior probability that the data search tool is the target data search tool when the search record item is the item value, and the second probability is a posterior probability that the data search tool is not the target data search tool when the search record item is the item value.

[0066] Specifically, the first probability that the data search tool is the target data search tool is obtained according to the first sub-probability and a prior probability that the data search tool is the target data search tool obtained in advance, and the second probability that the data search tool is not the target data search tool is obtained according to the second sub-probability and a prior probability that the data search tool is not the target data search tool obtained in advance.

[0067] In step S105, whether the data search tool is the target data search tool is identified according to a size relationship between the first probability and the second probability.

[0068] The size relationship is a size relationship between values of the first probability and the second probability.

[0069] Specifically, the size relationship between the first probability and the second probability is compared, and if the first probability is greater than the second probability, the data search tool is the target data search tool with a search efficiency lower than a preset search efficiency value.

[0070] In the data search tool identification method, the search record corresponding to the data search tool to be identified for the financial database is obtained, and the item values corresponding to the multiple search record items contained in the search record are obtained; according to the item values, the abnormal state values corresponding to the multiple search record items are obtained; the abnormal state values represent whether the multiple search record items are in an abnormal state; based on the abnormal state values, the first sub-probability and the second sub-probability corresponding to the multiple search record items are obtained; the first sub-probability represents the probability of the multiple search record items being in the abnormal state values when the data search tool is a target data search tool, and the second sub-probability represents the probability of the multiple search record items being in the abnormal state values when the data search tool is not the target data search tool; the target data search tool represents a data search tool with a search efficiency lower than a preset search efficiency value; according to the first sub-probabilities, the first probability that the data search tool is the target data search tool is obtained, and according to the second sub-probabilities, the second probability that the data search tool is not the target data search tool is obtained; according to the size relationship between the first probability and the second probability, whether the data search tool is the target data search tool is identified. According to the item values corresponding to the multiple search record items contained in the search record corresponding to the data search tool, the abnormal state values corresponding to the multiple search record items are obtained according to the item values, and the first sub-probability that the data search tool is the target data search tool and the second sub-probability that the data search tool is not the target data search tool are obtained based on the abnormal state values. According to the first sub-probability, the first probability that the data search tool is the target data search tool is obtained, and according to the second sub-probability, the second probability that the data search tool is not the target data search tool is obtained. Finally, according to the size relationship between the first probability and the second probability, whether the data search tool is the target data search tool is identified. The method can efficiently identify whether the data search tool is a target data search tool with low search efficiency.

[0071] In one embodiment, as shown in FIG. 1, before obtaining the search record corresponding to the data search tool to be identified for the financial database, the following steps are further included: Figure 2

[0072] Step S201, obtaining the historical search records corresponding to the multiple historical data search tools for the financial database, and the historical item values corresponding to the multiple search record items contained in the historical search records, and obtaining the historical abnormal state values corresponding to the multiple search record items based on the historical item values.

[0073] ​The plurality of historical data searching tools are inventory data searching tools whose searching efficiency has been determined, and the historical searching records are inventory searching records corresponding to the historical data searching tools, wherein one historical data searching tool corresponds to one historical searching record. The historical item value refers to inventory data corresponding to the searching record item, and the historical abnormal state value is a value obtained based on the historical item value and used to represent whether each searching record item in the historical data searching tool is abnormal.

[0074] Number 1 2 3 4 5 6 7 8 9 10 Time per thousand records (seconds) 1 2 3 2 5 1.5 2 4 2 1

[0075] Table 1: Time consumption per thousand record item value table

[0076]

[0077] Table 2: Historical abnormal state value table

[0078] Specifically, for example, the historical item value of time consumption per thousand records as one of the searching record items is shown in Table 1. The number in Table 1 is the number of the historical data searching tool, and each historical data searching tool corresponds to one historical item value of time consumption per thousand records. A plurality of historical searching records corresponding to the plurality of historical data searching tools are obtained, and each historical searching record contains the historical item value of time consumption per thousand records. Then, based on the historical item value of time consumption per thousand records, the historical abnormal state values corresponding to the 10 historical item values of time consumption per thousand records are obtained as shown in Table 2, wherein each historical item value corresponds to one historical abnormal state value. The historical abnormal state value 1 represents that the current historical item value is in an abnormal state, and the historical abnormal state value 0 represents that the current historical item value is not in an abnormal state. Meanwhile, Table 2 also records the historical abnormal state values corresponding to other searching record items such as index screening rate and index field number.

[0079] In step S202, based on the historical searching records, the first number of historical data searching tools that are target data searching tools and the second number of historical data searching tools that are not target data searching tools are obtained.

[0080] The first number is the number of historical data searching tools that are target data searching tools, and the second number is the number of historical data searching tools that are not target data searching tools.

[0081] Specifically, for example, as shown in the last column of Table 2, the historical lookup record contains a pre-obtained mark indicating whether the historical data lookup tool is the target data lookup tool, where mark 1 represents that the historical lookup record is the target data lookup tool, and mark 0 represents that the historical lookup record is not the target data lookup tool. According to the pre-obtained mark, the first number of historical data lookup tools that are the target data lookup tool is 3, and the second number of historical data lookup tools that are not the target data lookup tool is 7.

[0082] In step S203, according to the first number, the second number, and the historical abnormal state value corresponding to each lookup record item, the first sub-probability and the second sub-probability corresponding to each lookup record item are obtained.

[0083] Specifically, the number of current historical abnormal state values corresponding to the current lookup record item is obtained, and based on the number and the first number, the first sub-probability corresponding to the current lookup record item is obtained, and based on the number and the second number, the second sub-probability corresponding to the current lookup record item is obtained.

[0084] In the embodiment, by obtaining the historical item value historical abnormal state value corresponding to each lookup record item, and then obtaining the first number of historical data lookup tools that are the target data lookup tool and the second number of historical data lookup tools that are not the target data lookup tool, according to the first number, the second number, and the historical abnormal state value corresponding to each lookup record item, the first sub-probability and the second sub-probability corresponding to each lookup record item are obtained, which can accurately obtain the first sub-probability and the second sub-probability.

[0085] In one embodiment, as shown in Figure 3 According to the first number, the second number, and the historical abnormal state value corresponding to each lookup record item, the first sub-probability and the second sub-probability corresponding to each lookup record item are obtained, including the following steps:

[0086] In step S301, the first number of abnormal state values of the current historical abnormal state value of the current lookup record item corresponding to the historical data lookup tool in the case where the historical data lookup tool is the target data lookup tool is obtained, and the second number of abnormal state values of the current historical abnormal state value in the case where the historical data lookup tool is not the target data lookup tool is obtained.

[0087] Wherein, the current search record item is any one of the multiple search record items, and the current historical abnormal state value is any one of the two historical abnormal state values, and the first abnormal state value number is the number of the current historical abnormal state value when the historical data search tool is the target data search tool, and the second abnormal state value number is the number of the current historical abnormal state value when the historical data search tool is not the target data search tool.

[0088] Specifically, for example, as shown in Table 2, assuming that the current search record item is the index screening rate, and the current historical abnormal state value is 1, i.e. the historical abnormal state value is the index screening rate abnormality, and the first number of the historical data search tool being the target data search tool is 3, then the first abnormal state value number of the index screening rate abnormality when the historical data search tool is the target data search tool is 2, and the second number of the historical data search tool not being the target data search tool is 7, then the second abnormal state value number of the index screening rate abnormality when the historical data search tool is not the target data search tool is 5.

[0089] In step S302, the ratio of the first abnormal state value number to the first number is taken as the first sub-probability corresponding to the current search record item.

[0090] Specifically, for example, as shown in Table 2, the first abnormal state value number of the index screening rate abnormality is 2, and the first number of the historical data search tool being the target data search tool is 3, then the ratio of the first abnormal state value number to the first number is 2 / 3, i.e. the first sub-probability corresponding to the index screening rate abnormality.

[0091] In step S303, the ratio of the second abnormal state value number to the second number is taken as the second sub-probability corresponding to the current search record item.

[0092] Specifically, for example, as shown in Table 2, the second abnormal state value number of the index screening rate abnormality is 5, and the second number of the historical data search tool being the target data search tool is 7, then the ratio of the second abnormal state value number to the second number is 5 / 7, i.e. the second sub-probability corresponding to the index screening rate abnormality.

[0093] In this embodiment, by obtaining the first abnormal state value number of the current historical abnormal state value of the current search record item when the historical data search tool is the target data search tool, and obtaining the second abnormal state value number of the current historical abnormal state value when the historical data search tool is not the target data search tool, and then based on the ratio of the first abnormal state value number to the first number, and the ratio of the second abnormal state value number to the second number, the first sub-probability and the second sub-probability can be accurately obtained.

[0094] In one embodiment, as shown in FIG. 1, the data searching tool is a data searching tool for a target data searching tool according to each first sub-probability, and the data searching tool is not the target data searching tool according to each second sub-probability, including the following steps: Figure 4

[0095] In step S401, the first probability that the data searching tool is the target data searching tool is obtained based on each first sub-probability and a first ratio of the first number to the total number of historical data searching tools.

[0096] The total number of historical data searching tools is the number of historical data searching tools, for example, the total number of historical data searching tools shown in Table 2 is 10, and the first ratio is the ratio of the first number to the total number of historical data searching tools, that is, the prior probability that the historical data searching tool is the target data searching tool.

[0097] Specifically, for example, as shown in Table 2, the first ratio of the first number to the total number of historical data searching tools is 3 / 10, the product of each first sub-probability and the first ratio is obtained, and the first probability that the data searching tool is the target data searching tool is obtained based on the product.

[0098] In step S402, the second probability that the data searching tool is not the target data searching tool is obtained based on each second sub-probability and a second ratio of the second number to the total number of historical data searching tools.

[0099] The second ratio is the ratio of the second number to the total number of historical data searching tools, that is, the prior probability that the historical data searching tool is not the target data searching tool.

[0100] Specifically, for example, as shown in Table 2, the second ratio of the second number to the total number of historical data searching tools is 7 / 10, the product of each second sub-probability and the second ratio is obtained, and the second probability that the data searching tool is not the target data searching tool is obtained based on the product.

[0101] In this embodiment, by obtaining the first ratio of the first number to the total number of historical data searching tools and the second ratio of the second number to the total number of historical data searching tools, and then obtaining the product of the first ratio and each first sub-probability and the product of the second ratio and each second sub-probability, the first probability that the data searching tool is the target data searching tool and the second probability that the data searching tool is not the target data searching tool can be accurately obtained.

[0102] In one embodiment, based on each historical item value, the historical abnormal state value corresponding to each search record item is obtained, including the following steps:​

[0103] obtaining a historical item value corresponding to the current search record item for each historical data search tool; dividing the historical item value to obtain two historical item value sets corresponding to the current search record item; obtaining an average value of the historical item values in each historical item value set; and obtaining a set abnormal state value corresponding to each historical item value set based on the average value, and taking the set abnormal state value as a historical abnormal state value corresponding to each historical item value in each historical item value set.

[0104] The current search record item can be, for example, time consumption per thousand records, and the historical item value set is two sets of historical item values corresponding to the current search record item, wherein one historical item value cannot exist in both sets of historical item values. The average value is a mathematical average value of the historical item values in the historical item value set, and the set abnormal state value is used to represent whether all historical item values in each historical item value set are abnormal.

[0105] Specifically, for example, the current search record item can be time consumption per thousand records. As shown in Table 1, 10 historical item values of time consumption per thousand records are obtained, which are divided into two historical item value sets. The average value of the historical item values in each historical item value set is obtained, and the set abnormal state value of the historical item value set with a larger average value is 1, wherein the set abnormal state value 1 represents that all historical item values in each historical item value set are abnormal, that is, the historical abnormal state value corresponding to each historical item value in the historical item value set with a larger average value is 1.

[0106] In this embodiment, by dividing the historical item values of the current search record item into two sets and then comparing the average values of the two sets, the historical abnormal state values corresponding to each historical item value in the historical item value set can be accurately obtained through the comparison result of the average values of the two sets.

[0107] In one embodiment, dividing the historical item values to obtain two historical item value sets corresponding to the current search record item includes the following steps:

[0108] two initial historical item value sets based on the distance information; obtaining clustering centers corresponding to the initial historical item value sets based on the historical item values in the initial historical item value sets; repeating the step of obtaining the distance information of each historical item value to the initial clustering center until the initial clustering center and the clustering center are the same, and taking the initial historical item value set as the historical item value set corresponding to the current search record item.

[0109] wherein the initial clustering center is any two historical item values, the distance information is the distance of each historical item value to the initial clustering center, the initial historical item value set is two sets of historical item values divided based on the initial clustering center, and the clustering center is the average value of the elements in the initial historical item value set.

[0110] Specifically, from the historical item values corresponding to the current search record item, any two historical item values are selected as initial clustering centers, the distance of each historical item value to the two initial clustering centers is calculated, if the distance of the current historical item value to the current initial clustering center is small, the current historical item value is divided into the initial historical item value set corresponding to the initial clustering center, then the average value of the historical item values in each initial historical item value set is calculated, and the average value is taken as the clustering center corresponding to each initial historical item value set. The clustering center is taken as the initial clustering center, and the step of obtaining the distance information of each historical item value to the initial clustering center is repeated until the initial clustering center and the clustering center are the same, and the initial historical item value set is taken as the historical item value set corresponding to the current search record item.

[0111] In this embodiment, by performing clustering analysis on the historical item values corresponding to the current search record item, two historical item value sets corresponding to the current search record item can be accurately obtained.

[0112] In one embodiment, after obtaining the set abnormal state value corresponding to each historical item value set based on the average value, the following steps are further included:

[0113] obtaining two clustering centers corresponding to the two historical item value sets respectively; based on the set abnormal state value, dividing the two clustering centers into a target clustering center and a non-target clustering center; the target clustering center is the clustering center of the historical item value set with the abnormal state of the set abnormal state value.

[0114] wherein the target clustering center is the clustering center corresponding to the historical item value set with the abnormal state of the set abnormal state value.

[0115] Specifically, the cluster center of the set of historical item value sets with the abnormal state value as the abnormal state is taken as the target cluster center, and the cluster center of the set of historical item value sets with the abnormal state value not as the abnormal state is taken as the non-target cluster center.

[0116] In the embodiment, the cluster center as the non-target cluster center can be accurately obtained by taking the cluster center of the set of historical item value sets with the abnormal state value as the abnormal state as the target cluster center and taking the cluster center of the set of historical item value sets with the abnormal state value not as the abnormal state as the non-target cluster center.

[0117] In one embodiment, according to the item value, the abnormal state value corresponding to each search record item is obtained, including the following steps:

[0118] The first distance between the item value corresponding to the current search record item and the target cluster center and the second distance between the item value corresponding to the current search record item and the non-target cluster center are obtained; if the first distance is greater than the second distance, the abnormal state value corresponding to the current search record item represents that the current search record item is in the abnormal state; if the first distance is less than or equal to the second distance, the abnormal state value corresponding to the current search record item represents that the current search record item is in the normal state.

[0119] The first distance is the distance between the item value and the target cluster center, and the second distance is the distance between the item value and the non-target cluster center.

[0120] Specifically, the first distance and the second distance between the item value corresponding to the current search record item and the target cluster center and the non-target cluster center are obtained; if the first distance is greater than the second distance, the current search record item is in the abnormal state; if the first distance is less than or equal to the second distance, the current search record item is in the normal state.

[0121] In the embodiment, the first distance and the second distance between the item value corresponding to the current search record item and the target cluster center and the non-target cluster center are obtained, and then the first distance and the second distance are compared to obtain the abnormal state value corresponding to the current search record item.

[0122] In one embodiment, the current search record item is missing one or more historical item values; according to the first number, the second number, and the historical abnormal state value corresponding to each search record item, the first sub-probability and the second sub-probability corresponding to each search record item are obtained, including the following steps:

[0123] An initial number of historical abnormal state values corresponding to the current search record item is obtained, and the initial number is added to a preset number to obtain a target number corresponding to the initial number; based on the first number, the second number, and the target number, a first sub-probability and a second sub-probability corresponding to the current search record item are obtained.

[0124] The initial number is the number of historical abnormal state values corresponding to the current search record item, and the target number is an optimized number of historical abnormal state values after the initial number is added to the preset number.

[0125] Specifically, when the current search record item is missing one or more historical item values, in order to avoid the values of the first sub-probability and the second sub-probability being zero, the initial number is added to the preset number to obtain the target number, and then based on the first number, the second number, and the target number, the first sub-probability and the second sub-probability corresponding to the current search record item are obtained.

[0126] In this embodiment, by adding the initial number to the preset number to obtain the target number, and then based on the first number, the second number, and the target number, the first sub-probability and the second sub-probability corresponding to the current search record item are obtained, the case that the values of the first sub-probability and the second sub-probability are zero when the current search record item is missing one or more historical item values can be avoided.

[0127] In one embodiment, according to the size relationship between the first probability and the second probability, whether the data search tool is the target data search tool is identified, including the following steps:

[0128] If the first probability is greater than the second probability, the data search tool is the target data search tool; if the first probability is less than or equal to the second probability, the data search tool is not the target data search tool.

[0129] Specifically, if the first probability is greater than the second probability, the data search tool is the target data search tool, otherwise, the data search tool is not the target data search tool.

[0130] In this embodiment, by comparing the size of the first probability and the second probability, whether the data search tool is the target data search tool can be accurately identified.

[0131] In one embodiment, after identifying whether the data search tool is the target data search tool according to the size relationship between the first probability and the second probability, the following steps are further included:

[0132] In the case of identifying that the data search tool is the target data search tool, target data in the search range of the data search tool in the financial database is obtained; based on the target data, a database index of the data search tool searching the financial database is constructed.

[0133] Wherein, the target data is data that the data search tool can search in the financial database, and the database index is an index used to search the financial database.

[0134] Specifically, in the case that the data search tool is a target data search tool, the target data corresponding to the data search tool in the financial database is acquired, and then the database index of the financial database is constructed based on the target data.

[0135] In the embodiment, by identifying whether the data search tool is a target data search tool, in the case that the data search tool is a target data search tool, the target data is acquired, and the database index of the financial database is constructed, whether the database index of the financial database needs to be reconstructed can be accurately determined.

[0136] In one application embodiment, a low-efficiency data search tool identification method is provided, the low-efficiency data search tool being a data search tool with efficiency lower than a preset efficiency value, and the method specifically includes the following steps:

[0137] 1. Obtain a plurality of feature items in the search record of the data search tool, the feature items including but not limited to: index screening rate, index field number, index field number proportion of total field number, single data search tool execution duration, time consumption per thousand rows of records, average execution time, single execution time, access frequency, daily average execution frequency, average time consumption, average CPU time consumption (seconds), average IO time consumption (seconds), average application time consumption (seconds), average cluster time consumption (seconds), average logical read, and statement scan hit ratio, and then obtain the feature values corresponding to the feature items. Taking the time consumption per thousand rows of records as an example, the project value corresponding to the time consumption per thousand rows of records is shown in Table 3:

[0138] Number 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 Time per thousand records (seconds) 1 2 3 2 5 1.5 2 4 2 1 3 2 1 2 6 5 1.5 1

[0139] Table 3- Time consumption per thousand rows of records project value table

[0140] Here, the sample numbered as the data search tool is referred to as x i In the historical transaction data (the data has been appropriately deformed) sample set D collected above, D = {x1, x2,..., x m} contains m samples, and the cluster number is taken as 2, that is, the cluster division C = {C1, C2}, and two samples x5, x 17 are randomly selected as initial means, that is, μ1 = 5 and μ2 = 1.5.

[0141] Examine the sample x1 = 1, and the Euclidean distances of x1 from the current cluster centers x5 and x 17 are 4 and 0.5 respectively, so x1 will be classified into cluster C2. Similarly, after examining all the samples in the data set, the current cluster division is obtained as

[0142] C1 = {x5, x8, x 15 , x 16};

[0143] C2 = {x1, x2, x3, x4, x6, x7, x9, x 10 , x 11 , x 12 , x 13 , x 14 , x 17 , x 18};

[0144] Thus, new means (cluster centers) μ1' = 5 and μ'2 = 1.8 can be obtained from C1 and C2 respectively, and the new means are updated as the current means, and then the above process is repeatedly executed until the mean result no longer changes, and then the algorithm stops, and the final cluster division is obtained. In combination with the data characteristics, the cluster with a smaller mean value is marked as a normal value (marked as 0), and the cluster with a larger mean value is marked as an abnormal value (marked as 1).

[0145] 2. Probability statistical analysis is performed on the marked values of all feature items, a naive Bayes method is used in combination with a Laplace correction method, and a prediction model of the inefficient data search tool is established. The following takes the data set of Table 2 to demonstrate the process method. Here, 6 attributes (x1, x2, …, x6) are selected as examples, and 0 represents "no" and 1 represents "yes" in the corresponding attribute values. First, the prior probability P(c) of whether the target data search tool is an inefficient data search tool is calculated, then the conditional probability P(x|c) of whether each attribute x is an inefficient data search tool is calculated (that is, the probability of the occurrence of x under the condition of c, for example: the conditional probability P(x = 1 index screening rate abnormal IP | c = 1 inefficient data search tool) is the probability of the occurrence of index screening abnormality under the condition of the occurrence of an inefficient data search tool), and finally the posterior probability P(c|x) of whether the inefficient data search tool is calculated.

[0146] Based on Bayes' theorem, it can be known that

[0147] P(x) is a factor for normalization. For a given attribute x, the normalization factor P(x) is irrelevant to whether it is an inefficient data search tool, so the problem of predicting the posterior probability P(c|x) of whether it is an inefficient data search tool is converted into how to estimate the prior probability P(c) and the conditional probability P(x|c) based on the training data. However, the conditional probability P(x|c) is the joint probability on all attributes, and it is difficult to directly estimate it from limited samples. Therefore, the naive Bayes classifier is used in the scheme, and it is assumed that all attributes are independent, so the posterior probability P(c|x) of whether the target data search tool is an inefficient data search tool can be rewritten as:

[0148]

[0149] where d is the number of attributes, x i is the value of x on the i-th attribute, and is calculated as follows:

[0150] First, estimate the prior probability P(c) of the inefficient data search tool. According to the data in Table 2, P(c = 1 is an inefficient data search tool) = 3 / 10, P(c = 0 is not an inefficient data search tool) = 7 / 10. Then, estimate the conditional probability P(x i |c) for each attribute.

[0151]

[0152]

[0153]

[0154]

[0155]

[0156]

[0157]

[0158]

[0159] Thus, we have:

[0160]

[0161]

[0162] Since 0.026 > 0.00988, the naive Bayes classifier will judge the test sample (No. 1) as not an inefficient data search tool.

[0163] 3. If a certain feature attribute value has never appeared in the training set, the conditional probability of this feature attribute is 0. According to the following formula, the probability value of the posterior probability P is 0, which is obviously unreasonable. In order to avoid the information carried by the event attribute not occurring in the historical data set, a "smoothing" process is needed when estimating the probability value, so that the predicted result is more close to the actual value. Here, the "Laplace correction method" is adopted: add 1 to each count, so that zero does not occur. To balance this, the number of possible values of this attribute is added to the denominator, so the whole value will never be greater than 1.

[0164] For example, assuming that the sample set P (index filtering rate normal | inefficient data search tool) = 0 / 3, at this time, it is corrected by the "Laplace correction method": Using the Naive Bayes classifier learning method combined with the Laplace correction method can avoid the situation that the estimated probability is 0 due to insufficient training data set samples. When the training data set continues to grow, the probability value after the Laplace correction method also gradually tends to the actual probability value.

[0165] After training by the above large sample data, all the probability estimates involved are stored as a probability prediction model for finally judging whether the data search tool is inefficient, i.e. the inefficient data search tool prediction model. The inefficient data search tool prediction model is used to predict the data features of the data search tool to be identified, and an alarm information is generated if the data search tool is identified as inefficient.

[0166] The above embodiment obtains the inventory data of the search data corresponding to the data search tool, then obtains the item values of each feature item in the inventory data, performs clustering analysis on these item values, obtains the abnormal state values corresponding to each item value, and finally calculates the probability that the data search tool is an inefficient data search tool when a certain feature item is an abnormal state value. Then, whether the current data search tool to be identified is an inefficient data search tool is identified through these probabilities. The method can efficiently identify whether the data search tool is an inefficient data search tool.

[0167] It should be understood that although each step in the flowchart involved in the above embodiments is shown in sequence according to the arrow, the steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of the steps is not strictly limited in sequence, and the steps can be executed in other orders. Moreover, at least some of the steps in the flowchart involved in the above embodiments can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the steps or stages is not necessarily sequential, but can be alternately executed with at least some of the other steps or steps or stages in the other steps.

[0168] Based on the same inventive concept, the embodiments of the present application also provide a data search tool identification device for implementing the above-mentioned data search tool identification method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more data search tool identification device embodiments provided below can refer to the limitations of the data search tool identification method in the above text, which will not be repeated here.

[0169] In one embodiment, as shown in Figure 5 A data search tool identification device is provided, comprising: a project value acquisition module 501, an abnormal state value acquisition module 502, a sub-probability acquisition module 503, a target probability acquisition module 504, and a data search tool identification module, wherein:

[0170] The project value acquisition module 501 is configured to acquire the search records corresponding to the data search tool to be identified in the financial database, and acquire the project values corresponding to the plurality of search record projects contained in the search records, respectively.

[0171] The abnormal state value acquisition module 502 is configured to acquire the abnormal state values corresponding to the search record projects, respectively, according to the project values; the abnormal state values represent whether the search record projects are in an abnormal state.

[0172] The sub-probability acquisition module 503 is configured to acquire the first sub-probability and the second sub-probability corresponding to each search record project, respectively, based on the abnormal state values; the first sub-probability represents the probability of the search record project being in an abnormal state value under the condition that the data search tool is a target data search tool, and the second sub-probability represents the probability of the search record project being in an abnormal state value under the condition that the data search tool is not a target data search tool; the target data search tool represents a data search tool with a search efficiency lower than a preset search efficiency value.

[0173] The target probability obtaining module 504 is configured to obtain, according to each first sub-probability, a first probability that the data search tool is the target data search tool, and obtain, according to each second sub-probability, a second probability that the data search tool is not the target data search tool.

[0174] The data search tool identification module 505 is configured to identify, according to a size relationship between the first probability and the second probability, whether the data search tool is the target data search tool.

[0175] In one of the embodiments, the item value obtaining module 501 is further configured to obtain historical search records corresponding to a plurality of historical data search tools of the financial database respectively, and historical item values corresponding to a plurality of search record items included in each historical search record respectively, and obtain historical abnormal state values corresponding to each search record item based on the historical item values; obtain a first number of historical data search tools that are the target data search tool and a second number of historical data search tools that are not the target data search tool based on the historical search records; and obtain the first sub-probability and the second sub-probability corresponding to each search record item respectively according to the first number, the second number, and the historical abnormal state values corresponding to each search record item.

[0176] In one of the embodiments, the item value obtaining module 501 is further configured to obtain a first number of abnormal state values of a current historical abnormal state value of a current search record item corresponding to the historical data search tool in a case where the historical data search tool is the target data search tool, and obtain a second number of abnormal state values of the current historical abnormal state value in a case where the historical data search tool is not the target data search tool; take a ratio of the first number of abnormal state values to the first number as the first sub-probability corresponding to the current search record item; and take a ratio of the second number of abnormal state values to the second number as the second sub-probability corresponding to the current search record item.

[0177] In one of the embodiments, the target probability obtaining module 504 is further configured to obtain the first probability that the data search tool is the target data search tool based on each first sub-probability and a first ratio of the first number to a total number of historical data search tools; and obtain the second probability that the data search tool is not the target data search tool based on each second sub-probability and a second ratio of the second number to the total number of historical data search tools.

[0178] In one of the embodiments, the item value obtaining module 501 is further configured to obtain historical item values corresponding to the current search record item from each historical data search tool; divide the historical item values to obtain two historical item value sets corresponding to the current search record item; obtain average values of the historical item values in each historical item value set; and obtain set abnormal state values corresponding to each historical item value set based on the average values, and take the set abnormal state values as historical abnormal state values corresponding to each historical item value in each historical item value set.

[0179] In one of the embodiments, the item value obtaining module 501 is further configured to randomly select two historical item values from the historical item values as initial clustering centers; obtain distance information of each historical item value to the initial clustering centers, and divide each historical item value into two initial historical item value sets based on the distance information; obtain clustering centers corresponding to each initial historical item value set based on the historical item values in each initial historical item value set; repeat the step of obtaining the distance information of each historical item value to the initial clustering centers until the initial clustering centers and the clustering centers are the same, and take the initial historical item value sets as the historical item value sets corresponding to the current search record item.

[0180] In one of the embodiments, the item value obtaining module 501 is further configured to obtain two clustering centers corresponding to the two historical item value sets respectively; divide the two clustering centers into a target clustering center and a non-target clustering center based on the set abnormal state values; and the target clustering center is the clustering center of the historical item value set with the abnormal state.

[0181] In one of the embodiments, the abnormal state value obtaining module 502 is further configured to obtain a first distance between the item value corresponding to the current search record item and the target clustering center, and a second distance between the item value corresponding to the current search record item and the non-target clustering center; if the first distance is greater than the second distance, the abnormal state value corresponding to the current search record item represents that the current search record item is in the abnormal state; and if the first distance is less than or equal to the second distance, the abnormal state value corresponding to the current search record item represents that the current search record item is in the normal state.

[0182] In one of the embodiments, the item value obtaining module 501 is further configured to obtain an initial number of historical abnormal state values corresponding to the current search record item, and add a preset number to the initial number to obtain a target number corresponding to the initial number; and obtain a first sub-probability and a second sub-probability corresponding to the current search record item based on the first number, the second number and the target number.

[0183] In one of the embodiments, the data search tool identification module 505 is further configured to identify the data search tool as a target data search tool if the first probability is greater than the second probability, and identify the data search tool as a non-target data search tool if the first probability is less than or equal to the second probability.

[0184] In one of the embodiments, the data search tool identification module 505 is further configured to, in a case where the data search tool is identified as the target data search tool, acquire target data within a search range of the data search tool in the financial database, and construct a database index of the financial database searched by the data search tool based on the target data.

[0185] The modules in the data search tool identification apparatus can be implemented by software, hardware, or a combination thereof. The modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in the computer device in software form, so as to be invoked and executed by the processor to perform operations corresponding to the modules.

[0186] In one of the embodiments, a computer device is provided, which can be a terminal. An internal structure diagram of the computer device can be as shown in FIG. 8. Figure 6 The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved through WIFI, mobile cellular network, NFC (near field communication), or other technologies. The computer program is executed by the processor to implement a data search tool identification method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball, or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0187] Those skilled in the art can understand that Figure 6 The structure shown in FIG. 8 is only a block diagram of part of the structure related to the scheme of the present application, and does not limit the computer device to which the scheme of the present application is applied. Specifically, the computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0188] In an embodiment, a computer device is also provided, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above method embodiments when executing the computer program.

[0189] In an embodiment, a computer readable storage medium is provided, storing a computer program, and the computer program implementing the steps in the above method embodiments when executed by a processor.

[0190] In an embodiment, a computer program product is provided, comprising a computer program, and the computer program implementing the steps in the above method embodiments when executed by a processor.

[0191] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties.

[0192] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0193] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0194] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A data search tool identification method, characterized by, The method comprises: obtaining a search record corresponding to a data search tool to be identified for a financial database, and obtaining item values corresponding to a plurality of search record items contained in the search record; obtaining an abnormal state value corresponding to each of the search record items according to the item values; the abnormal state value represents whether each of the search record items is in an abnormal state; obtaining a first sub-probability and a second sub-probability corresponding to each of the search record items according to the abnormal state value; the first sub-probability represents a probability of each of the search record items being in the abnormal state value when the data search tool is a target data search tool, and the second sub-probability represents a probability of each of the search record items being in the abnormal state value when the data search tool is not the target data search tool; the target data search tool represents a data search tool with a search efficiency lower than a preset search efficiency value; obtaining a first probability that the data search tool is the target data search tool according to each of the first sub-probabilities, and obtaining a second probability that the data search tool is not the target data search tool according to each of the second sub-probabilities; identifying whether the data search tool is the target data search tool according to a size relationship between the first probability and the second probability; wherein, before the obtaining of the search record corresponding to the data search tool to be identified for the financial database, the method further comprises: obtaining historical search records corresponding to a plurality of historical data search tools for the financial database, and historical item values corresponding to a plurality of search record items contained in each of the historical search records, and obtaining historical abnormal state values corresponding to each of the search record items based on each of the historical item values; obtaining a first number of the historical data search tools being the target data search tool and a second number of the historical data search tools not being the target data search tool based on each of the historical search records; obtaining a first sub-probability and a second sub-probability corresponding to each of the search record items according to the first number, the second number, and the historical abnormal state values corresponding to each of the search record items.

2. The method of claim 1, wherein, The obtaining of the first sub-probability and the second sub-probability corresponding to each of the search record items according to the first number, the second number, and the historical abnormal state values corresponding to each of the search record items comprises: obtaining a first number of abnormal state values of a current historical abnormal state value of a current search record item corresponding to the historical data search tools when the historical data search tools are the target data search tools, and obtaining a second number of abnormal state values of the current historical abnormal state value when the historical data search tools are not the target data search tools; taking a ratio of the first number of abnormal state values to the first number as the first sub-probability corresponding to the current search record item; taking a ratio of the second number of abnormal state values to the second number as the second sub-probability corresponding to the current search record item.

3. The method of claim 1, wherein, The first probability that the data searching tool is the target data searching tool is obtained based on each first sub-probability and a first ratio of the first number to the total number of the historical data searching tools. The second probability that the data searching tool is not the target data searching tool is obtained based on each second sub-probability and a second ratio of the second number to the total number of the historical data searching tools. The historical abnormal state value corresponding to each searching record item is obtained based on each historical item value, including:

4. The method of claim 1, wherein, An historical item value corresponding to a current searching record item is obtained for each historical data searching tool. The historical item values are divided to obtain two historical item value sets corresponding to the current searching record item. An average value of the historical item values in each historical item value set is obtained. Based on the average value, a set abnormal state value corresponding to each historical item value set is obtained, and the set abnormal state value is taken as the historical abnormal state value corresponding to each historical item value in each historical item value set. The historical item values are divided to obtain two historical item value sets corresponding to the current searching record item, including:

5. The method of claim 4, wherein, Two historical item values are randomly selected as initial clustering centers from the historical item values. Distance information of each historical item value to the initial clustering centers is obtained, and each historical item value is divided into two initial historical item value sets based on the distance information. Based on the historical item values in each initial historical item value set, a clustering center corresponding to each initial historical item value set is obtained. The clustering center is taken as an initial clustering center, and the step of obtaining distance information of each historical item value to the initial clustering center is repeated until the initial clustering center and the clustering center are the same, and the initial historical item value set is taken as the historical item value set corresponding to the current searching record item. After the set abnormal state value corresponding to each historical item value set is obtained based on the average value, the following steps are further included:

6. The method of claim 4, wherein, Two clustering centers corresponding to the two historical item value sets are obtained. Based on the set abnormal state value, the two clustering centers are divided into a target clustering center and a non-target clustering center; the target clustering center is the clustering center of the historical item value set with an abnormal state. The abnormal state value corresponding to each searching record item is obtained based on the item value, including:

7. The method of claim 6, wherein, A first distance between the item value corresponding to the current searching record item and the target clustering center, and a second distance between the item value corresponding to the current searching record item and the non-target clustering center are obtained. If the first distance is greater than the second distance, the abnormal state value corresponding to the current searching record item represents that the current searching record item is in an abnormal state. ​ If the first distance is less than or equal to the second distance, the abnormal state value corresponding to the current search record item represents that the current search record item is in a normal state.

8. The method of claim 1, wherein, The current search record item is missing one or more historical item values; and the first sub-probability and the second sub-probability corresponding to each search record item are obtained based on the first number, the second number, and the historical abnormal state values corresponding to each search record item, including: An initial number of historical abnormal state values corresponding to the current search record item is obtained, and a preset number is added to the initial number to obtain a target number corresponding to the initial number; The first sub-probability and the second sub-probability corresponding to the current search record item are obtained based on the first number, the second number, and the target number.

9. The method of claim 1, wherein, The size relationship between the first probability and the second probability is used to identify whether the data search tool is the target data search tool, including: If the first probability is greater than the second probability, the data search tool is the target data search tool; If the first probability is less than or equal to the second probability, the data search tool is not the target data search tool.

10. The method of claim 1, wherein, After identifying whether the data search tool is the target data search tool based on the size relationship between the first probability and the second probability, the method further includes: In a case where it is identified that the data search tool is the target data search tool, target data within a search range of the data search tool in the financial database is obtained; Based on the target data, a database index for searching the financial database by the data search tool is constructed.

11. A data lookup tool identification apparatus, comprising: The apparatus includes: An item value obtaining module configured to obtain search records corresponding to a data search tool to be identified for a financial database, and obtain item values corresponding to a plurality of search record items included in the search records; An abnormal state value obtaining module configured to obtain abnormal state values corresponding to each search record item based on the item values; the abnormal state values represent whether each search record item is in an abnormal state; A sub-probability obtaining module configured to obtain first sub-probabilities and second sub-probabilities corresponding to each search record item based on the abnormal state values; the first sub-probabilities represent probabilities of each search record item being in the abnormal state values in a case where the data search tool is a target data search tool, and the second sub-probabilities represent probabilities of each search record item being in the abnormal state values in a case where the data search tool is not the target data search tool; the target data search tool represents a data search tool with a search efficiency lower than a preset search efficiency value; A target probability obtaining module configured to obtain a first probability that the data search tool is the target data search tool based on each first sub-probability, and obtain a second probability that the data search tool is not the target data search tool based on each second sub-probability; and The target probability obtaining module configured to obtain a first probability that the data search tool is the target data search tool based on each first sub-probability, and obtain a second probability that the data search tool is not the target data search tool based on each second sub-probability; and The data search tool identification module is configured to identify whether the data search tool is the target data search tool according to a size relationship between the first probability and the second probability. Before the obtaining of the search record corresponding to the data search tool to be identified for the financial database, the apparatus further comprises: obtaining a plurality of historical search records respectively corresponding to a plurality of historical data search tools for the financial database, and a plurality of historical item values respectively corresponding to a plurality of search record items included in each of the historical search records, and obtaining a historical abnormal state value corresponding to each of the search record items based on each of the historical item values; obtaining a first number of the historical data search tools that are the target data search tool and a second number of the historical data search tools that are not the target data search tool based on each of the historical search records; and obtaining a first sub-probability and a second sub-probability respectively corresponding to each of the search record items according to the first number, the second number, and the historical abnormal state value corresponding to each of the search record items.

12. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor, when executing the computer program, implements the steps of the method of any one of claims 1 to 10.

13. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 10.

14. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 10. The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 10.

Citation Information

Patent Citations

  • Strict mode setting method and tool for front-end item and computer equipment

    CN115061675A

  • Systems and methods for providing faster data access using lookup and relationship tables

    US20230145273A1