A Method for Optimizing the Efficiency of Intelligent Retrieval of Computer Data

By adjusting the radius parameters of the core objects in the DBSCAN algorithm and classifying them based on the continuity and differences of similarity of text data, the problem of poor clustering in traditional data retrieval is solved, and the accuracy and efficiency of data retrieval is improved.

CN120067283BActive Publication Date: 2025-07-11HUNAN INT ECONOMICS UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510549808.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-11
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

When traditional data retrieval methods process large-scale data and complex information, the search range is too large and there is too much unrelated data, which is not conducive to the rapid progress of accurate retrieval. The DBSCAN algorithm uses Euclidean distance, resulting in the incomplete classification of clustering results, which affects the accuracy of retrieval.

Method used

Using a method based on intelligent computer data retrieval, by obtaining the keywords and similarity values of text data, using the DBSCAN algorithm for clustering, adjusting the radius parameters of the core object, and classifying them according to the degree of continuity and differences of similarity of text data.

Benefits of technology

The fine division of text data of different similarity degrees is achieved, and the accuracy of clustering and data retrieval efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067283B_ABST
    Figure CN120067283B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing, and particularly relates to a method for optimizing the intelligent retrieval efficiency of computer data, including: obtaining keywords of computer text data; obtaining the similarity degree between text data according to the keywords; determining core objects according to the DBSCAN algorithm and the similarity degree; screening the sample radius interval according to the core objects, and obtaining the sample similarity continuity degree according to the number difference of sample points included in the sample radius interval; dividing the core objects into a first sample and a second sample according to the sample radius interval; obtaining the difference between the two samples according to the sample radius interval and the distances between the first sample and the second sample; obtaining the necessity of adjusting the radius parameter according to the sample similarity continuity degree and the difference between the two samples, and adjusting the radius parameter; classifying the text data according to the adjusted radius parameter and completing the retrieval of the text data. The present invention realizes the fine division of text data and improves the retrieval efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a method for optimizing the intelligent retrieval efficiency of computer data. Background Art

[0002] In the traditional data retrieval method, when dealing with large-scale data and complex information, the search range is too large and there are too many irrelevant data, which is not conducive to the rapid progress of accurate retrieval. Therefore, in order to improve the retrieval efficiency and optimize the user experience, the DBSCAN algorithm is usually used to improve the retrieval efficiency at present. This algorithm classifies the data before retrieval, so as to search for relevant data more specifically during retrieval, thereby improving the retrieval efficiency. However, the DBSCAN algorithm uses the Euclidean distance as the default distance metric method, which not only cannot comprehensively consider the feature relationships between data, but also causes the problem that the clustering result division is not fine enough due to the large differences between data. For example, when retrieving text data, on the one hand, the Euclidean distance cannot be used to measure the distance of text data. At the same time, the content information contained in different text data is rich and diverse, which leads to the fact that the division is not fine enough when using the DBSCAN algorithm for clustering, and thus inaccurate retrieval occurs during retrieval. Summary of the Invention

[0003] The present invention provides a method for optimizing the intelligent retrieval efficiency of computer data to solve the above problems.

[0004] The method for optimizing the intelligent retrieval efficiency of computer data according to the present invention adopts the following technical solutions:

[0005] On the one hand, an embodiment of the present invention provides a method for optimizing the intelligent retrieval efficiency of computer data, and the method includes the following steps:

[0006] Obtain the text data in the computer data and the keywords of the text data;

[0007] Obtain the similarity degree between text data according to the number of occurrences of keywords in the text data and the similarity value of the keyword text vectors;

[0008] Use the text data as sample points and use the DBSCAN algorithm to cluster the sample points to obtain the core objects during the clustering process, and the metric distance used during the clustering process is the similarity degree between text data;

[0009] Obtain the sample similarity continuity degree of the core object according to the sample radius interval selected by the core object and the quantity difference of the two types of sample points included in the sample radius interval;

[0010] Divide the core object into a first sample and a second sample according to the selected sample radius interval. The distance between the sample points in the first sample and the core object is greater than the sample radius interval, and the distance between the sample points in the second sample and the core object is less than or equal to the sample radius interval.

[0011] Obtain the difference between the first sample and the second sample of the core object according to the selected sample radius interval, the shortest distance between the first sample and the second sample, and the distances from each sample point in the first sample and the second sample to the core object.

[0012] Obtain the necessity of adjusting the radius parameter of the core object according to the degree of continuity of the sample similarity of the core object and the difference between the first sample and the second sample.

[0013] Adjust the radius parameter of the core object according to the necessity of adjusting the radius parameter of the core object.

[0014] Classify the text data according to the adjusted radius parameter of the core object, and complete the text data retrieval according to the classification result of the text data.

[0015] Furthermore, the similarity degree between the text data includes the following specific steps:

[0016] Calculate the product of the average weight of the keywords in the th combination of the keywords of any two text contents and the cosine similarity value of the keyword text vectors in the th combination, which is denoted as the first product of the th combination. Take the mean value of the first products of all combinations as the similarity degree between any two text data.

[0017] Furthermore, the metric distance used in the clustering process is , where exp( ) represents the exponential function with the natural constant as the base, and represents the similarity degree between any two text data.

[0018] Furthermore, the method for obtaining the degree of continuity of the sample similarity of the core object according to the selected sample radius interval and the difference in the number of two types of sample points included in the sample radius interval includes the following specific steps:

[0019] The distance from each sample point in the neighborhood of the core object to the core object is the sample point radius.

[0020] Sort the radii of all sample points in the neighborhood of the core object in ascending order, and calculate the difference between adjacent sample point radii after sorting.

[0021] Select B maximum adjacent sample point radius differences with a preset quantity from the adjacent sample point radius differences, and denote them as the sample radius interval.

[0022] For the B sample radius intervals selected for the i-th core object, calculate the b-th sample radius interval selected for the i-th core object The difference in the number of sample points contained in the two sample points corresponding to the b-th sample radius interval The ratio of the absolute value is used as the first ratio of the b-th sample radius interval, and the inverse proportional normalization value of the mean of the first ratios of all sample radius intervals is used as the adjustment coefficient of the maximum sample point radius within the neighborhood of the i-th core object;

[0023] Obtain the sample similarity continuity degree of the core object according to the adjustment coefficient of the maximum sample point radius within the neighborhood of the core object and the maximum sample point radius within the neighborhood of the core object.

[0024] Furthermore, the specific steps for obtaining the sample similarity continuity degree of the core object according to the adjustment coefficient of the maximum sample point radius within the neighborhood of the core object and the maximum sample point radius within the neighborhood of the core object are as follows:

[0025] The normalization value of the product of the maximum sample point radius within the neighborhood of the i-th core object and the adjustment coefficient of the maximum sample point radius within the neighborhood of the i-th core object is used as the sample similarity continuity degree of the i-th core object.

[0026] Furthermore, the specific steps for obtaining the difference between the first sample and the second sample of the core object according to the selected sample radius interval, the shortest distance between the first sample and the second sample, and the distances from each sample point in the first sample and the second sample to the core object are as follows:

[0027] Under any one of the sample radius intervals selected for the i-th core object, calculate the mean of all the shortest distances in the shortest distance sequence between the first sample and the second sample, and calculate the average distance from each sample point in the first sample to the i-th core object The product of the standard deviation of the distances from each sample point in the first sample to the core object is denoted as the second product, and the average distance from each sample point in the second sample to the core object The product of the standard deviation of the distances from each sample point in the second sample to the core object is denoted as the second product, and the sum of the first product and the second product is used as the first sum value, and the ratio of the mean of all the shortest distances to the first sum value is denoted as the second ratio;

[0028] The normalization value of the maximum value among the second ratios under all the sample radius intervals selected for the i-th core object is used as the difference between the first sample and the second sample of the i-th core object.

[0029] Further, the specific method for obtaining the shortest distance sequence between the first sample and the second sample is as follows:

[0030] Traverse the shortest distances between any two sample points in the first sample and the second sample. The any two sample points are not selected repeatedly. Arrange the shortest distances between any two sample points from low to high to obtain the shortest distance sequence between the first sample and the second sample.

[0031] Further, the obtaining of the necessity for adjusting the radius parameter of the core object according to the sample similarity continuity degree of the core object and the difference between the first sample and the second sample includes the following specific steps:

[0032] Calculate the difference of 1 minus the sample similarity continuity degree of the i-th core object, and denote the product of the difference and the difference between the first sample and the second sample of the i-th core object as the necessity for adjusting the radius parameter of the i-th core object.

[0033] Further, the adjusting of the radius parameter of the core object according to the necessity for adjusting the radius parameter of the core object includes the following specific steps:

[0034] Calculate the difference of 1 minus the necessity for adjusting the radius parameter of the i-th core object that needs to be adjusted, and denote the product of the difference and the preset radius as the radius parameter of the adjusted core object.

[0035] Further, the classifying of the text data according to the radius parameter of the adjusted core object and the completion of the text data retrieval according to the classification result of the text data include the following specific contents:

[0036] Cluster the text data by the radius parameter of the adjusted core object to obtain each cluster of text data after final clustering, and label them;

[0037] Traverse each cluster of text data with labels in sequence, retrieve the keywords of the text data in each cluster of text data. If the text data in the cluster does not contain the keyword to be retrieved, then retrieve the text data in the next cluster.

[0038] The beneficial effects of the technical solution of the present invention are: By performing similarity measurement on the text data and applying the DBSCAN algorithm, the core objects of the text data can be determined. Based on the continuity of the sample similarity within the preset neighborhood of the core object and the analysis of the difference between the high-similarity and low-similarity samples of the core object, the necessity for adjusting the radius parameter of the core object can be determined. According to this analysis result, the radius parameter of the core object can be adaptively adjusted, thereby improving the text data clustering method using the Euclidean distance as the distance metric. This method realizes the fine division of text data with different similarity degrees, improves the accuracy of clustering, and further improves the efficiency of data retrieval. Brief Description of the Drawings

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 It is a flowchart of the steps of a method for optimizing the intelligent retrieval efficiency of computer data according to the present invention. Detailed Embodiments

[0041] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the drawings and preferred embodiments, will detail the specific embodiments, structures, features and effects of a method for optimizing the intelligent retrieval efficiency of computer data according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0043] The following will specifically describe the specific solution of a method for optimizing the intelligent retrieval efficiency of computer data provided by the present invention in conjunction with the drawings.

[0044] Please refer to Figure 1 , which shows a flowchart of the steps of a method for optimizing the intelligent retrieval efficiency of computer data provided by an embodiment of the present invention. The method includes the following steps:

[0045] Step S001: Obtain text data in computer data.

[0046] There are various types of computer data, such as text, audio, video, etc. In this embodiment, taking the intelligent retrieval optimization of text data as an example, the text data is obtained by accessing the database. The Oracle database is used in this embodiment, and the text data stored in the database is news data. In other embodiments, other databases or other text data can be used, such as the Sql database and thesis data, etc. The text data includes the text title and the text content.

[0047] Step S002: Obtain the similarity degree between text data.

[0048] For the text title, keywords in the text title are extracted through TextRank. Considering that there may be multiple keywords in some titles, the top two nodes ranked by TextRank are selected as keywords. At the same time, for the extracted keywords, they are converted into text vectors through the word embedding method, which can obtain the semantic and syntactic relationships between keywords and help improve the accuracy of subsequent similarity analysis. Among them, TextRank and the word embedding method are well-known technologies, and the specific methods are not elaborated in this embodiment.

[0049] Since there may be multiple keywords in the text title, in this embodiment, the similarity measurement between text data is analyzed by considering the weights and similarities between keywords. Specifically, the keywords extracted from the text title in the text data are used to measure the similarity between text data. The similarity between text data is characterized by the cosine similarity of text vectors, and the feature weights between keywords are also considered. For each keyword in the text title, the number of times it appears in the text content is counted, and the ratio of the number of times the current keyword appears to the sum of the number of times all keywords appear is used as the weight corresponding to each keyword in the text title.

[0050] This embodiment gives a method for obtaining the similarity between text data as follows:

[0051] Calculate the product of the average weight of keywords in the m-th combination of keywords of any two text contents and the cosine similarity value of the text vectors of keywords in the n-th combination, denoted as the first product of the m-th combination. The average value of the first products of all combinations is used as the similarity between any two text data. and the -th combination, and record it as the first product of the -th combination. The average value of the first products of all combinations is used as the similarity between any two text data.

[0052] It should be noted that: the calculation formula for the similarity between any two text data is:

[0053]

[0054] In the formula: represents the similarity between any two text data; M represents the number of combinations of keywords in the text titles of any two text data (pairwise combination of keywords in different text titles); represents the average weight of keywords in the m-th combination of keywords of any two text titles; represents the cosine similarity value of the text vectors of keywords in the m-th combination of keywords of any two text titles.

[0055] This calculation method pairs the keywords of any two text titles pairwise, and uses the average weight of the keywords in the combination to Adjustment is made because the weights of different keywords are different, and the confidence levels corresponding to the similarity degrees analyzed in the combination are different. When the weights of the two keywords in the combination are larger, it indicates that the keywords in the current combination are more representative of the two text data, and thus the confidence level of the similarity degree of the text data analyzed is greater. The closer the value of the cosine similarity is to 1, the greater the corresponding similarity degree.

[0056] In another embodiment, keywords in the text content are used to calculate the similarity degree between any two text data. The specific method is to count the number of occurrences of keywords in the text content, and use the ratio of the number of occurrences of the current keyword to the sum of the number of occurrences of all keywords as the weight corresponding to each keyword in the text content. A method for obtaining the similarity degree between text data is given as follows:

[0057]

[0058] In the formula: represents the similarity degree between any two text data; represents the number of combinations of keywords in the text content of any two text data (pairwise combinations of keywords in different text contents); represents the th combination of keywords in any two text contents; represents the average weight of the keywords in the th combination of keywords in any two text contents;

[0059] Thus, the similarity degree between any two text data is obtained.

[0060] Step S003: Determine core objects for the text data according to the DBSCAN algorithm and the similarity degree between text data, and obtain the sample similarity continuity degree of the core objects and the differences between high - and low - similarity samples.

[0061] Regarding each text data as a sample point, cluster all sample points through the DBSCAN algorithm. During the clustering process, the metric distance between any two text data is ; The core points (core objects) in the DBSCAN algorithm refer to the points that contain a sufficient number (greater than or equal to the minimum number of points MinPts) of sample points within a given radius eps (i.e., within the neighborhood of the sample points). These points are considered to be the internal points of the cluster in the DBSCAN algorithm and are density-reachable. The principle of the DBSCAN algorithm and the concept of the core points in the DBSCAN algorithm are well-known, and the specific implementation process will not be elaborated in this embodiment. This embodiment performs subsequent analysis on the text data based on the core points in the DBSCAN algorithm. In this embodiment, the preset neighborhood radius size eps = 6 and the minimum neighborhood sample number threshold MinPts = 5. In other embodiments, the numerical settings can be made according to specific situations, and this embodiment does not make specific limitations.

[0062] Analyze the distribution continuity of the text data for the neighborhood corresponding to the preset radius eps of the core object. The more continuous the distribution of the text data, the better the similarity and continuity performance between the data, and the greater the possibility of corresponding to the same type of text data; the more discontinuous the distribution, the more likely the data may have a break, indicating that there are text data of different categories within the neighborhood of the current core object. At the same time, considering the difference between the high-similarity and low-similarity samples within the neighborhood of the core object on the basis of the similarity and continuity, the farther the distance between the two and the denser their respective distributions, the greater the possibility that there are text data of different categories within the neighborhood of the current core object.

[0063] It should be noted that for the similarity and continuity between data, it is reflected by the continuous degree of sample similarity of the core object. Specifically, taking the distance from each sample point within the neighborhood of the core object to the core point as the sample point radius, the same sample point radius may correspond to multiple sample points (these sample points are all on the circle with the current core point as the center and the radius as the sample point radius). Statistically sort the sample point radii in ascending order (repeated ones are also sorted), and at the same time calculate the difference between adjacent sample point radii. In this embodiment, for the sorted sample point radii, calculate the difference between adjacent sample point radii by subtracting the previous one from the next one, where the last one does not participate in the calculation. From the obtained differences between adjacent sample point radii, screen out B largest differences between adjacent sample point radii, denoted as the sample radius interval. This embodiment will be described with B = 10 as an example. In other embodiments, other values can be used, and this embodiment does not make specific limitations. This embodiment gives a method for obtaining the sample similarity and continuity degree of the core object as follows:

[0064] Calculate the b-th sample radius interval selected for the i-th core object And the difference in the number of sample points contained in the two sample points corresponding to the b-th sample radius interval The absolute value of the ratio is used as the first ratio of the b-th sample radius interval, and the inverse normalization value of the mean of the first ratios of all sample radius intervals is used as the adjustment coefficient of the maximum sample point radius within the neighborhood of the i-th core object;

[0065] The normalized value of the product of the maximum sample point radius within the neighborhood of the i-th core object and the adjustment coefficient of the maximum sample point radius within the neighborhood of the i-th core object is used as the sample similarity continuity degree of the i-th core object.

[0066] It should be noted that: The calculation formula for the sample similarity continuity degree of the i-th core object is:

[0067]

[0068] In the formula: represents the adjustment coefficient of the maximum sample point radius within the neighborhood of the i-th core object.

[0069]

[0070] In the formula: represents the sample similarity continuity degree of the i-th core object; represents the sample point radius within the neighborhood of the i-th core object; represents the maximum sample point radius within the neighborhood of the i-th core object; B represents the B sample radius intervals selected for the i-th core object; represents the b-th selected sample radius interval; the b-th sample radius interval corresponds to two sample points, and there are multiple sample points on the circle (the circle corresponding to the sample point radius with the i-th core as the center) where each of these two sample points is located. The difference in the number of sample points on the circles corresponding to the two sample points is denoted as ; In this embodiment, is briefly recorded as the difference in the number of sample points included in the two sample points corresponding to the b-th sample radius interval, where adding 1 in is to avoid the denominator being 0; max( ) represents taking the maximum value; | | represents taking the absolute value; represents the linear normalization function.

[0071] This calculation method reflects the maximum continuous value of the current core object through The larger its value, the greater the possible sample similarity continuity degree in the end. Because the overall formula is adjusted on the basis of .

[0072] In the formula, is used to determine the overall adjustment coefficient of the formula, can only reflect the maximum continuous value of the current core object, but cannot determine the continuous situation during the process of reaching the maximum continuous value. Therefore, in this embodiment, multiple maximum sample radius intervals are selected and the continuous situation at each sample interval is analyzed. The more continuous each sample interval is, the more continuous the whole is, and the larger the adjustment coefficient value is; Denote the b-th sample radius interval selected. The larger this value is, the more discontinuous the distribution of sample points is here (since the sample points are sorted, the sample radius interval value here can reflect the continuous state); the continuity is further reflected by The smaller the difference in the number of the two types of samples that make up the current sample radius interval is, the more it can indicate that there is a fault here as a whole, that is, the distribution of sample points here is more discontinuous; The larger it is, the more discontinuous it means. By Adjusting the corresponding relationship, so the continuous state is reflected by each sample radius interval. The larger the value is, the better the continuous state is, and the larger the adjustment coefficient value is.

[0073] Furthermore, according to the B sample radius intervals selected above, the core objects are divided into high-similarity samples (denoted as the first samples) and low-similarity samples (denoted as the second samples). Specifically, for the b-th sample radius interval among the B sample radius intervals selected according to the i-th core point, when the distance between any other core point and the i-th core point is greater than the b-th sample radius interval, this core point is a high-similarity sample point of the i-th core point, and all high-similarity sample points form the high-similarity samples of the i-th core point, that is, the high-similarity samples of the i-th core object; when the distance between any other core point and the i-th core point is less than or equal to the b-th sample radius interval, this core point is a low-similarity sample point of the i-th core point, and all low-similarity sample points form the low-similarity samples of the i-th core point, that is, the low-similarity samples of the i-th core object.

[0074] Traverse the shortest distances between any two sample points in the high- and low-similarity samples (the points taken are not repeated, and the shortest distance is taken as the criterion), and arrange them from low to high to obtain a sequence of the shortest distances of the high- and low-similarity samples. For example: calculate the distance between sample point qa and sample point qb in the high- and low-similarity samples. Next time, calculate the distance between any two sample points among the remaining sample points except sample point qa and sample point qb until all sample points in the high- and low-similarity samples are involved in the calculation.

[0075] Under any one of the sample radius intervals selected for the i-th core object, calculate the mean value of all the shortest distances in the shortest distance sequence of the first samples and the second samples, and multiply the average distance from each sample point in the first samples to the i-th core object by the standard deviation of the distances from each sample point in the first samples to the core object , and denote it as the second product. Multiply the average distance from each sample point in the second samples to the core object by the standard deviation of the distances from each sample point in the second samples to the core object The product is denoted as the second product, and the sum of the first product and the second product is taken as the first sum value. The ratio of the mean value of all the shortest distances to the first sum value is denoted as the second ratio;

[0076] The normalized value of the maximum value among the second ratios under all the sample radius intervals selected by the i-th core object is taken as the difference between the first sample and the second sample of the i-th core object.

[0077] It should be noted that: The following is a calculation method for obtaining the difference between high- and low-similarity samples given in this embodiment:

[0078]

[0079] In the formula: represents the difference between the high- and low-similarity samples of the i-th core object, that is, the difference between the first sample and the second sample; B represents the preset number B of sample radius intervals selected by the i-th core object; K represents the number of the shortest distances in the shortest distance sequence of the high- and low-similarity samples (the first sample and the second sample); represents the k-th shortest distance in the shortest distance sequence of the high- and low-similarity samples (the first sample and the second sample); represents the average distance from each sample point in the first sample (high-similarity sample) to the i-th core object; represents the standard deviation of the distances from each sample point in the first sample (high-similarity sample) to the core object; represents the average distance from each sample point in the second sample (low-similarity sample) to the core object; represents the standard deviation of the distances from each sample point in the second sample (high-similarity sample) to the core object; max( ) represents taking the maximum value; represents the linear normalization function.

[0080] This calculation method passes through represents the distance between the high- and low-similarity samples, which is mainly determined by accumulating and averaging the shortest distances corresponding to the high- and low-similarity sample ranges. When the shortest distance is far enough, the actual distance is naturally farther; represents the comprehensive density of the high- and low-similarity samples. Among them, represents the density of the high-similarity samples. Taking the core object as a reference, the smaller the standard deviation, the more consistent the distances from each sample point to the core object, but it cannot explain the specific distance degree, so it is reflected by the average distance. Combining the two together reflects the density of the high-similarity samples; For similarly, it should be supplemented that the analysis of this method mainly considers the overall density and does not consider the density in the internal local area. Finally, the two are accumulated to determine the comprehensive density of the high- and low-similarity samples; Through Reflect the difference between high- and low-similarity samples corresponding to one sample radius interval. When the distance between high- and low-similarity samples is larger, the comprehensive density of high- and low-similarity samples is smaller, indicating that the difference between high- and low-similarity samples corresponding to the current sample radius interval is larger. The maximum difference between high- and low-similarity samples is selected through max() as the difference between high- and low-similarity samples of the current core object.

[0081] Thus, the sample similarity continuity of the core object of the sample data and the difference between high- and low-similarity samples of the core object are obtained.

[0082] Step S004: Obtain the necessity of adjusting the radius parameter of the core object according to the sample similarity continuity of the core object and the difference between high- and low-similarity samples, and adjust the radius parameter of the core object according to the necessity of adjustment.

[0083] The above process analyzes the sample similarity continuity of the core object and the difference between high- and low-similarity samples, and uses the sample similarity continuity of the core object as a weight to adjust the difference between high- and low-similarity samples thereof to judge the necessity of adjusting the radius parameter of the core object.

[0084] Calculate the difference obtained by subtracting the sample similarity continuity of the i-th core object from 1, and record the product of the difference and the difference between the first sample and the second sample of the i-th core object as the necessity of adjusting the radius parameter of the i-th core object.

[0085] Specifically, the calculation formula for the necessity of adjustment is as follows:

[0086]

[0087] In the formula: represents the necessity of adjusting the radius parameter of the i-th core object; represents the sample similarity continuity of the i-th core object; represents the difference between high- and low-similarity samples of the i-th core object.

[0088] This calculation method shows that the smaller the sample similarity continuity of the core object, the greater the possibility that there is a fault in the sample distribution, the greater the necessity of adjusting the final radius parameter, and the greater the determined weight; and at the same time, it is adjusted in combination with the difference between high- and low-similarity samples of the core object.

[0089] Thus, the necessity of adjusting the radius parameter of the core object is obtained. For those greater than or equal to the threshold, their radius parameters are adjusted, otherwise no adjustment is made. In this embodiment, the threshold is set to 0.6, and other implementation manners can select other values according to actual situations, which are not limited in this embodiment. One adjustment method given in this embodiment is:

[0090] Calculate the difference between 1 and the adjustment necessity of the radius parameter of the i-th core object to be adjusted, and denote the product of the difference and the preset radius as the radius parameter of the adjusted core object.

[0091] Specifically, the calculation formula for the radius parameter of the adjusted core object is as follows:

[0092]

[0093] In the formula, represents the radius parameter of the adjusted core object; represents the preset radius; represents the adjustment necessity of the radius parameter of the i-th core object to be adjusted.

[0094] Step S005: Perform final clustering on the text data according to the radius parameter of the adjusted core object, and perform retrieval according to the clustering.

[0095] By adjusting the radius parameter of the core object through the above steps (the radius parameter of the core object that has not been adjusted is equal to the preset radius ), the text data of each cluster after final clustering is obtained accordingly, and it is numbered.

[0096] When performing text data retrieval subsequently, traverse each cluster of text data with numbers in turn, retrieve the keywords of the text data in each cluster of text data. If there are relevant keywords to be retrieved in the cluster, perform text data retrieval within the current cluster, and do not perform text data retrieval in other clusters. If there are no relevant keywords in the current cluster, then perform screening of the next cluster. This method can achieve more targeted search and screening of relevant data within a small range, improving the retrieval efficiency of the data.

[0097] In other embodiments, after retrieving the text data in the current cluster, it is also possible to continue to perform screening retrieval of the next cluster, and finally multiple text data of different cluster types can be retrieved simultaneously.

[0098] So far, this embodiment is completed.

[0099] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for optimizing the intelligent retrieval efficiency of computer data, characterized in that, The method includes the following steps: Obtain the text data in the computer data and the keywords of the text data; Obtain the similarity degree between text data according to the number of occurrences of keywords in the text data and the similarity value of the keyword text vectors; Use the text data as sample points and cluster the sample points using the DBSCAN algorithm to obtain the core objects during the clustering process, and the metric distance used during the clustering process is the similarity degree between text data; Obtain the sample similarity continuity degree of the core object according to the sample radius interval selected based on the core object and the difference in the number of two types of sample points included in the sample radius interval; Divide the core object into a first sample and a second sample according to the selected sample radius interval. For the first sample, the distance between the sample points and the core object is greater than the sample radius interval, and for the second sample, the distance between the sample points and the core object is less than or equal to the sample radius interval; Obtain the difference between the first sample and the second sample of the core object according to the selected sample radius interval, the shortest distance between the first sample and the second sample, and the distances from each sample point in the first sample and the second sample to the core object; Obtain the necessity of adjusting the radius parameter of the core object according to the sample similarity continuity degree of the core object and the difference between the first sample and the second sample; Adjust the radius parameter of the core object according to the necessity of adjusting the radius parameter of the core object; Classify the text data according to the adjusted radius parameter of the core object, and complete the text data retrieval according to the classification result of the text data; The step of obtaining the necessity of adjusting the radius parameter of the core object according to the sample similarity continuity degree of the core object and the difference between the first sample and the second sample includes the following specific steps: Calculate the difference of 1 minus the sample similarity continuity degree of the i-th core object, and record the product of the difference and the difference between the first sample and the second sample of the i-th core object as the necessity of adjusting the radius parameter of the i-th core object.

2. The method for optimizing the intelligent retrieval efficiency of computer data according to claim 1, wherein The similarity degree between text data includes the following specific steps: Calculate the product of the average weight of the keywords in the th combination of the keywords of any two text contents and the cosine similarity value of the keyword text vectors in the th combination, which is denoted as the first product of the th combination. Take the mean value of the first products of all combinations as the similarity degree between any two text data.

3. The method for optimizing the intelligent retrieval efficiency of computer data according to claim 1, wherein The metric distance used in the clustering process is , where exp( ) represents the exponential function with the natural constant as the base, indicating the similarity degree between any two text data.

4. The method for optimizing the intelligent retrieval efficiency of computer data according to claim 1, wherein, The step of obtaining the sample similarity continuity degree of the core object according to the sample radius interval selected based on the core object and the difference in the number of two types of sample points included in the sample radius interval includes the following specific steps: The distance from each sample point in the neighborhood of the core object to the core object is the sample point radius; Sort all the sample point radii in the neighborhood of the core object in ascending order, and calculate the difference between adjacent sample point radii after sorting; Select B maximum adjacent sample point radius differences with a preset quantity from the adjacent sample point radius differences, and record them as the sample radius interval; For the B sample radius intervals selected for the i-th core object, calculate the b-th sample radius interval selected for the i-th core object The difference in the number of sample points contained in the two sample points corresponding to the b-th sample radius interval The ratio of the absolute value is used as the first ratio of the b-th sample radius interval. The inverse normalization value of the mean of the first ratios of all sample radius intervals is used as the adjustment coefficient of the maximum sample point radius within the neighborhood of the i-th core object; Obtain the sample similarity continuity degree of the core object according to the adjustment coefficient of the maximum sample point radius in the neighborhood of the core object and the maximum sample point radius in the neighborhood of the core object.

5. The method for optimizing the intelligent retrieval efficiency of computer data according to claim 4, wherein The step of obtaining the sample similarity continuity degree of the core object according to the adjustment coefficient of the maximum sample point radius in the neighborhood of the core object and the maximum sample point radius in the neighborhood of the core object includes the following specific steps: Use the normalized value of the product of the maximum sample point radius in the neighborhood of the i-th core object and the adjustment coefficient of the maximum sample point radius in the neighborhood of the i-th core object as the sample similarity continuity degree of the i-th core object.

6. The method for optimizing the intelligent retrieval efficiency of computer data according to claim 1, characterized in that, Obtaining the difference between the first sample and the second sample of the core object based on the selected sample radius interval, the shortest distance between the first sample and the second sample, and the distances from each sample point in the first sample and the second sample to the core object, includes the following specific steps: At any sample radius interval selected from the i-th core object, calculate the mean of all the shortest distances in the shortest distance sequence between the first sample and the second sample, and calculate the average distance from each sample point in the first sample to the i-th core object and the standard deviation of the distances from each sample point in the first sample to the core object The product is denoted as the second product. Calculate the average distance from each sample point in the second sample to the core object and the standard deviation of the distances from each sample point in the second sample to the core object The product is denoted as the second product. Take the sum of the first product and the second product as the first sum value, and take the ratio of the mean of all the shortest distances to the first sum value as the second ratio; Taking the normalized value of the maximum value among the second ratios under all sample radius intervals selected for the i-th core object as the difference between the first sample and the second sample of the i-th core object.

7. The method for optimizing the intelligent retrieval efficiency of computer data according to claim 6, characterized in that The specific method for obtaining the shortest distance sequence between the first sample and the second sample is as follows: Traversing the shortest distances between any two sample points in the first sample and the second sample, where the any two sample points are not selected repeatedly, and arranging the shortest distances between any two sample points from low to high to obtain the shortest distance sequence between the first sample and the second sample.

8. The method for optimizing the intelligent retrieval efficiency of computer data according to claim 1, wherein, Adjusting the radius parameter of the core object according to the necessity of adjusting the radius parameter of the core object, includes the following specific steps: Calculating the difference between 1 and the necessity of adjusting the radius parameter of the i-th core object that needs to be adjusted, and taking the product of the difference and the preset radius as the radius parameter of the adjusted core object.

9. The method for optimizing the intelligent retrieval efficiency of computer data according to claim 1, characterized in that Classifying the text data according to the radius parameter of the adjusted core object, and completing the text data retrieval according to the classification result of the text data, including the following specific contents: Clustering the text data through the radius parameter of the adjusted core object to obtain each cluster of text data after final clustering, and numbering them; Traversing each cluster of numbered text data in turn, retrieving the keywords of the text data in each cluster of text data. If the text data in the cluster does not contain the keyword to be retrieved, then retrieving the text data in the next cluster.

Citation Information

Patent Citations

  • An improved algorithm for carrying out abnormal mining on density irregular data based on DBSCAN

    CN109669990A

  • Intelligent screening method and system for approval workflow data based on machine learning

    CN118312656A