A data query method based on representative skyline point selection

By using the STV method in the data query method and calculating the number of dominant points and similarity of skyline points, the problem of too many result sets and insufficient representativeness in skyline query is solved, and a more accurate and stable selection of representative skyline points is achieved.

CN119377491BActive Publication Date: 2025-05-06ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411976069.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In the prior art, when processing large-scale data sets, especially in high-dimensional spaces, skyline queries may produce a large number of result sets, and there are problems of instability and underrepresentation in the selection of representative skyline points.

Method used

A data query method based on a single transferable voting (STV) method is used to calculate the number of dominant points and similarity of each skyline point, a preference sequence is generated, and a vote is conducted to select representative skyline points, ensuring that the selected points are both proportional and representative.

Benefits of technology

It effectively reduces the instability in the selection of representative skyline points, improves the accuracy and representativeness of representative skyline points, and thus optimizes the quality of data query results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119377491B_ABST
    Figure CN119377491B_ABST
Patent Text Reader

Abstract

The present invention discloses a data query method based on representative skyline point selection, comprising: obtaining a product data set to be queried, performing skyline query on the data set to obtain a number of skyline points, calculating the number of dominant points of each skyline point and the similarity between skyline point pairs, further calculating the preference values ​​of all skyline points to the remaining skyline points, and obtaining the preference sequence of each skyline point to the remaining skyline points; based on a single transferable voting method, taking all skyline points as voters and voted persons, performing a round of voting according to the preference sequence of each skyline point to the remaining skyline points, obtaining the first participant or the second participant in the voting round, removing the first participant or the second participant from the voted persons and performing corresponding ballot transfer, and repeating the process of selecting the first participant or the second participant and ballot transfer until a predetermined number of first participants are obtained, that is, obtaining a data query result for product recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of databases, and in particular relates to a data query method based on representative skyline point selection. Background Art

[0002] Recommending partial (data or entity) products that match the user's interests and goals based on the user's preferences and needs is a widely studied topic involving machine learning, data mining, recommendation systems and other fields. In practical applications, product recommendations can be applied to various industries and fields. For example, on e-commerce platforms, relevant products or promotions can be recommended based on the user's purchase history and browsing behavior. Skyline query is a common recommendation method that queries the optimal points in different dimensions and recommends them to users.

[0003] However, when dealing with large-scale data sets, especially in high-dimensional space, skyline queries may produce a large number of result sets. To solve this problem, the concept of representative skyline was first proposed in "Xuemin Lin, Yidong Yuan, Qing Zhang, and YingZhang. 2007. Selecting Stars: The k Most Representative Skyline Operator. InProceedings of the 23rd International Conference on Data Engineering, ICDE2007, Rada Chirkova, Asuman Dogac, M. Tamer Özsu, and Timos K. Sellis (Eds.).IEEE Computer Society, 86–95", which aims to identify k skyline points that dominate the largest number of points. Subsequently, "Yufei Tao, Ling Ding, Xuemin Lin, and Jian Pei. 2009. Distance-Based Representative Skyline. In Proceedings of the25th International Conference on Data Engineering, ICDE 2009, Yannis E.Ioannidis, Dik Lun Lee, and Raymond T. Ng (Eds.). IEEE Computer Society, 892–903" pointed out that the above method may lead to some almost unrepresentative skyline points due to its instability, and proposed a distance-based representative skyline calculated based on the k-center problem to better capture the outline of the skyline.

[0004] "Rui Mao, Taotao Cai, Rong-Hua Li, Jeffrey Xu Yu, and Jianxin Li. 2017. Efficient distance-based representative skyline computation in 2D space. World Wide Web 20, 4 (2017), 621–638) and (Sergio Cabello. 2023. Faster distance-based representative skyline and k-center along pareto front in the plane. J. Glob. Optim. 86, 2 (2023), 441–466" further reduced the time complexity of the distance-based representative skyline algorithm in two-dimensional space. "Atish Das Sarma, Ashwin Lall, DanuponNanongkai, Richard J. Lipton, and Jun (Jim) Xu. 2011. Representative skylinesusing threshold-based preference distributions. In Proceedings of the 27thInternational Conference on Data Engineering, ICDE 2011, Serge Abiteboul,Klemens Böhm, Christoph Koch, and Kian-Lee Tan (Eds.). IEEE Computer Society, 387–398》 defines user preference as a threshold function and selects k representative skyline points to maximize the likelihood of a random user click. Matteo Magnani, Ira Assent, and Michael L. Mortensen. 2014.Taking the Big Picture: representative skylines based on significance and diversity. VLDB J. 23, 5 (2014), 795–815》 considers importance and diversity when selecting representative skyline points.

[0005] "Malene Søholm, Sean Chester, and Ira Assent. 2016. Maximum Coverage Representative Skyline. In Proceedings of the 19th International Conference on Extending Database Technology, EDBT 2016, Evaggelia Pitoura, Sofian Maabout, Georgia Koutrika, Amélie Marian, Letizia Tanca, Ioana Manolescu, and Kostas Stefanidis (Eds.). OpenProceedings.org, 702–703" proposed the concept of maximum coverage representative skyline, which aims to identify k points that jointly dominate the largest area of ​​the data space in two dimensions. "MeiBai, Junchang Xin, Guoren Wang, Luming Zhang, Roger Zimmermann, Ye Yuan, andXindong Wu. 2016. Discovering the k Representative Skyline Over a SlidingWindow. IEEE Trans. Knowl. Data Eng. 28, 8 (2016), 2041–2056" studies the problem of computing k representative skylines over data streams.

[0006] Our focus is on finding representative skyline points in a general sense. Existing methods have advantages, but also disadvantages in some cases. The representative skyline in "Xuemin Lin, Yidong Yuan, Qing Zhang, and YingZhang. 2007. Selecting Stars: The k Most Representative Skyline Operator. InProceedings of the 23rd International Conference on Data Engineering, ICDE2007, Rada Chirkova, Asuman Dogac, M. Tamer Özsu, and Timos K. Sellis (Eds.).IEEE Computer Society, 86–95" only considers the number of non-skyline points, so it is greatly affected by the distribution of non-skyline points. When the non-skyline points deviate greatly from the distribution of skyline points, their representativeness is poor. In contrast, the representative skyline in Yufei Tao, Ling Ding, Xuemin Lin, and Jian Pei. 2009. Distance-BasedRepresentative Skyline. In Proceedings of the 25th International Conferenceon Data Engineering, ICDE 2009, Yannis E. Ioannidis, Dik Lun Lee, and Raymond T. Ng (Eds.). IEEE Computer Society, 892–903 only focuses on the outlines of skyline points and tends to select representatives uniformly along the skyline. It does not consider the relative distribution of skyline points. Summary of the invention

[0007] The purpose of the present invention is to address the deficiencies of the prior art and to provide a data query method based on representative skyline point selection, which takes into account both the proportion of skyline points and the number of non-skyline points dominated by them when selecting representative skyline points.

[0008] According to a first aspect of an embodiment of the present application, a data query method based on representative skyline point selection is provided, comprising:

[0009] (1) Obtain a product dataset to be queried, perform a skyline query on the product dataset to obtain a number of skyline points, and calculate the number of dominant points of each skyline point and the similarity between skyline point pairs;

[0010] (2) Based on the number of dominant points of each skyline point and the similarity between skyline point pairs, the preference values ​​of all skyline points to the remaining skyline points are calculated to obtain the preference sequence of each skyline point to the remaining skyline points;

[0011] (3) Based on the single transferable voting method, all skyline points are regarded as voters and voters, and a round of voting is performed according to the preference sequence of each skyline point for the remaining skyline points to obtain the first participant or the second participant in the voting round. The first participant or the second participant is removed from the voters and the corresponding vote is transferred. The process of selecting the first participant or the second participant and transferring the vote is repeated until a predetermined number of first participants are obtained, that is, a predetermined number of representative skyline points are selected, and all representative skyline points are used as the data query results for product recommendation.

[0012] Furthermore, in step (1), the number of dominant points of each skyline point is calculated as follows:

[0013] Get each skyline point The number of dominant non-skyline points , for the number of non-skyline points Normalize to get the skyline point The leading points :

[0014]

[0015] in is the number of skyline points in the product dataset.

[0016] Furthermore, in step (1), the similarity between skyline point pairs is calculated as follows:

[0017] Skyline point pair The distance between , calculate the similarity between skyline point pairs :

[0018] .

[0019] Furthermore, in step (2), the skyline point Point to the skyline Preference value Calculated by the following formula:

[0020]

[0021] in is a parameter that adjusts the weights of the two calculation factors. is the similarity between skyline point pairs, Skyline Point The leading points.

[0022] Furthermore, step (3) includes:

[0023] (3.1) Each voter votes for the candidate with the highest preference value;

[0024] (3.2) Determine whether the number of votes obtained by each voted person is greater than or equal to the predetermined quota;

[0025] (3.3) If there are several voters whose votes are greater than or equal to the quota, one of the voters is selected as the first participant, and based on the preference sequence of the voters who voted for the first participant, the part of the votes of the first participant that exceeds the quota is transferred to other voters, and the remaining votes of these voters are updated;

[0026] (3.4) If the number of votes obtained by all the voted persons is less than the quota, the voted person with the lowest number of votes is selected as the second participant, and based on the preference sequence of the voters who voted for the second participant, the votes of the second participant are transferred to other voted persons, and the remaining votes of these voters are updated;

[0027] (3.5) Repeat the above steps (3.2)-(3.4) until a predetermined number of first participants are selected, that is, a predetermined number of representative skyline points are selected, and all representative skyline points are used as data query results.

[0028] Furthermore, the quota ,in Gathering for voters, is the predetermined number of first participants to be selected.

[0029] According to a second aspect of an embodiment of the present application, there is provided a data query system based on representative skyline point selection, comprising a first database, a second database and a server;

[0030] The first database is used to store product data sets;

[0031] The server is used to perform data query based on the selection of representative skyline points according to the method of the first aspect to obtain data query results for product recommendations;

[0032] The second database is used to store the data query results of the product recommendation.

[0033] According to a third aspect of an embodiment of the present application, a computer program product is provided, comprising a computer program / instruction, which implements the method described in the first aspect when executed by a processor.

[0034] According to a fourth aspect of an embodiment of the present application, there is provided an electronic device, including:

[0035] one or more processors;

[0036] A memory for storing one or more programs;

[0037] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect.

[0038] According to a fifth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in the first aspect are implemented.

[0039] The technical solution provided by the embodiments of the present application may have the following beneficial effects:

[0040] It can be seen from the above embodiments that the present application adopts the STV (Single transferable vote) method to select representative skyline points. The skyline points act as both voters and voted for. A similarity function is defined based on the characteristics of the skyline scene to derive a preference sequence. Each voter selects those voted for who are most similar to him in terms of function value. The similarity function takes into account the distance between skyline points and the number of non-skyline points dominated by each skyline point. Skyline points that dominate more points are more likely to receive more votes and become the first participant, thereby obtaining a set of representative skyline points with good proportional representativeness, that is, a number of more representative products are selected from the product data set for recommendation to users.

[0041] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0043] Figure 1 The figure is a schematic diagram showing a data query system based on representative skyline point selection according to an exemplary embodiment.

[0044] Figure 2The figure is a schematic diagram of hotel addresses and prices according to an exemplary embodiment.

[0045] Figure 3 The present invention is a flowchart showing a data query method based on representative skyline point selection according to an exemplary embodiment.

[0046] Figure 4 The present invention is a block diagram showing a data query device based on representative skyline point selection according to an exemplary embodiment.

[0047] Figure 5 The diagram is a schematic diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0048] Here, exemplary embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application.

[0049] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0050] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0051] The method provided in this application can be based on Figure 1 The database system shown is implemented, and the database system may include: a first database, a second database and a server; the first database is used to store and provide product data sets, the server is used to process the product data sets provided by the first database, query to obtain representative skyline points as data query results, and the second database is used to store the data query results obtained by the server query.

[0052] In this example, the first database and the second database can be any type of relational databases. The data maintained in the first database exists in the form of a table, and each record in the data set is a row of data or a tuple in the table. For example, the storage form of the above-mentioned hotel table in the first database. The second database also maintains a table for representing a representative skyline point set of recommended products.

[0053] The server is a physical server including an independent host or a virtual server hosted by a host cluster. During operation, the server obtains a product data set from the first database, obtains a representative skyline point set through query, and stores it in the second database.

[0054] like Figure 3 As shown, the present application provides a data query method based on representative skyline point selection, which is applied to the server.

[0055] The following table 1 and Figure 2 The present method is described in detail in the following embodiments.

[0056] (1) Obtain the product dataset to be queried , perform skyline query on the product dataset to obtain several skyline points, and calculate the number of dominant points of each skyline point and the similarity between skyline point pairs ;

[0057] Specifically, firstly, according to the needs of the user, a skyline query is performed on the product dataset to obtain a number of skyline points.

[0058] In the database field, the definition of skyline is: given a A collection of dimensional points , each point It can be expressed as ,in yes of Features. A point is called dominating another point ,if For each And at least one dimension Make .exist The set of all points in the graph that are not dominated by any other point is called the skyline set .

[0059] Taking hotel recommendation as an example, Table 1 lists 10 hotels, namely , each hotel has two characteristics: distance from the destination and price. These hotels are skyline points, such as Figure 2This means that in terms of distance from the destination and price, there is no other choice better than these hotels.

[0060] Table 1 Hotel accommodation and price list

[0061]

[0062] In order to apply the STV process to skyline points, it is necessary to obtain the preference ranking of each skyline point to other skyline points, which will be used to establish the voting order. Regarding the voting preference ranking of skyline points and their distribution, it is first necessary to determine the similarity between skyline points. Define the similarity function , for two given points , ,Require .

[0063] The similarity is obtained by measuring the distance between skyline points, that is, the closer the distance, the greater the similarity (closer to 1), and the similarity function is defined as ,in It is a skyline point and The Euclidean distance between and The distance between calculate, Skyline Point Include Dimensional Features , Same reason.

[0064] For skyline points , Product Dataset The number of skyline points, is a collection of skyline points, express Zhongyou The set of dominant non-skyline points. Leading The number of non-skyline points in is recorded as .

[0065] The number of non-skyline points dominated by the skyline point It is also a factor that needs to be considered in the calculation of preference value, which shows that the number of data of skyline points is better than that of other points, such as the hotel in Table 1 The quality of the proposed method exceeds that of many other available options. It should be noted that this application does not require that the distribution of representative skyline points and non-skyline points be proportional. This application does not calculate the difference in the number of non-skyline points dominated by two skyline points, but directly considers the number of non-skyline points dominated by each skyline point.

[0066] The range of variation of may be large, so it should be normalized to be consistent with the range of the similarity function. Based on this requirement, The number of dominant points can be calculated by the following formula:

[0067]

[0068] (2) Based on the number of dominant points of each skyline point and the similarity between skyline point pairs, the preference values ​​of all skyline points to all skyline points are calculated to obtain the preference sequence of each skyline point to the remaining skyline points;

[0069] For all the skyline point pairs described above, according to the voting function The preference value is calculated by the definition of , and the preference sequence of each skyline point to all skyline points is obtained. Based on the preference sequence, the scores of all current skyline points are counted, and the winner is selected or the loser is eliminated according to the rule of single transferable vote, and the excess votes are transferred in the order of preference sequence. This step is continued until the winner is selected. Finally, the Shapley value is used to analyze the contribution of all eigenvalues ​​in the product dataset to the skyline point algorithm results.

[0070] Specifically, according to the similarity function value between the skyline point pairs obtained in step (1), And the number of non-skyline points dominated by each skyline point, calculate the preference value of all skyline point pairs, where right The preference value is calculated by the following formula:

[0071]

[0072] Relatively speaking, right The preference value is

[0073]

[0074] in is a parameter for adjusting the weights of the two calculation factors, which can be set according to actual conditions. For example, in the embodiment shown in Table 1, , which means that only the distances between skyline points are considered in the voting function.

[0075] Based on skyline points The preference values ​​of the remaining skyline points are sorted to obtain the preference sequence for the remaining skyline points.

[0076] (3) Based on the single transferable voting method, all skyline points are regarded as voters and voters, and a round of voting is performed according to the preference sequence of each skyline point for the remaining skyline points to obtain the first participant or the second participant in the voting round. The first participant or the second participant is removed from the voters and the corresponding vote is transferred. The process of selecting the first participant or the second participant and transferring the vote is repeated until the predetermined number of participants is obtained. The first participant gets the predetermined amount Each representative skyline point, all representative skyline points are taken as data query results;

[0077] use Represents a group of voters, Represents a set of voters. All skyline points are considered both voters and voters. Each voter Each person has one ballot and votes for the candidate with the highest preference value in the preference sequence.

[0078] Specifically, using the symbol To express the opinion of any voted person , does not exist and In other words, Contains several middle The candidate with the highest preference value.

[0079] Ordered partitions of called Votes , if it meets the following conditions:

[0080]

[0081]

[0082] In the STV system, the voted parties are strictly ranked, and each voter can only vote for one voted party in each round. In the skyline scenario, if each node only votes for itself at the beginning, then no node can exceed the quota. and win. To solve this problem, Include voters myself and The other voters with the highest preference value. Then, the voter At the same time These voters vote.

[0083] It should be noted that it can also be set that voters cannot vote for themselves. Only voters are included in Other voters with the highest preference values.

[0084] The standard STV protocol does not allow ties in rankings, but in practice there may be multiple voters with the same similarity as a single voter. Therefore, when there are tied voters, each voter can vote for all voters in the ranking at the same time.

[0085] The following first describes the parameters involved in the subsequent steps:

[0086] : The voted candidate with the highest preference score for each voter in the current round;

[0087] : Each voter The current number of remaining votes. Initially, the value of each voter is 1;

[0088] :Vote for each voted person the list of voters;

[0089] : Each voter The number of votes received.

[0090] The quota in this application That is, under this rule, no more than Voters will receive .

[0091] Combined with the single transferable voting algorithm, a round of voting is performed based on the preference sequence of each skyline point for the remaining skyline points, the first participant or the second participant in the round is calculated, and the excess votes are transferred among the voted according to the preference sequence. The process is repeated until the first participant is selected. Representative skyline points, specifically, Figure 3 As shown, each round can include the following processes:

[0092] (3.1) Each voter Vote for the voted person with the highest preference value;

[0093] According to the preference value, the votes of each skyline point are distributed to several points with the highest preference value.

[0094] (3.2) Judge each voted person Number of votes received Is it greater than or equal to the quota? ;

[0095] (3.3) If there are several votes obtained The first participant is selected as one of the voters, for example, based on the number of non-skyline points dominated by the candidate point or other custom functions to measure the quality of skyline points, and the votes of the first participant are divided into the first participant and the second participant is ... Exceeding quota Transfer part of the votes to other voters and update the remaining votes of these voters ;

[0096] If the voters Receive the highest number of votes and the number of votes exceeds the quota , then the voted will be selected as the first participant, and any excess votes will be transferred. Since it is possible to vote for multiple voters at the same time, Each vote ,from Delete and revoke their vote. At this time, if If it is empty, you need to add the next point with the highest preference value to it. The points in vote with the remaining votes and update .

[0097] It should be noted that when the first participant transfers his votes, the votes transferred are those that exceed the quota. If the first participant gets the number of votes , no vote transfer will be made.

[0098] (3.4) If the number of votes received by all voters is All less than the quota , then the voted person with the lowest number of votes is selected as the second participant, and the votes of the second participant are increased based on the preference sequence of the voters who voted for the second participant. Transfer to other voters and update the remaining votes of these voters ;

[0099] In the event that no vote exceeds the quota, the voter with the least votes shall be Eliminated, transferred by the second participant When the number of votes is Voters , will be Delete .if If empty, the voter The one with the highest preference value among the remaining voters joins middle, The points in will be voted and updated .

[0100] (3.5) Repeat the above steps (3.2)-(3.4) until the predetermined number is selected. First participant;

[0101] It should be noted that in the specific implementation, there is no way to select In the case of the first participant, if If a first participant is selected, but a new first participant cannot be selected in the next voting round, the first participant should be eliminated according to the rules until only one of the voted participants remains. When the number of votes is greater than 1, the remaining voters will be selected as the first participants.

[0102] The above process is further described below in conjunction with an embodiment. In the embodiment shown in Table 1, the quota Q is , that is, you need to select 3 representative hotels from all skyline points.

[0103] After initialization, the preference sequence of each voter is obtained by calculating the voting function between each point, as shown in Table 2.

[0104] Table 2 Voter preference sequence table

[0105]

[0106] The first round of ballots is shown in Figure 3.

[0107] Table 3 First round of ballots

[0108]

[0109] In the first round, there were three voters who met the quota, namely , , Assumptions Selected as the first participant.

[0110] because Just got to Q, no votes left to transfer. , , All voted, leaving 0 votes. This will result in the set The candidate This means that When there are multiple candidates, even if the voters in the current round cast one vote for these multiple candidates at the same time, for example At the same time, and , but in the end If you win, the votes will be consumed, and then Tied Voters no longer has that ballot.

[0111] The new values ​​of votes are shown in Table 4.

[0112] Table 4 Second round of ballots

[0113]

[0114] In the second round, now only arrive ,so Was selected as the first participant. There were no extra votes to transfer this time, so vote for ,Right now , and The votes have been used up. All voters in of will become 0.

[0115] At this stage, there is no voted candidate with a score of Q, so the voted candidate with the least votes will be eliminated, i.e. , , .

[0116] Since all three points All are 0, indicating that no votes can be transferred, so all three points are eliminated in the third to fifth rounds listed in Table 4.

[0117] Likewise, in rounds 6 and 7, 1 and However, their votes will not be transferred because the voters who voted for them Not empty.

[0118] Finally, in the eighth round, was eliminated due to having the fewest votes. Its votes were transferred to ,result is declared as the first participant. Table 5 shows the number of votes for each voted participant in each round. of changes.

[0119] Table 5 Each voted person in each round Table of changes in

[0120]

[0121] The Shapley value is a concept in cooperative game theory that fairly measures the contribution of each participant to the jointly generated utility. It is the only function that can satisfy the four ideal properties of explanatory model predictions, namely, balance, symmetry, additivity, and having zero elements. The Shapley value has been shown to play an important role in solving the interpretability problem. In the setting of this application, the contribution of the features of the skyline points to the calculation results is evaluated. Therefore, in this application, the Shapley value can also be used to analyze and calculate the contribution of each feature in the product dataset to each representative skyline point obtained in step (3).

[0122] After selecting a set of skyline points as representatives, we can further calculate the contribution of each feature of the representative skyline point to its representativeness. Although existing algorithms have explored various criteria for identifying representative skyline points, there is no explanation for the results. In this application, a competitive relationship emerges between skyline points and their features. The formation of representativeness depends on the combined effect of multiple features. One way to measure the contribution of a point's features to the selection result is to use the concept of Shapley value.

[0123] Given a set Skyline Point , so that each skyline point has a feature set .

[0124] In the feature subset The representative skyline point found on the skyline point of is expressed as , it only considers Features in. Skyline Points Features Contribution to this point becoming a representative point Calculated by the following formula:

[0125]

[0126] in is a feature subset, is the utility function, and the utility function is:

[0127]

[0128] in Defined as a skyline point The maximum number of votes obtained during the voting process, that is, if , Equal to the point The first participant selected and before the remaining votes are transferred .

[0129] For a specific point , calculate the contribution of all its features to whether the point can be selected as a representative skyline, that is, For all According to the efficiency axiom of Shapley value, This holds true for all skyline points. Therefore,

[0130]

[0131] Based on this, in a specific implementation, the contribution of each feature of a representative skyline point to the point becoming a representative point can be displayed in the query result, which can serve as a further explanation of the query result.

[0132] It should be noted that the product data set can be a data set of any product, such as hotels, food, tour groups, courses, etc., that is, this method can be used to recommend any product, and can also be used on various shopping platforms. After obtaining the product categories required by the user, more suitable products can be recommended to the user. Through the above method, firstly, skyline points (that is, relatively better products in the product data set) are obtained from the product data set according to user needs, and representative skyline points that consider the proportion of skyline points and the number of dominant points are further selected from the skyline points by the STV method, so as to obtain products that meet the user's needs and have certain representativeness, and these products are returned to the user as product recommendation query results.

[0133] Corresponding to the aforementioned embodiment of the data query method based on representative skyline point selection, the present application also provides an embodiment of a data query device based on representative skyline point selection.

[0134] Figure 4 is a block diagram of a data query device based on representative skyline point selection according to an exemplary embodiment. Figure 4 , the device may include:

[0135] The dominant point number and similarity calculation module 11 is used to obtain a product data set to be queried, perform a skyline query on the product data set to obtain a number of skyline points, and calculate the dominant point number of each skyline point and the similarity between skyline point pairs;

[0136] A preference value calculation module 12 is used to calculate the preference values ​​of all skyline points to the remaining skyline points according to the number of dominant points of each skyline point and the similarity between skyline point pairs, so as to obtain the preference sequence of each skyline point to the remaining skyline points;

[0137] The voting module 13 is used to use a single transferable voting method to take all skyline points as voters and voters, conduct a round of voting according to the preference sequence of each skyline point for the remaining skyline points, obtain the first participant or the second participant in the voting round, remove the first participant or the second participant from the voters and perform the corresponding ballot transfer, and repeat the process of selecting the first participant or the second participant and the ballot transfer until a predetermined number of first participants are obtained, that is, a predetermined number of representative skyline points are selected, and all the representative skyline points are used as data query results.

[0138] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0139] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0140] Accordingly, the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the above-mentioned data query method based on representative skyline point selection.

[0141] Accordingly, the present application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned data query method based on representative skyline point selection. Figure 5 As shown, it is a hardware structure diagram of a data query device based on representative skyline point selection provided by an embodiment of the present invention, in which any device with data processing capability is located, except Figure 5 In addition to the processor, memory and network interface shown, any device with data processing capability in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capability, which will not be described in detail.

[0142] Accordingly, the present application also provides a computer-readable storage medium on which computer instructions are stored, and when the instructions are executed by the processor, the data query method based on the selection of representative skyline points as described above is implemented. The computer-readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card (Flash Card), etc. equipped on the device. Furthermore, the computer-readable storage medium can also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or is to be output.

[0143] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.

Claims

1. A data query method based on representative skyline point selection, characterized in that: include: (1) Obtain a product dataset to be queried, perform a skyline query on the product dataset to obtain a number of skyline points, and calculate the number of dominant points of each skyline point and the similarity between skyline point pairs; (2) Based on the number of dominant points of each skyline point and the similarity between skyline point pairs, the preference values ​​of all skyline points to the remaining skyline points are calculated to obtain the preference sequence of each skyline point to the remaining skyline points; (3) Based on the single transferable voting method, all skyline points are regarded as voters and voters, and a round of voting is performed according to the preference sequence of each skyline point for the remaining skyline points to obtain the first participant or the second participant in the voting round. The first participant or the second participant is removed from the voters and the corresponding vote is transferred. The process of selecting the first participant or the second participant and transferring the vote is repeated until a predetermined number of first participants are obtained, that is, a predetermined number of representative skyline points are selected, and all representative skyline points are used as the data query results for product recommendation.

2. The method according to claim 1, characterized in that In step (1), the number of dominant points of each skyline point is calculated as follows: Get each skyline point The number of dominant non-skyline points , for the number of non-skyline points Normalize to get the skyline point The leading points : , in is the number of skyline points in the product dataset.

3. The method according to claim 2, characterized in that In step (1), the similarity between skyline point pairs is calculated as follows: Skyline point pair The distance between , calculate the similarity between skyline point pairs : 。 4. The method according to claim 3, characterized in that In step (2), the skyline point Point to the skyline Preference value Calculated by the following formula: , in is a parameter that adjusts the weights of the two calculation factors. is the similarity between skyline point pairs, Skyline Point The leading points.

5. The method according to claim 1, characterized in that Step (3) includes: (3.1) Each voter votes for the candidate with the highest preference value; (3.2) Determine whether the number of votes obtained by each voted person is greater than or equal to the predetermined quota; (3.3) If there are several voters whose votes are greater than or equal to the quota, one of the voters is selected as the first participant, and based on the preference sequence of the voters who voted for the first participant, the part of the votes of the first participant that exceeds the quota is transferred to other voters, and the remaining votes of these voters are updated; (3.4) If the number of votes obtained by all the voted persons is less than the quota, the voted person with the lowest number of votes is selected as the second participant, and based on the preference sequence of the voters who voted for the second participant, the votes of the second participant are transferred to other voted persons, and the remaining votes of these voters are updated; (3.5) Repeat the above steps (3.2)-(3.4) until a predetermined number of first participants are selected, that is, a predetermined number of representative skyline points are selected, and all representative skyline points are used as data query results.

6. The method according to claim 5, characterized in that Said quota ,in Gathering for voters, is the predetermined number of first participants to be selected.

7. A data query system based on representative skyline point selection, characterized in that: comprising a first database, a second database and a server; The first database is used to store product data sets; The server is used to perform data query based on representative skyline point selection according to the method described in any one of claims 1 to 6 to obtain data query results for product recommendation; The second database is used to store the data query results of the product recommendation.

8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Traffic alarm condition level predication method based on distance metric learning

    CN104834977A

  • A digital city skyline extraction method

    CN109285177A