Data screening method, device, computer equipment and storage medium
By constructing the correlation structure chart between commodity nodes and merging feature, and combining the GCN model for convolutional classification, the accuracy and calculation time problems of commodity sales data classification screening in the retail industry are solved, and efficient and accurate data screening is achieved.
Patent Information
- Application Number
- CN202110425406.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-04-20
AI Technical Summary
In the classification and screening of sales data of goods in the retail industry, the prior art is difficult to meet the needs of large-scale data and complex relationships, resulting in low classification accuracy and long calculation time.
By obtaining multiple shopping lists, using the Apriori algorithm to find the double-item frequent set, calculate the support degree and confidence, build the association structure chart between product nodes, obtain the node feature vector and merge it, and finally input the association structure chart and node feature vector into the GCN model for convolutional classification.
It reduces the complex relationships and data volume between products, reduces the calculation time of data, achieves more accurate classification results, optimizes the classification process, and reduces the scale and computational complexity of classified data.
Smart Images

Figure CN113032648B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a data screening method, device, computer equipment and storage medium. Background Art
[0002] At present, in the process of data analysis, it is a common technical means to classify and filter data according to certain rules. For example, in the retail industry, sometimes in order to study the sales and replenishment of goods or to portray the image of consumer users based on the goods, goods are classified and filtered. However, due to the large variety of goods and complex influencing relationships, it is difficult to have an efficient classification and screening method, which often relies on manual work, is time-consuming and labor-intensive, and inefficient.
[0003] In the prior art, commonly used methods for screening (classifying) effective data include XGBoost, SVM, random forest, CNN and other methods. These classification methods have good screening effects on small-scale data and usually there is no correlation between the classified objects. However, in the retail industry, the sales data of commodities has the characteristics of large data scale, multiple feature dimensions, and complex mutual influence relationships. If conventional machine learning classification algorithms are used, they cannot meet actual needs. On the one hand, when the amount of data is large, the parameter optimization process will be more cumbersome and the calculation time will be long; on the other hand, because the input of the model does not take into account the mutual influence between commodities, the classification accuracy is not high. Summary of the invention
[0004] The purpose of the present invention is to provide a data screening method, device, computer equipment and storage medium, aiming to solve the problem that the accuracy of the classification and screening of commodities in the prior art needs to be improved.
[0005] In order to solve the above technical problems, the object of the present invention is to achieve the following technical solutions: to provide a data screening method, comprising:
[0006] Acquire multiple shopping lists, and obtain a transaction data set according to the shopping lists, wherein each shopping list is a transaction, and a commodity in the shopping list is a single item set in the transaction;
[0007] Using the Apriori algorithm to find the two-item frequent set from the transaction data set;
[0008] Calculating the support and confidence of each two-item set in the two-item frequent set, and constructing a correlation structure graph between commodity nodes according to the support and confidence of the two-item set;
[0009] Obtaining important feature vectors, secondary feature vectors, and external feature vectors of each commodity node in the association structure graph, and merging the important feature vectors, secondary feature vectors, and external feature vectors to obtain a node feature vector;
[0010] The association structure graph and node feature vector are input into the GCN model for convolution classification, and the classification result is output.
[0011] In addition, the technical problem to be solved by the present invention is to provide a data screening device, comprising:
[0012] A data acquisition unit, configured to acquire a plurality of shopping lists and obtain a transaction data set according to the shopping lists, wherein each shopping list is a transaction and a commodity in the shopping list is a single item set in the transaction;
[0013] An algorithm unit, used for finding a two-item frequent set from the transaction data set using an Apriori algorithm;
[0014] A construction unit, used to calculate the support and confidence of each two-item set in the two-item frequent set, and to construct an association structure graph between commodity nodes according to the support and confidence of the two-item set;
[0015] A vector acquisition unit, used to acquire important feature vectors, secondary feature vectors and external feature vectors of each commodity node in the association structure diagram, and perform feature merging on the important feature vectors, secondary feature vectors and external feature vectors to obtain a node feature vector;
[0016] The classification unit is used to input the association structure graph and the node feature vector into the GCN model for convolution classification and output the classification result.
[0017] In addition, an embodiment of the present invention provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the data screening method described in the first aspect when executing the computer program.
[0018] In addition, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the data screening method described in the first aspect above.
[0019] The embodiment of the present invention discloses a data screening method, device, computer equipment and storage medium. The method includes obtaining multiple shopping lists, obtaining a transaction data set according to the shopping lists; finding a two-item frequent set from the transaction data set using the Apriori algorithm; calculating the support and confidence of each two-item set in the two-item frequent set, and constructing a correlation structure graph between commodity nodes according to the support and confidence of the two-item set; obtaining a node feature vector in the correlation structure graph; inputting the correlation structure graph and the node feature vector into a GCN model for convolution classification, and outputting the classification result. The embodiment of the present invention analyzes the sales data of various commodities, only filters out commodities with correlation relationships, so as to reduce the complex relationship and data volume between commodities, and then constructs a correlation structure graph between commodities, and then performs high-level feature extraction on commodities with correlation relationships, which can reduce the feature dimension of data, thereby reducing the calculation time of data, and finally realizes data screening and classification through the GCN model, realizes the advantages of accurate classification, and optimizes the classification process, reduces the scale of classification data, and reduces the complexity of classification calculation. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0021] Figure 1 A schematic diagram of a flow chart of a data screening method provided by an embodiment of the present invention;
[0022] Figure 2 A schematic diagram of a sub-flow chart of step S104 of the data screening method provided by an embodiment of the present invention;
[0023] Figure 3 An exemplary association structure diagram provided for an embodiment of the present invention;
[0024] Figure 4 A structural diagram of a graph convolutional neural network provided by an embodiment of the present invention;
[0025] Figure 5 A schematic block diagram of a data screening device provided by an embodiment of the present invention;
[0026] Figure 6 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0029] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0030] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0031] See also Figure 1 , Figure 1 A schematic diagram of a flow chart of a data screening method provided by an embodiment of the present invention;
[0032] like Figure 1 As shown, the method includes steps S101 to S105.
[0033] S101. Acquire multiple shopping lists, and obtain a transaction data set according to the shopping lists, wherein each shopping list is a transaction, and a commodity in the shopping list is a single item set in the transaction.
[0034] In this embodiment, multiple shopping list information of consumer users is obtained, a preset number of shopping lists are extracted, and a transaction data set is obtained according to the shopping lists. For example, the transaction data set is shown in Table 1 below:
[0035] Table 1
[0036] Shopping list number Product ID 1 a, c, e 2 b, d 3 b, c 4 a, b, c, d … …
[0037] Table 1 contains 4 shopping lists and 5 different commodities, which are denoted as a, b, c, d, and e. The data in a shopping list number is a transaction, that is, Table 1 contains 4 transactions. A commodity in the shopping list is a single item set in the transaction, that is, a, b, c, d, and e are single item sets in the transaction. The commodities are listed by obtaining the transaction data set for subsequent classification and screening.
[0038] S102: Using the Apriori algorithm, find out the two-item frequent set from the transaction data set.
[0039] Specifically, step S102 includes:
[0040] The support of each of the single-item sets is calculated, and the single-item sets whose support is greater than a preset minimum support threshold are constructed as single-item frequent sets; all the single-item sets in the single-item frequent sets are combined in pairs to obtain multiple double-item sets, the support of each of the double-item sets is calculated, and the double-item sets whose support is greater than the minimum support threshold are constructed as double-item frequent sets.
[0041] In this embodiment, the support of the single item set refers to the probability that the single item set exists in a transaction, that is, the probability that a certain commodity appears in a consumption list. The support of each single item set is calculated. If the support of the single item set is greater than the preset minimum support threshold, it means that the commodity corresponding to the single item set is a frequently sold commodity. Such frequently sold commodities can be grouped, that is, the single item sets whose support is greater than the preset minimum support threshold are constructed as single item frequent sets; the commodities corresponding to the single item frequent sets obtained in this way are more valuable for analysis. Specifically, all single-item sets in the single-item frequent set are combined in pairs to obtain multiple double-item sets, and then the support of each double-item set is calculated. The support of the double-item set refers to the probability that the double-item set exists in a transaction at the same time, that is, the probability that the corresponding two commodities appear in a consumption list. If the support of the double-item set is greater than the preset minimum support threshold, it means that the two commodities corresponding to the double-item set are commodities that are frequently purchased at the same time. Such commodities that are frequently purchased at the same time can be grouped, that is, the double-item sets whose support is greater than the minimum support threshold are constructed as double-item frequent sets. The two commodities corresponding to the double-item frequent set obtained in this way can better reflect the association relationship between the two commodities.
[0042] More specifically, the support of a single item set is calculated as follows:
[0043] Calculate the number of transactions that contain the specified single item set in the transaction data set, and calculate the ratio of the number of transactions that contain the specified single item set to the number of all transactions, and use the obtained ratio as the support of the specified single item set. Taking Table 1 above as an example, assuming that the specified single item set corresponds to commodity a, the number of transactions of the specified single item set is 2, and the number of all transactions is 4, then the support of the specified single item set is 0.5.
[0044] The support of a two-item set is calculated as follows:
[0045] The number of transactions in the transaction data set that include the specified double-item set is calculated, and the ratio of the number of transactions that include the specified double-item set to the number of all transactions is calculated, and the obtained ratio value is used as the support of the specified double-item set; for example, if the specified double-item set includes single-item set A and single-item set B, it can be calculated according to the following formula:
[0046]
[0047] Taking Table 1 above as an example, assuming that single item set A is commodity a, single item set B is commodity b, the number of transactions containing the specified double item set is 1, and the number of all transactions is 4, then the support degree containing the specified double item set is 0.25.
[0048] S103: Calculate the support and confidence of each two-item set in the two-item frequent set, and construct a correlation structure graph between commodity nodes according to the support and confidence of the two-item set.
[0049] In this embodiment, the confidence of the two-item set refers to the probability that when one single item set in the two-item set already exists in the transaction, there is another single item set, that is, when the corresponding one commodity already exists in the shopping list, there is the probability that another corresponding commodity also exists.
[0050] Specifically, the confidence of the two-item set is calculated as follows:
[0051] Calculate the number of transactions in the transaction data set that include the specified double-item set and the number of transactions that include the target single-item set in the specified double-item set; then calculate the ratio of the number of transactions that include the specified double-item set to the number of transactions that include the target single-item set in the specified double-item set, and use the obtained ratio as the confidence of the specified double-item set. It should be noted that the confidence of the double-item set here has a two-way calculation situation. For example, the double-item set includes single-item set A and single-item set B, and the confidence can be calculated according to the following formula:
[0052]
[0053] Taking Table 1 above as an example, assuming that item set A is product a and item set B is product b, the confidence in the direction from product a to product b is: the ratio of the number of transactions containing product a and product b (i.e. 1) to the number of transactions containing product a (i.e. 2); that is, the confidence in the direction from product a to product b is 0.5.
[0054] The confidence level in the direction from product b to product a is: the ratio of the number of transactions that include product b and product a (i.e., 1) to the number of transactions that include product b (i.e., 3); that is, the confidence level in the direction from product b to product a is 0.33.
[0055] Furthermore, after calculating and obtaining the support and confidence of each two-item set in the two-item frequent set, the degree of association between two commodities can be determined according to the support and confidence of each two-item set, and two commodities whose association reaches a preset degree are connected, thereby constructing a correlation structure diagram between commodity nodes.
[0056] In one embodiment, constructing a correlation structure graph between commodity nodes according to the support and confidence of the two-item set includes:
[0057] Determine whether the conditions that the support of the double-item set is greater than the preset minimum support and the confidence is greater than the preset minimum confidence are met at the same time. If so, associate the two commodity nodes corresponding to the double-item set, take the larger value of the support and the confidence of the double-item set as the weight of the connecting edge of the two commodity nodes, and construct an association structure diagram between the commodity nodes.
[0058] In this embodiment, one of the corresponding commodity nodes in the two-item set is recorded as i, the other commodity node is recorded as j, and the support degree is recorded as S ij , the confidence level is recorded as C ij , the weight of the connecting edge between two commodity nodes is recorded as A ij , then A ij is defined as follows:
[0059]
[0060] That is, when the two-item set satisfies the conditions of support not less than 0.5 and confidence not less than 0.6 at the same time, it is considered that the corresponding two commodity nodes have a large mutual influence relationship, and the larger value of the support and confidence is taken as the weight of the connection edge connecting the two commodity nodes; if it does not meet the conditions, it is considered that the corresponding two commodity nodes have a small mutual influence relationship and can be ignored. In this way, each commodity node can be connected into an association structure diagram according to the definition of the weight between two points, which can be referred to as Figure 3 Example association structure diagram.
[0061] S104: Obtain important feature vectors, secondary feature vectors, and external feature vectors of each commodity node in the association structure diagram, and merge the important feature vectors, secondary feature vectors, and external feature vectors to obtain a node feature vector.
[0062] In this embodiment, the important feature vectors are the existing replenishment parameters in the automatic replenishment system, including one or more of the store number, fast-selling degree, predicted daily average sales, and actual daily average sales; the secondary feature vectors are the attribute characteristics of the goods, including one or more of the category, shelf life, and store arrival time; the external feature vectors include one or more of holidays and promotional activities. The node feature vector can be obtained by merging the important feature vectors, secondary feature vectors, and external feature vectors. This embodiment takes all these features into account as node feature vectors, fully considering various factors, and providing a more scientific and effective basis for subsequent data classification.
[0063] In one embodiment, if Figure 2 As shown, the step S104 includes:
[0064] S201, obtaining an important feature vector, a secondary feature vector, and an external feature vector of each commodity node in the association structure graph;
[0065] S202, performing normalization and one-hot encoding on the important feature vectors;
[0066] S203, extracting features from the processed important feature vectors according to the following formula:
[0067]
[0068] Among them, z is the important feature vector after feature extraction, p*q is the size of the convolution kernel, f is the nonlinear activation function, and w i is the weight, v i is the important feature vector of the input, b is the bias;
[0069] S204: Merge the important feature vector after feature extraction with the secondary feature vector and the external feature vector to obtain a node feature vector.
[0070] In this embodiment, the important feature vector has the greatest impact on the model; this part of the feature vector is generated based on the automatic replenishment system, so the vector dimension is large and the information contained is relatively complex. In order to capture more useful information, the CNN model is used to extract feature information from this part of the feature vector before classification to extract higher-level features, specifically:
[0071] First, the important feature vectors are normalized according to the following formula to standardize the order of magnitude:
[0072]
[0073] Where x is the input value before normalization, x′ is the value after normalization, and x min is the minimum value of the important eigenvector, x max is the maximum value of the important eigenvector; x, x′, x max Substituting the specific value of into the above formula for calculation can realize normalization and obtain the important eigenvector after normalization.
[0074] Secondly, it should be noted that one-hot encoding (also known as one-bit effective encoding, which is the representation of categorical variables as binary vectors) is required for important feature vectors of some text classes. For example, for the feature vector of the text class of fast-selling degree, the feature vector is described according to whether the product belongs to fast-selling, ordinary selling, or slow-selling. After one-hot encoding, the feature can be changed from the original one-column feature to three columns, and only the position corresponding to a specific class is 1. That is, if the product is a fast-selling product, only the position corresponding to the fast-selling column is 1, and the other positions are 0. The final result is a three-dimensional sparse matrix. The following three-dimensional matrix after one-hot encoding can be referred to:
[0075] Secondly, a two-layer convolutional layer is used to extract deep features from the important feature vectors after normalization and one-hot encoding. Specifically, assuming that the number of products is n and the feature vector corresponding to each product is m-dimensional, the input of the CNN model is an n×m matrix, such as Figure 4 As shown in the figure, the entire feature extraction process goes through two convolutional layers, followed by a pooling layer, and finally a dropout layer with an empirical value of probability 0.3 is added. The dropout layer randomly deletes some neurons during training to prevent overfitting. That is, the specific feature extraction process: p*q, f, w are respectively i 、v i Substituting the values of , b into the above formula for calculation, the important feature vector after deep feature extraction can be output.
[0076] Finally, the important feature vector after feature extraction is merged with the secondary feature vector and the external feature vector to obtain the final node feature vector.
[0077] S105: Input the association structure graph and node feature vector into the GCN model for convolution classification, and output the classification result.
[0078] Specifically, the step S105 includes:
[0079] Calculate and output the classification result Z according to the following formula:
[0080] L 0 =X;
[0081]
[0082]
[0083]
[0084] Among them, L 0 is the initial input; X represents all commodity nodes; L 1 and L 2 are the outputs of layer 1 and layer 2 in the GCN model respectively; A is the adjacency matrix composed of the weights of the connecting edges; is the adjacency matrix obtained by normalizing the adjacency matrix, D is the degree matrix, where D ii =∑ j A ij ; W1 is the weight matrix of the first layer, W2 is the weight matrix of the second layer, and ρ is the activation function Relu.
[0085] L 0 ,X,L 1 , L 2 , Substituting the values of W1, W2, and ρ into the above formula for calculation, the classification result Z can be output, thereby completing the classification of the products.
[0086] It should be noted that the input of the CNN model used for image classification before all belongs to the data of Euclidean space, and the characteristic of the data is that the structure is very regular. However, the research object of the present invention is two commodities of different categories, and the relationship structure between the two commodities is irregular, which is a more complex graph structure. The structure of the graph is very irregular and can be considered as a kind of infinite-dimensional data, so it has no translation invariance. The surrounding structure of each commodity node is unique. For data of this structure, it is obviously inappropriate to use traditional CNN and RNN to perform data screening. Therefore, the present invention adopts a GCN model that specifically processes the input of this graph structure data to realize the data classification work of the automatic replenishment system.
[0087] In one embodiment, after step S105, the following steps are included:
[0088] Input the classification results into the following loss function formula to optimize the parameters of the GCN model:
[0089]
[0090] Among them, L is the error value, y is the true label, is the classification result Z predicted by the GCN model.
[0091] In this embodiment, in order to further optimize the GCN model to improve the accuracy of the classification results, y and The value of is substituted into the above formula for calculation, thereby optimizing the parameters of the GCN model.
[0092] The embodiment of the present invention further provides a data screening device, which is used to execute any embodiment of the above-mentioned data screening method. Figure 5 , Figure 5 It is a schematic block diagram of a data screening device provided by an embodiment of the present invention.
[0093] like Figure 5 As shown, the data screening device 500 includes: a data acquisition unit 501, an algorithm unit 502, a construction unit 504, a vector acquisition unit 505 and a classification unit 505.
[0094] The data acquisition unit 501 is used to acquire multiple shopping lists and obtain a transaction data set according to the shopping lists, wherein each shopping list is a transaction and a commodity in the shopping list is a single item set in the transaction;
[0095] An algorithm unit 502, configured to find a two-item frequent set from the transaction data set using an Apriori algorithm;
[0096] A construction unit 504 is used to calculate the support and confidence of each two-item set in the two-item frequent set, and to construct an association structure graph between commodity nodes according to the support and confidence of the two-item set;
[0097] The vector acquisition unit 504 is used to acquire the important feature vector, the secondary feature vector and the external feature vector of each commodity node in the association structure diagram, and perform feature merging on the important feature vector, the secondary feature vector and the external feature vector to obtain a node feature vector;
[0098] The classification unit 505 is used to input the association structure graph and the node feature vector into the GCN model for convolution classification and output the classification result.
[0099] The device constructs an irregular association structure diagram containing the mutual influence relationship between commodities through the strong association rules generated by the Apriori algorithm, fully considers various factors that may affect the classification results, and performs high-level feature extraction to obtain node feature vectors. Finally, the GCN model is used to achieve more effective and accurate data screening and classification, which has the advantage of improving classification accuracy.
[0100] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0101] The above data screening device can be implemented in the form of a computer program. The computer program can be Figure 6 Runs on the computer device shown.
[0102] See also Figure 6 , Figure 6 600 is a schematic block diagram of a computer device provided by an embodiment of the present invention. The computer device 600 is a server, which can be an independent server or a server cluster composed of multiple servers.
[0103] See also Figure 6 The computer device 600 includes a processor 602 , a memory and a network interface 605 connected via a system bus 601 , wherein the memory may include a non-volatile storage medium 603 and an internal memory 604 .
[0104] The non-volatile storage medium 603 may store an operating system 6031 and a computer program 6032. When the computer program 6032 is executed, the processor 602 may execute a data screening method.
[0105] The processor 602 is used to provide computing and control capabilities to support the operation of the entire computer device 600 .
[0106] The internal memory 604 provides an environment for the operation of the computer program 6032 in the non-volatile storage medium 603. When the computer program 6032 is executed by the processor 602, the processor 602 can execute the data screening method.
[0107] The network interface 605 is used for network communication, such as providing data information transmission, etc. Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device 600 to which the solution of the present invention is applied. The specific computer device 600 may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0108] Those skilled in the art will understand that Figure 6 The embodiments of the computer device shown in the figure do not constitute a limitation on the specific composition of the computer device. In other embodiments, the computer device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. For example, in some embodiments, the computer device may only include a memory and a processor. In such embodiments, the structure and function of the memory and the processor are the same as those of the embodiment of the present invention. Figure 6 The embodiments shown are consistent and will not be described again here.
[0109] It should be understood that in the embodiment of the present invention, the processor 602 may be a central processing unit (CPU), and the processor 602 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0110] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the data screening method of the embodiment of the present invention is implemented.
[0111] The storage medium is a physical, non-transient storage medium, for example, it can be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk, etc., which can store program codes.
[0112] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0113] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A data screening method, characterized in that: include: Acquire multiple shopping lists, and obtain a transaction data set according to the shopping lists, wherein each shopping list is a transaction, and a commodity in the shopping list is a single item set in the transaction; Using the Apriori algorithm to find the two-item frequent set from the transaction data set; Calculating the support and confidence of each two-item set in the two-item frequent set, and constructing an association structure graph between commodity nodes according to the support and confidence of the two-item set, specifically including: judging whether the conditions that the support of the two-item set is greater than a preset minimum support and the confidence is greater than a preset minimum confidence are simultaneously met, and if so, associating the two commodity nodes corresponding to the two-item set, and taking the larger value of the support and the confidence of the two-item set as the weight of the connecting edge of the two commodity nodes, and constructing an association structure graph between commodity nodes; Obtaining the important feature vectors, secondary feature vectors and external feature vectors of each commodity node in the association structure diagram, and merging the important feature vectors, secondary feature vectors and external feature vectors to obtain the node feature vector; specifically including: obtaining the important feature vectors, secondary feature vectors and external feature vectors of each commodity node in the association structure diagram; normalizing and one-hot encoding the important feature vectors; according to the formula Extract features from the processed important feature vector; wherein z is the important feature vector after feature extraction, is the size of the convolution kernel, f is the nonlinear activation function, wi is the weight, vi is the input important feature vector, and b is the bias; merge the important feature vector after feature extraction with the secondary feature vector and the external feature vector to obtain a node feature vector; The association structure graph and node feature vector are input into the GCN model for convolution classification, and the classification result is output.
2. The data screening method according to claim 1, characterized in that: The method of finding a two-item frequent set from the transaction data set using the Apriori algorithm includes: The support of each of the single-item sets is calculated, and the single-item sets whose support is greater than a preset minimum support threshold are constructed as single-item frequent sets; all the single-item sets in the single-item frequent sets are combined in pairs to obtain multiple double-item sets, the support of each of the double-item sets is calculated, and the double-item sets whose support is greater than the minimum support threshold are constructed as double-item frequent sets.
3. The data screening method according to claim 2, characterized in that: The support of a single item set is calculated as follows: Calculate the number of transactions that contain the specified single itemset in the transaction data set, calculate the ratio of the number of transactions of the specified single itemset to the number of all transactions, and use the obtained ratio as the support of the specified single itemset; The support of a two-item set is calculated as follows: The number of transactions containing the specified double-item set in the transaction data set is calculated, and the ratio of the number of transactions of the specified double-item set to the number of all transactions is calculated, and the obtained ratio value is used as the support of the specified double-item set.
4. The data screening method according to claim 1, characterized in that: The step of inputting the association structure graph and the node feature vector into the GCN model for convolution classification and outputting the classification result includes: Calculate and output the classification result Z according to the following formula: L 0 =X; Among them, L 0 is the initial input, X represents all commodity nodes, L 1 and L 2 are the outputs of the first and second layers in the GCN model, respectively. A is the adjacency matrix composed of the weights of the connecting edges. is the adjacency matrix obtained by normalizing the adjacency matrix, W1 is the first layer weight matrix, W2 is the second layer weight matrix, and ρ is the activation function Relu.
5. The data screening method according to claim 1, characterized in that: After the step of inputting the association structure graph and the node feature vector into the GCN model for convolution classification and outputting the classification result, the method further comprises: Input the classification results into the following loss function formula to optimize the parameters of the GCN model: Among them, L is the error value, y is the true label, is the classification result Z predicted by the GCN model.
6. A data screening device, characterized in that: include: A data acquisition unit, configured to acquire a plurality of shopping lists and obtain a transaction data set according to the shopping lists, wherein each shopping list is a transaction and a commodity in the shopping list is a single item set in the transaction; An algorithm unit, used for finding a two-item frequent set from the transaction data set using an Apriori algorithm; A construction unit is used to calculate the support and confidence of each two-item set in the two-item frequent set, and construct an association structure diagram between commodity nodes according to the support and confidence of the two-item set, specifically comprising: judging whether the conditions that the support of the two-item set is greater than a preset minimum support and the confidence is greater than a preset minimum confidence are simultaneously met, and if so, associating the two commodity nodes corresponding to the two-item set, and taking the larger value of the support and the confidence of the two-item set as the weight of the connection edge of the two commodity nodes, and constructing an association structure diagram between commodity nodes; The vector acquisition unit is used to acquire the important feature vectors, secondary feature vectors and external feature vectors of each commodity node in the association structure diagram, and merge the important feature vectors, secondary feature vectors and external feature vectors to obtain the node feature vector; specifically, it includes: acquiring the important feature vectors, secondary feature vectors and external feature vectors of each commodity node in the association structure diagram; normalizing and one-hot encoding the important feature vectors; according to the formula Extract features from the processed important feature vector; wherein z is the important feature vector after feature extraction, is the size of the convolution kernel, f is the nonlinear activation function, wi is the weight, vi is the input important feature vector, and b is the bias; merge the important feature vector after feature extraction with the secondary feature vector and the external feature vector to obtain a node feature vector; The classification unit is used to input the association structure graph and the node feature vector into the GCN model for convolution classification and output the classification result.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the data screening method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the data screening method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Probability graph-based transformer state association rule mining method
CN106649479A
Method and device for determining degree of association of commodities
CN107944896A
Commodity recommendation method and system based on user session and graph convolutional neural network
CN110490717A