Analysis Method and Device for Regional Education Demand
By constructing a binary graph of e-commerce platform data and a graph embedding model for self-supervised task training, combined with the attractive model, the high cost and long-term problems of educational demand prediction are solved, low-cost and short-term educational demand analysis are realized, and data support is provided for educational institutions' site selection and planning.
Patent Information
- Application Number
- CN202210378977.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-04-12
AI Technical Summary
The existing technology consumes a lot of manpower, material resources and financial resources in the forecast of educational demand, with a long cycle and difficulty in considering the impact of migrant population, which leads to difficulties in analyzing educational needs and is unable to effectively guide educational institutions' site selection and planning.
By building a binary graph based on e-commerce platform data, using the graph automatic coding network for self-supervising task training, combining the attractive model to predict educational needs, and analyzing regional educational needs based on user purchasing behavior.
It realizes low-cost and short-cycle educational needs analysis, provides data basis for site selection and planning of educational institutions, reduces human, material and financial resources investment, and can effectively capture changes in regional educational needs.
Smart Images

Figure CN114782228B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for analyzing regional education needs. Background Art
[0002] The siting of educational infrastructure is an important issue in urban planning. The prerequisite for solving this problem is to predict and analyze the educational needs in the city. Considering the influence of the floating population, it is very difficult to predict educational needs. Currently, the prediction and analysis of educational needs are mainly achieved through methods such as manual visits and questionnaire statistics.
[0003] In the process of implementing the present invention, the inventors found the following defects in the prior art:
[0004] 1) Each research cycle is long, consuming a large amount of manpower, material resources and financial resources, with high costs and long cycles;
[0005] 2) Due to the influence of the floating population, there are complex spatio-temporal correlations in the changes of educational needs. It is difficult to estimate the real educational needs of the local area only based on past data. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a method and device for analyzing regional education needs, which can save manpower, material resources and financial resources, have low costs and short cycles, solve the problem of difficult demand analysis, and provide a data basis for the siting and planning arrangement of educational institutions.
[0007] To achieve the above object, according to one aspect of the embodiments of the present invention, a method for analyzing regional education needs is provided.
[0008] A method for analyzing regional education needs includes:
[0009] Constructing a bipartite graph of regions and commodities based on e-commerce platform data;
[0010] Inputting the bipartite graph into a graph autoencoder network, and training a graph embedding model through a self-supervised task. The graph embedding model is used to represent the association relationship between regions and commodities;
[0011] Constructing an attraction model based on historical enrollment data. The attraction model is used to predict the future enrollment numbers of each region;
[0012] Calculating the educational needs of each region based on the output results of the graph embedding model and the attraction model.
[0013] Optionally, constructing a bipartite graph of regions and commodities based on e-commerce platform data includes:
[0014] Clean the e-commerce platform data for address correction and deletion of abnormal user data. The e-commerce platform data includes product data, regional data, and order data;
[0015] Extract features from the product data to construct a product feature vector. The product data includes product self-information and age information;
[0016] Extract features from the regional data to construct a regional feature vector. The regional data includes the regional information of the delivery address of the order;
[0017] Take regions and products as nodes of a bipartite graph respectively, take the regional feature vector and the product feature vector as node feature vectors, form edges by associating regional nodes and product nodes through order data, and take the number of orders for products purchased within a region as the weight of the edge to construct a bipartite graph.
[0018] Optionally, after extracting features from the product data to construct a product feature vector, it further includes:
[0019] For the product data, aggregate the products according to the product categories they belong to, and construct a bipartite graph based on the aggregated products.
[0020] Optionally, the steps for address correction include:
[0021] Use the geographical parsing services of multiple geographical information service providers to parse the longitude and latitude of educational institutions;
[0022] Obtain the voting scores of each parsing result by voting according to the distribution of the parsing results;
[0023] Determine the final result of the address of the educational institution according to the voting scores and preset rules for address correction.
[0024] Optionally, taking the number of orders for products purchased within a region as the weight of the edge includes:
[0025] Bin the number of orders for products purchased within a region to obtain multiple bin values, and establish a mapping relationship between the order quantity and the bin values;
[0026] Take the bin value corresponding to the order quantity as the weight of the edge.
[0027] Optionally, inputting the bipartite graph into a graph autoencoder network and training through a self-supervised task to obtain a graph embedding model includes:
[0028] Input the bipartite graph into the graph encoder of the graph autoencoder network, extract structural embedding information through graph convolution, and fuse the structural predecessor information with the features of the nodes through a fully connected layer to obtain the hidden layer representation of the nodes;
[0029] Input the hidden layer representation of the node into the graph decoder of the graph auto-encoder network, and train it by providing a supervision signal through a link prediction task to obtain the reconstructed association relationship between the regional node and the commodity node, so as to obtain a graph embedding model.
[0030] Optionally, the attraction model is constructed based on the indicators of four dimensions: distance, scale, education quality, and quality of life. The attraction P of educational institution i to any location j ij is expressed as:
[0031]
[0032]
[0033] where c i represents the total number of people in educational institution i, which is used to describe the scale of the educational institution; ra i represents the teacher-student ratio of educational institution i, which is used to describe the education quality of the educational institution; rb i represents the per capita area of educational institution i, which is used to describe the quality of life of students; ω1, ω2, ω3 are the weights of the corresponding factors; R ij represents the distance index from educational institution i to location j, and this distance index is expressed as:
[0034]
[0035] where D ij represents the straight-line distance from educational institution i to location j; D avg is the average distance from educational institutions within the optional range of the city to any location.
[0036] According to another aspect of the embodiments of the present invention, an analysis device for regional education demand is provided.
[0037] An analysis device for regional education demand includes:
[0038] A bipartite graph construction module for constructing a bipartite graph of regions and commodities according to e-commerce platform data;
[0039] A graph embedding model training module for inputting the bipartite graph into a graph auto-encoder network and training to obtain a graph embedding model through a self-supervised task, where the graph embedding model is used to represent the association relationship between regions and commodities;
[0040] An attraction model construction module for constructing an attraction model according to historical enrollment data, where the attraction model is used to predict the future enrollment numbers of each region;
[0041] An education demand calculation module for calculating the education demand of each region based on the output results of the graph embedding model and the attraction model.
[0042] According to another aspect of the embodiments of the present invention, there is provided an electronic device for analyzing regional education needs.
[0043] An electronic device for analyzing regional education needs includes: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the analysis method for regional education needs provided by the embodiments of the present invention.
[0044] According to still another aspect of the embodiments of the present invention, there is provided a computer-readable medium.
[0045] A computer-readable medium has a computer program stored thereon, and when the program is executed by a processor, it implements the analysis method for regional education needs provided by the embodiments of the present invention.
[0046] One embodiment of the above invention has the following advantages or beneficial effects: By constructing a bipartite graph of regions and commodities based on e-commerce platform data; inputting the bipartite graph into a graph autoencoder network, and training through a self-supervised task to obtain a graph embedding model, which is used to represent the association relationship between regions and commodities; constructing an attraction model based on historical enrollment data, which is used to predict the future enrollment numbers of each region; calculating the education needs of each region based on the output results of the graph embedding model and the attraction model. Relying on the high user coverage of the e-commerce platform, the algorithm describes the education needs of residents in the region through users' purchase behaviors, no longer requiring a large amount of cost for demand research, saving manpower, material resources and financial resources, with low cost and short cycle; proposing an attraction operator to model the decision-making mode of users when choosing an educational institution, and describing the historical enrollment numbers of each region by constraining the interaction under the supply-demand relationship, realizing data-driven analysis of regional education needs, constructing a sequence prediction task by combining multiple environmental factors to analyze regional education needs, solving the problem of difficult demand analysis, and providing a data basis for the site selection of educational institutions and the planning and arrangement of educational institutions.
[0047] The further effects of the above non-conventional optional methods will be described in conjunction with specific embodiments hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention. Among them:
[0049] Figure 1 is a schematic diagram of the main steps of the analysis method for regional education needs according to the embodiments of the present invention;
[0050] Figure 2 is a schematic diagram of the implementation principle of the embodiments of the present invention;
[0051] Figure 3 It is a schematic diagram of the principle of the network embedding model according to an embodiment of the present invention;
[0052] Figure 4 It is a schematic diagram of the network structure of the graph encoder according to an embodiment of the present invention;
[0053] Figure 5 It is a schematic diagram of the network structure of the graph decoder according to an embodiment of the present invention;
[0054] Figure 6 It is a schematic diagram of the main modules of the analysis device for regional education needs according to an embodiment of the present invention;
[0055] Figure 7 It is an exemplary system architecture diagram to which an embodiment of the present invention can be applied;
[0056] Figure 8 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. Detailed implementation manners
[0057] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0058] Regarding the prior art technical solutions for realizing education demand prediction and analysis through methods such as manual field surveys and questionnaire statistics, the inventors have analyzed and found that they mainly have the following several defects:
[0059] 1) The research requires a large amount of manpower, material resources and financial resources, with high costs and long cycles;
[0060] 2) Due to the influence of the floating population, the changes in education demand have complex spatio-temporal correlations, are highly correlated with the local development status, economy and geographical environment, which brings difficulties to demand analysis;
[0061] 3) The spatial distribution of educational infrastructure (supply side) will affect the distribution and change trend of demand. For example, policies such as zoned school districts and neighborhood enrollment, etc. Conversely, the change in demand will also lead to changes in the distribution of educational facilities, and it is difficult to consider the interaction of the supply-demand relationship;
[0062] 4) Poor replicability. The long working cycle results in the research results in the current city being unable to be applied to new cities.
[0063] When modeling an education demand prediction analysis model by combining with the currently popular machine learning, only modeling based on past years' data will lack relevant data, resulting in the inability to use more comprehensive data for modeling; the change of education demand is affected by various environmental factors, with complex spatio-temporal dynamic relationships, and the data accumulation is small, making it difficult to construct a solution for sequence prediction tasks.
[0064] In view of the above disadvantages, the present invention designs a data-driven education demand prediction analysis method, which can analyze the education demand of each region based on e-commerce data and school enrollment information. This method can be applied to the site selection and school district planning of kindergartens, primary schools, middle schools, and interest class training institutions. A data-driven education demand prediction analysis method of the present invention, specifically, models the impact of the resident age distribution on enrollment by constructing the correlation between the historical enrollment numbers of each region and the region representation, so as to predict the future education demand quantity. This method has the following advantages: 1) Relying on the high user coverage of the e-commerce platform, the algorithm describes the education demand situation of the residents living in the region through the purchase behavior of users, and there is no need to invest a large amount of cost in demand research; 2) An attraction operator is proposed to model the decision-making mode of users when choosing schools, and the historical enrollment numbers of each region are described by constraining the interaction under the supply-demand relationship; 3) High scalability, only need to collect local order information and school enrollment information to apply, and the functional interaction information described only based on order data is still available.
[0065] Figure 1 It is a schematic diagram of the main steps of the analysis method for regional education demand according to an embodiment of the present invention. As Figure 1 shown, the analysis method for regional education demand according to the embodiment of the present invention mainly includes the following steps S101 to step S104.
[0066] Step S101: Construct a bipartite graph of regions and commodities according to e-commerce platform data;
[0067] Step S102: Input the bipartite graph into a graph autoencoder network, and train a graph embedding model through a self-supervised task. The graph embedding model is used to represent the association relationship between regions and commodities;
[0068] Step S103: Construct an attraction model according to historical enrollment data. The attraction model is used to predict the future enrollment numbers of each region;
[0069] Step S104: Calculate the education demand of each region based on the output results of the graph embedding model and the attraction model.
[0070] Due to the influence of the floating population, it is very difficult to directly investigate educational needs. Although data sources implicitly related to enrollment needs include medical, birth, vaccination data, etc., such data highly involves user privacy and is very difficult to obtain. After investigation, the inventor found that educational needs are closely related to the residents in the region, including their types of work, economic conditions, historical purchased goods, etc. Based on this, it is proposed to use self-supervised tasks to model the representation of the region and capture the educational needs of each region. Accordingly, the present invention provides an analysis method for regional educational needs. The algorithm inputs include local order data, regional portrait data, commodity information data, school enrollment data, etc. The algorithm outputs the number of people with educational needs in each plot this year. The processing process is mainly divided into three parts: 1) data preprocessing, 2) network embedding, and 3) educational needs prediction. When training the model, data from previous years is used. First, an attraction model is used to construct an estimate of the enrollment number in each region; then, after preprocessing, the order data, user portraits (including delivery addresses), commodity information data, etc. are used to form a bipartite graph, which is used as the input of the embedding network for training. After training, the representation vectors of each region are obtained. Finally, the relationship between the representation vectors and the enrollment number is captured through a neural network model. When testing the model, current e-commerce data is used to predict the educational needs of each region in the next year.
[0071] Figure 2 It is a schematic diagram of the implementation principle of an embodiment of the present invention. As Figure 2 shown, which shows the implementation principle of the analysis method for regional educational needs of the present invention. The entire algorithm framework mainly includes three parts:
[0072] 1) Data preprocessing: This part is mainly used to clean, mine, and construct a weighted bipartite graph of users and commodities. In practical applications, to ensure the scale of the graph, commodities can be aggregated by tertiary categories (commodity categories). The region and the commodity are used as the nodes of the bipartite graph respectively, and the regional features and commodity features are used as node features. Order-related nodes form edges, and the weight of the edge is the binned value of the order quantity;
[0073] 2) Network embedding: This module uses a self-supervised graph autoencoder network to represent the relationship between the node at each position and the commodity node, so as to model the initial representation of each region;
[0074] 3) Educational needs prediction: This module includes two parts. First, the enrollment number in each region is estimated through an attraction model. Then, a fully connected network is constructed to capture the relationship between the representation of each region and the enrollment number, so as to predict the future enrollment number in each region.
[0075] According to an embodiment of the present invention, in the data preprocessing stage, a bipartite graph of regions and commodities is mainly constructed based on e-commerce platform data. Specifically, constructing a bipartite graph of regions and commodities based on e-commerce platform data may specifically include the following steps:
[0076] Step S1011: Clean the e-commerce platform data to correct addresses and delete abnormal user data. The e-commerce platform data includes commodity data, regional data, and order data;
[0077] Step S1012: Extract features from the commodity data to construct commodity feature vectors. The commodity data includes commodity self-information and age information;
[0078] Step S1013: Extract features from the regional data to construct regional feature vectors. The regional data includes regional information to which the delivery address of the order belongs;
[0079] Step S1014: Use regions and commodities as nodes of the bipartite graph respectively, use regional feature vectors and commodity feature vectors as node feature vectors, form edges by associating regional nodes and commodity nodes through order data, and use the number of orders for purchasing commodities within a region as the weight of the edge to construct the bipartite graph.
[0080] According to one embodiment of the present invention, the steps for address correction in step S1011 include: using the geocoding services of multiple geographic information service providers to perform longitude and latitude parsing on educational institutions; obtaining the voting scores of each parsing result through voting based on the distribution of the parsing results; determining the final result of the address of the educational institution according to the voting scores and preset rules for address correction. Taking kindergartens as an example of educational institutions, through exploratory analysis of the data, the inventors found that there are significant drifts in the longitude and latitude of kindergartens in the orders, manifested as exceeding the administrative region of the city where they are located, or multiple kindergartens overlapping at the same location, etc. Here, a voting algorithm is used: First, use the geocoding services of multiple (for example: 3) map service providers to perform longitude and latitude parsing on the kindergarten name (longitude and latitude parsing is also called reverse geocoding, input an address and return longitude and latitude, here the input is the address of the kindergarten), and then vote according to the distribution of the parsing results. For example: for a parsing result, if other results are within 50m of it, record the number of votes of this result as 1. Traverse the parsing results of these 3 map service providers respectively, and select the result with the highest number of votes as the final parsing result. If there are ties, by default, use the parsing result of service provider 1 as the final parsing result. If the number of votes is 0 for all (that is: the distances between the other two results and this result are both outside 50m), then record this address in the results that cannot be parsed, and finally process it manually. This method can efficiently handle the vast majority of address anomalies.
[0081] In an embodiment of the present invention, when deleting abnormal user data, it can be achieved by analyzing the user's order data (the time span of the order data is one year). Some users have a very large number of delivery addresses. Thus, a discrete distribution of delivery address vectors is formed according to each user's orders. Users with a distribution interval exceeding 4 and the probability of the highest interval being lower than 0.4 are defined as abnormal users (that is: when the number of delivery addresses used by a user in the past is greater than 4, and the proportion of the number of times of the most frequently used delivery address is lower than 0.4, then the user is determined to be an abnormal user). Through investigation, it is found that such users are more likely to be company or proxy accounts. Their purchase behavior cannot be associated with the age of their children, so they are omitted.
[0082] In step S1012 when constructing the product feature vector, the product data includes product self-information and age information. The product self-information includes the first-level and second-level category types, sales volume, price, etc. of the product; the age information corresponding to the product refers to the age information applicable to the product. When obtaining the age information, since the age information of the product will be marked in the product details page information and the product name, it can be extracted by regular matching. Among them, the age information is an interval, such as 3 - 6 years old, 3 - 15 months old, etc. When performing feature extraction of the age information, a mapping table is maintained for the age information of the product, and the irregular age information can be uniformly mapped into an 18-dimensional indicator vector. Each dimension of this vector represents 0 - 1 year old, 1 - 2 years old, and so on. For products with multiple age information, the present invention performs a bitwise OR operation on the vectors corresponding to each age to obtain the intermediate result of each product. For example: using an 18-dimensional vector to represent 0 - 18 years old, for 3 - 6 years old (including 6 years old), it can be represented as an 18-dimensional vector: [0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0].
[0083] According to another embodiment of the present invention, after step S1012 extracts features from the product data to construct the product feature vector, the product data can also be aggregated according to the product category to which the product belongs, and a bipartite graph is constructed based on the aggregated products. When constructing the bipartite graph, in order to avoid the bipartite graph structure being too large, resulting in sparse edges and reducing the effect of the self-supervised task, the present invention aggregates the products into category products according to the third-level category, calculates the average value of the features of the nodes for aggregation, and directly omits the categories with sparse quantities.
[0084] In an embodiment of the present invention, when extracting features from the regional data to construct the regional feature vector, each region represents a community and is initialized as a one-hot vector. The regional data can be extracted from the user information of the e-commerce platform or can be regional data such as the delivery address obtained in combination with the order data.
[0085] According to an embodiment of the present invention, taking the order quantity of purchased goods within a region as the weight of an edge includes: binning the order quantity of purchased goods within the region to obtain multiple bin values, and establishing a mapping relationship between the order quantity and the bin values; taking the bin value corresponding to the order quantity as the weight of the edge. In the embodiment of the present invention, after determining the graph nodes, edges are constructed according to the purchase relationship, and the initial weight of the edge is the frequency of purchasing a certain category of goods at a certain location. The edges of the bipartite graph are represented by the purchase relationship. If a certain community purchases a certain category 150 times, there will be an edge with a weight of 150 between these two corresponding nodes. Since the decoder of the self-supervised task corresponds to a multi-classification task, the frequency will undoubtedly increase its training difficulty and make the supervision signal sparser. Therefore, in the present invention, binning is performed according to the frequency, divided into 5 levels, and the bin value is taken as the weight of the edge, which can improve the stability of subsequent embedding learning.
[0086] According to another embodiment of the present invention, when step S102 inputs the bipartite graph into the graph autoencoder network and trains to obtain a graph embedding model through a self-supervised task, it may specifically include:
[0087] Input the bipartite graph into the graph encoder of the graph autoencoder network, extract structural embedding information through graph convolution, and fuse the structural predecessor information and the features of the nodes through a fully connected layer to obtain the hidden layer representation of the nodes;
[0088] Input the hidden layer representation of the nodes into the graph decoder of the graph autoencoder network, and provide a supervision signal through a link prediction task for training to obtain the reconstructed association relationship between the regional nodes and the commodity nodes, so as to obtain a graph embedding model.
[0089] In the embodiment of the present invention, the graph autoencoder network is, for example, a model referred to in "Graph Convolutional Matrix Completion". Figure 3 It is a schematic diagram of the principle of the network embedding model in the embodiment of the present invention. The network embedding model includes a graph encoder and a graph decoder. Input the weighted bipartite graph into the graph encoder to obtain the hidden layer representation of the nodes, and then provide a supervision signal through a link prediction task for training. The reconstruction loss is the cross-entropy on the link prediction task (for multi-classification on each edge). The link prediction task refers to predicting whether there is an edge between the commodity node and the location node and what the value of the edge is. The physical meaning is to predict the frequency level of a certain community purchasing a certain commodity. The reconstruction loss represents the loss function of this self-supervised model. Since it involves an encoder-decoder structure here, it is called the reconstruction loss.
[0090] The encoder part extracts the structural embedding information through graph convolution to implement message passing. This information is fused with the features of the node itself through a fully connected layer to obtain the embedding representation of the node. The embedding representation is the output of the graph encoder, representing the vector obtained by transforming the original features. This vector can be understood as the key information obtained by refining and summarizing the original features.
[0091] Figure 4 It is a schematic diagram of the network structure of the graph encoder in an embodiment of the present invention. As Figure 4 , the adjacency matrix of the input bipartite graph and the feature matrix of the nodes (side information, relative to interaction information, representing the information on the node side in the bipartite graph) are input. The graph convolution uses a simplified version designed based on the message passing mechanism, expressed as:
[0092]
[0093] where r represents the scoring level, that is, the result of binning during network construction, and levels 1 - 5 can be considered. c ij represents the normalization factor, and here symmetric normalization is used ( N i represents the degree of the current node, N j represents the degree of the corresponding node of the current node), and W r is the learnable parameter matrix at this scoring level. X i is the feature of the current node (one - hot vector). This formula means that for node j, the information passed from its adjacent node i at level r is μ i→j,r . For information at different scoring levels, it is necessary to aggregate it, expressed as:
[0094]
[0095] accum(·) represents the aggregation operation, and here summation is used. Finally, a low - dimensional mapping is performed on the message value:
[0096] h j = σ(W r h j ).
[0097] In the graph encoder, two convolutional layers are stacked to extract the structural information of the graph, and dropout is performed before the input of each convolutional layer to ensure the regularization performance of the model.
[0098] In addition, in order to model the feature information of the nodes, the present invention designs a fully connected layer to fuse the structural features and the node features, expressed as:
[0099] z j = σ(Wh j + W2f j ), f j = σ(W1X j + b), j ∈ N L ;
[0100] f j represents the representation of the features of node j itself after passing through a fully connected layer, and z j represents the final representation after fusing the structural information h j and its own features f j for describing the age-related embeddings of each node in the bipartite graph. After the training of the entire GCMC model is completed, this embedding is used as the input of the attraction operator.
[0101] In the embodiments of the present invention, the decoder part inputs the embedding representations of each node and outputs the reconstructed relational matrix The entire model is trained by cross-entropy on multi-classification.
[0102] Figure 5 is a schematic diagram of the network structure of the graph decoder in the embodiments of the present invention. As Figure 5 , the matrix factorization module trains a relational matrix for correlating the embedding information of positions and commodities. Generally, a separate relational matrix needs to be trained for each rating level. However, due to the large difference in the amount of edge data on each rating level, the training difficulty for rating levels with scarce data will be very high. The present invention proposes a way of matrix factorization to improve the performance by weighted representation of a trainable basic parameter matrix, expressed as:
[0103]
[0104] Finally, the correlation result calculated by the node calculates the distribution of the current relational rating level through softmax, and the probability for level r can be expressed as:
[0105]
[0106] The optimization objective of the entire model is to minimize the following formula:
[0107]
[0108] Among them, I[·] is an indicator function used to identify the edges that exist during data collection but are masked during training to provide a supervision signal. These edges are represented as 1, otherwise 0. For these edges, if viewed individually according to the scoring level, the optimization objective of the present invention is to minimize the negative log-likelihood of a certain level of this edge. If the scoring is regarded as a classification, the optimization objective of the present invention is to minimize the cross-entropy between the true category and the calculation result.
[0109] According to another embodiment of the present invention, when predicting and analyzing educational needs, an attraction operator is introduced in the present invention to model the preference relationship of users when choosing educational institutions, so as to estimate the enrollment numbers in each region. The attraction model of the present invention is constructed based on indicators in four dimensions: distance, scale, educational quality, and quality of life. The attraction P of educational institution i to any location j ij is expressed as:
[0110]
[0111]
[0112] where c i represents the total number of people in educational institution i and is used to describe the scale of the educational institution; ra i represents the teacher-student ratio of educational institution i and is used to describe the educational quality of the educational institution; rb i represents the per capita area of educational institution i and is used to describe the quality of life of students; these three-dimensional indicators are respectively normalized, and ω1, ω2, ω3 are the weights of the corresponding factors (empirical parameters, default settings are 0.6, 0.2, 0.2); R ij represents the distance exponent from educational institution i to location j, and this distance exponent is expressed as:
[0113]
[0114] where D ij represents the straight-line distance from educational institution i to location j, with the unit of meter; D avg is the average distance of educational institutions within the optional range of this city (for example: 10 kilometers) to any location. The optional range is the maximum limit distance when people choose educational institutions. D avg is the tolerable range when choosing educational institutions in this city, also known as the truncation distance. It means that when the distance from an educational institution to a location exceeds the truncation distance, the attraction of this educational institution to users at the corresponding location will immediately decrease, which is reflected in the denominator value in the formula being greater than 1, causing the attraction value to decrease. On the contrary, when the denominator value is less than 1, the attraction value increases. The Bin(·) function bins the attraction according to the numerical value when the denominator value is less than 1 to prevent the attraction from expanding due to the square operation. P ijRepresents the selection probability calculated based on the attraction. From this, we can list the equation. Let the number of students admitted to each community be x, and the attraction of the i-th educational institution (e.g., kindergarten) to the j-th community be p ij , the number of students admitted to each school is y, y is known, there are m schools and n communities, then the equation is listed as:
[0115]
[0116] where x ji is the component of the j-th community in the i-th educational institution, representing the number of students from the j-th community attending school in the i-th educational institution. If considering the school district constraint, this equation can be separately listed for each school district to solve for the number of students admitted in each community.
[0117] When making demand predictions, for each community, a model is established using the representation vector of the previous year and the number of students admitted in the next year. The model is implemented using a two-layer fully connected network, and finally, the educational demand for each community in the next year is predicted based on the representation vectors of each community in the current year.
[0118] Figure 6 is a schematic diagram of the main modules of the regional educational demand analysis device according to an embodiment of the present invention. As Figure 6 shown, the regional educational demand analysis device 600 according to an embodiment of the present invention mainly includes a bipartite graph construction module 601, a graph embedding model training module 602, an attraction model construction module 603, and an educational demand calculation module 604.
[0119] The bipartite graph construction module 601 is used to construct a bipartite graph of regions and commodities based on e-commerce platform data;
[0120] The graph embedding model training module 602 is used to input the bipartite graph into a graph autoencoder network and train a graph embedding model through a self-supervised task. The graph embedding model is used to represent the association relationship between regions and commodities;
[0121] The attraction model construction module 603 is used to construct an attraction model based on historical enrollment data. The attraction model is used to predict the future enrollment numbers in each region;
[0122] The educational demand calculation module 604 is used to calculate the educational demand in each region based on the output results of the graph embedding model and the attraction model.
[0123] According to an embodiment of the present invention, the bipartite graph construction module 601 can also be used for:
[0124] Clean the e-commerce platform data for address correction and deletion of abnormal user data. The e-commerce platform data includes commodity data, regional data, and order data;
[0125] Extract features from the commodity data to construct a commodity feature vector, where the commodity data includes commodity self-information and age information;
[0126] Extract features from the regional data to construct a regional feature vector, where the regional data includes the regional information to which the delivery address of the order belongs;
[0127] Take the region and the commodity as the nodes of the bipartite graph respectively, take the regional feature vector and the commodity feature vector as the node feature vectors, form edges by associating the regional node and the commodity node through the order data, and take the number of orders for purchasing commodities within the region as the weight of the edge to construct a bipartite graph.
[0128] According to another embodiment of the present invention, after the bipartite graph construction module 601 extracts features from the commodity data to construct a commodity feature vector, it can also be used for:
[0129] For the commodity data, aggregate the commodities according to the commodity categories to which they belong, and construct a bipartite graph based on the aggregated commodities.
[0130] According to still another embodiment of the present invention, when the bipartite graph construction module 601 corrects the address, it can also be used for:
[0131] Use the geographical parsing services of multiple geographical information service providers to parse the longitude and latitude of educational institutions;
[0132] Obtain the voting scores of each parsing result by voting according to the distribution of the parsing results;
[0133] Determine the final result of the address of the educational institution according to the voting scores and preset rules to correct the address.
[0134] According to still another embodiment of the present invention, when the bipartite graph construction module 601 takes the number of orders for purchasing commodities within the region as the weight of the edge, it can also be used for:
[0135] Bin the number of orders for purchasing commodities within the region to obtain multiple bin values, and establish a mapping relationship between the order quantity and the bin values;
[0136] Take the bin value corresponding to the order quantity as the weight of the edge.
[0137] According to still another embodiment of the present invention, the graph embedding model training module 602 can also be used for:
[0138] Input the bipartite graph into the graph encoder of the graph auto-encoding network, extract the structural embedding information through graph convolution, and fuse the structural predecessor information with the features of the nodes through a fully connected layer to obtain the hidden layer representation of the nodes;
[0139] Input the hidden layer representation of the node into the graph decoder of the graph auto-encoder network, and train it by providing a supervision signal through the link prediction task to obtain the association relationship between the reconstructed regional node and the commodity node, so as to obtain the graph embedding model.
[0140] According to another embodiment of the present invention, the attraction model is constructed based on the indexes of four dimensions: distance, scale, education quality, and quality of life. The attraction P of educational institution i to any location j ij is expressed as:
[0141]
[0142]
[0143] where c i represents the total number of people in educational institution i, which is used to describe the scale of the educational institution; ra i represents the teacher-student ratio of educational institution i, which is used to describe the education quality of the educational institution; rb i represents the per capita area of educational institution i, which is used to describe the quality of life of students; ω1, ω2, ω3 are the weights of the corresponding factors; R ij represents the distance index from educational institution i to location j, and this distance index is expressed as:
[0144]
[0145] where D ij represents the straight-line distance from educational institution i to location j; D avg is the average distance of the educational institutions within the optional range of the city to any location.
[0146] According to the technical solution of the embodiment of the present invention, a bipartite graph of regions and commodities is constructed based on e-commerce platform data; the bipartite graph is input into a graph auto-encoding network, and a graph embedding model is obtained through self-supervised task training. The graph embedding model is used to represent the association relationship between regions and commodities; an attraction model is constructed based on historical enrollment data, and the attraction model is used to predict the future enrollment numbers of each region; based on the output results of the graph embedding model and the attraction model, the educational demands of each region are calculated. Relying on the high user coverage of the e-commerce platform, the algorithm describes the educational demand situation of the residents in the region through the purchase behavior of users, and there is no need to invest a large amount of cost in demand research, saving manpower, material resources and financial resources, with low cost and short cycle; an attraction operator is proposed, which is used to model the decision-making mode of users when choosing educational institutions, and describes the historical enrollment numbers of each region by constraining the interaction under the supply-demand relationship, realizing data-driven analysis of regional educational demands, constructing a sequence prediction task by combining various environmental factors to analyze regional educational demands, solving the problem of difficult demand analysis, and providing a data basis for the site selection of educational institutions and the planning and arrangement of educational institutions.
[0147] Figure 7 An exemplary system architecture 700 is shown that can apply the analysis method of regional educational demand or the analysis device of regional educational demand according to the embodiment of the present invention.
[0148] As Figure 7 shown, the system architecture 700 may include terminal devices 701, 702, 703, a network 704, and a server 705. The network 704 is used to provide a medium for communication links between the terminal devices 701, 702, 703 and the server 705. The network 704 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0149] Users can use the terminal devices 701, 702, 703 to interact with the server 705 through the network 704 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 701, 702, 703, such as shopping applications, e-commerce platform applications, search applications, data processing tools, social platform software, etc. (only as examples).
[0150] The terminal devices 701, 702, 703 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0151] The server 705 may be a server that provides various services. For example, it may be a back-end management server (only for illustration) that supports shopping websites browsed by users using terminal devices 701, 702, and 703. The back-end management server may construct a bipartite graph of regions and commodities based on e-commerce platform data for the received data such as analysis requests for regional education needs; input the bipartite graph into a graph auto-encoding network, and obtain a graph embedding model through self-supervised tasks. The graph embedding model is used to represent the association relationship between regions and commodities; construct an attraction model based on historical enrollment data. The attraction model is used to predict the future enrollment numbers of each region; calculate the education needs of each region and other processes based on the output results of the graph embedding model and the attraction model, and feedback the processing results (such as the education needs of each region - only for illustration) to the terminal device.
[0152] It should be noted that the method for analyzing regional education needs provided by the embodiments of the present invention is generally executed by the server 705. Correspondingly, the device for analyzing regional education needs is generally set in the server 705.
[0153] It should be understood that Figure 7 the numbers of terminal devices, networks, and servers in
[0154] are merely illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers. Figure 8 Reference is now made to Figure 8 which shows a schematic structural diagram of a computer system 800 suitable for implementing the terminal device or server of the embodiments of the present invention.
[0155] As Figure 8 shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage section 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the system 800 are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.
[0156] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as required. A removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 810 as required so that a computer program read therefrom is installed into the storage section 808 as required.
[0157] Specifically, according to the embodiments disclosed by the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed by the present invention include a computer program product which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by a central processing unit (CPU) 801, the above-described functions defined in the system of the present invention are performed.
[0158] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0160] The units or modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described units or modules can also be provided in a processor. For example, it can be described as: a processor includes a bipartite graph construction module, a graph embedding model training module, an attraction model construction module, and an education demand calculation module. Among them, the names of these units or modules do not constitute a limitation to the units or modules themselves in some cases. For example, the bipartite graph construction module can also be described as "a module for constructing a bipartite graph of regions and commodities according to e-commerce platform data".
[0161] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device includes: constructing a bipartite graph of regions and commodities according to e-commerce platform data; inputting the bipartite graph into a graph autoencoder network, and training through a self-supervised task to obtain a graph embedding model, where the graph embedding model is used to represent the association relationship between regions and commodities; constructing an attraction model according to historical enrollment data, where the attraction model is used to predict the future enrollment numbers of each region; and calculating the education demand of each region based on the output results of the graph embedding model and the attraction model.
[0162] According to the technical solution of the embodiments of the present invention, by constructing a bipartite graph of regions and commodities according to e-commerce platform data; inputting the bipartite graph into a graph autoencoder network, and training through a self-supervised task to obtain a graph embedding model, where the graph embedding model is used to represent the association relationship between regions and commodities; constructing an attraction model according to historical enrollment data, where the attraction model is used to predict the future enrollment numbers of each region; and calculating the education demand of each region based on the output results of the graph embedding model and the attraction model, relying on the high user coverage of the e-commerce platform, the algorithm describes the education demand situation of the residents in the region through the purchase behavior of users, no longer requires a large amount of cost for demand research, saves manpower, material resources and financial resources, has low cost and short cycle; proposes an attraction operator for modeling the decision-making mode of users when choosing an educational institution, and describes the historical enrollment numbers of each region by restricting the interaction under the supply-demand relationship, realizes data-driven regional education demand analysis, constructs a sequence prediction task by combining various environmental factors to analyze regional education demand, solves the problem of difficult demand analysis, and provides a data basis for the location selection of educational institutions and the planning and arrangement of educational institutions.
[0163] The above specific embodiments do not limit the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for analyzing regional education needs, characterized in that, Including: Constructing a bipartite graph of regions and commodities based on e-commerce platform data, where the e-commerce platform data includes commodity data, regional data, and order data; Constructing a bipartite graph of regions and commodities based on e-commerce platform data includes: extracting features from the commodity data to construct commodity feature vectors, extracting features from the regional data to construct regional feature vectors, using regions and commodities as nodes of the bipartite graph respectively, using the regional feature vectors and commodity feature vectors as node feature vectors, associating the regional nodes and commodity nodes through order data to form edges, using the number of orders for purchasing commodities within a region as the weight of the edge, and constructing the bipartite graph; the commodity data includes commodity self-information and age information applicable to the commodity; Inputting the bipartite graph into a graph autoencoder network and training through a self-supervised task to obtain a graph embedding model, where the graph embedding model is used to represent the association relationship between regions and commodities; Constructing an attraction model based on historical enrollment data, where the attraction model is used to predict the future enrollment numbers of each region, and the attraction model is a model constructed based on four-dimensional indicators including the scale of educational institutions, the educational quality of educational institutions, the quality of life of students in educational institutions, and the distance index from educational institutions to regions for determining the attraction of educational institutions to users at corresponding locations; among them, the total number of people in the educational institution is used to describe the scale of the educational institution; the student-teacher ratio of the educational institution is used to describe the educational quality of the educational institution; the per capita area of the educational institution is used to describe the quality of life of students; Calculating the educational demand of each region based on the output results of the graph embedding model and the attraction model, including: obtaining the representation vectors of each region based on the graph embedding model, estimating the enrollment numbers of each region through the attraction model, capturing the relationship between the representation vectors of each region and the enrollment numbers through a fully connected network, and predicting the future enrollment numbers of each region accordingly.
2. The method according to claim 1, wherein Constructing a bipartite graph of regions and commodities based on e-commerce platform data further includes: Performing data cleaning on the e-commerce platform data to correct addresses and delete abnormal user data; The commodity data includes commodity self-information and age information; The regional data includes the regional information to which the delivery address of the order belongs.
3. The method according to claim 2, wherein After extracting features from the commodity data to construct commodity feature vectors, it further includes: Aggregating the commodities according to the commodity categories to which they belong for the commodity data, and constructing a bipartite graph based on the aggregated commodities.
4. The method according to claim 2, characterized in that, The steps for address correction include: Using the geographical parsing services of multiple geographical information service providers to perform longitude and latitude parsing on educational institutions; Obtaining the voting scores of each parsing result by voting according to the distribution of the parsing results; Determining the final result of the address of the educational institution according to the voting scores and preset rules for address correction.
5. The method according to claim 2, wherein Using the number of orders for purchasing commodities within a region as the weight of the edge includes: Binning the number of orders for purchasing commodities within a region to obtain multiple bin values, and establishing a mapping relationship between the order quantity and the bin values; Using the bin value corresponding to the order quantity as the weight of the edge.
6. The method according to claim 1, wherein Inputting the bipartite graph into a graph autoencoder network and training through a self-supervised task to obtain a graph embedding model includes: Input the bipartite graph into the graph encoder of the graph auto-encoder network, extract the structural embedding information through graph convolution, and fuse the structural embedding information with the features of the nodes through a fully connected layer to obtain the hidden layer representation of the nodes; Input the hidden layer representation of the nodes into the graph decoder of the graph auto-encoder network, and train it through a link prediction task to provide a supervision signal to obtain the reconstructed association relationship between the regional nodes and the commodity nodes, so as to obtain a graph embedding model.
7. The method according to claim 1, characterized in that The attractiveness P of educational institution i to any position j ij is expressed as: ; ; Among them, represents the total number of people in educational institution i, which is used to describe the scale of the educational institution; represents the teacher-student ratio of educational institution i, which is used to describe the educational quality of the educational institution; represents the per capita area of educational institution i, which is used to describe the living quality of students; is the weight of the corresponding factor; represents the distance index from educational institution i to location j, and this distance index is expressed as: ; Among them, represents the straight-line distance from educational institution i to location j; is the average distance from educational institutions within the optional range of the current city to any location.
8. An analysis device for regional education needs, characterized in that It includes: A bipartite graph construction module for constructing a bipartite graph of regions and commodities according to e-commerce platform data, where the e-commerce platform data includes commodity data, regional data, and order data; Constructing a bipartite graph of regions and commodities according to e-commerce platform data includes: extracting features from the commodity data to construct commodity feature vectors, extracting features from the regional data to construct regional feature vectors, using regions and commodities as nodes of the bipartite graph respectively, using regional feature vectors and commodity feature vectors as node feature vectors, associating regional nodes and commodity nodes through order data to form edges, and using the number of orders for purchasing commodities within a region as the weight of the edges to construct a bipartite graph; the commodity data includes commodity own information and age information applicable to the commodity; A graph embedding model training module for inputting the bipartite graph into a graph auto-encoder network and training through a self-supervised task to obtain a graph embedding model, where the graph embedding model is used to represent the association relationship between regions and commodities; An attraction model construction module for constructing an attraction model according to historical enrollment data, where the attraction model is used to predict the future enrollment numbers of each region, and the attraction model is a model for determining the attraction of an educational institution to users at a corresponding location based on four-dimensional indicators of the scale of the educational institution, the educational quality of the educational institution, the quality of life of the students in the educational institution, and the distance index from the educational institution to the region; among them, the total number of people in the educational institution is used to describe the scale of the educational institution; the student-teacher ratio of the educational institution is used to describe the educational quality of the educational institution; the per capita area of the educational institution is used to describe the quality of life of the students; An education demand calculation module for calculating the education demand of each region based on the output results of the graph embedding model and the attraction model, including: obtaining the representation vectors of each region based on the graph embedding model, estimating the enrollment numbers of each region through the attraction model, and capturing the relationship between the representation vectors of each region and the enrollment numbers through a fully connected network, so as to predict the future enrollment numbers of each region.
9. An electronic device for analyzing regional education needs, characterized in that, It includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
An educational big data analysis method based on artificial intelligence
AU2020103529A4
Regional feature acquisition method and device, computer equipment and medium
CN111125272A