Grid tag classification method and device, electronic equipment and storage medium

By constructing a first heterogeneous graph of job seekers and positions and a second heterogeneous graph of recruiters and positions, determining grid labels and performing clustering, the problem of poor matching of regional recommendation strategies in recruitment software is solved, and more accurate traffic estimation and recommendation effects are achieved.

CN119397321BActive Publication Date: 2025-10-17QIAN JIN NETWORK INFORMATION TECH SHANGHAI LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411502261.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2025-10-17
Estimated Expiration
2044-10-25

AI Technical Summary

Technical Problem

Existing recruitment software has poor matching of recommendation strategies in different regions, resulting in mediocre recommendation results and difficulty in accurately estimating the number of potential job seekers.

Method used

By constructing a first heterogeneous graph between job seekers and positions and a second heterogeneous graph between recruiters and positions, the vector information of nodes is obtained, the grid labels are determined, and the target clusters are generated based on the clustering algorithm, thereby improving the flexibility of area division and the accuracy of traffic estimation.

Benefits of technology

It achieves more accurate regional division and traffic estimation, reduces sales costs, and improves the convenience and recommendation effect of advertising sales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119397321B_ABST
    Figure CN119397321B_ABST
Patent Text Reader

Abstract

The application discloses a grid label classification method and device, electronic equipment and a storage medium. The method comprises: obtaining a first heterogeneous graph corresponding to job-seeking behavior information and a second heterogeneous graph corresponding to recruitment behavior information, and determining vector information of each node in each heterogeneous graph; determining a grid label associated with each position node included in the first heterogeneous graph and the second heterogeneous graph; determining a grid label associated with each job seeker node included in the first heterogeneous graph and the second heterogeneous graph; for each grid label, selecting N target position nodes and M target job seeker nodes; determining an aggregated vector of each grid label according to the vector information of the N target position nodes and the vector information of the M target job seeker nodes corresponding to each grid label; and performing clustering processing on the plurality of aggregated vectors and determining a target number of target clustering clusters. The embodiment of the application can improve the flexibility of regional division and is beneficial to improving the accuracy of estimated traffic in the divided region.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a grid label classification method and device, electronic equipment and a storage medium. BACKGROUND

[0002] With the popularity of the Internet and the increasing maturity of data processing technology, more and more convenience is brought to the life and work of users. For example, in the recruitment industry, recruiters and job seekers can use recruitment software to find the required objects. At present, in order to improve the convenience of recruiters and job seekers using the recruitment software, the recruitment software can provide business services such as information promotion, for example, improving the display position, display times, etc. of the positions published by the recruiters on the home page.

[0003] The skilled in the art find that the distribution quantity and the intensive degree of different positions in different cities, districts, counties and other regions are slightly different. Therefore, the recommendation strategies used by different regions are basically the same when recommending positions, which leads to obvious differences in promotion effect of the same position in different regions. In the related technology, different regions are often finely divided in order to more targetedly promote information, but since the number of positions covered by the finely divided regions is small, it is more difficult to accurately estimate the traffic of the regions, in other words, it is difficult to estimate the number of potential job seekers that can be selected by the recruiters, therefore, the matching degree between the current recommendation strategy and the region is poor, which leads to general recommendation effect. SUMMARY

[0004] Therefore, the embodiments of the present application provide a grid label classification method and device, electronic equipment and a storage medium, which can improve the flexibility of region division, and are beneficial to improve the accuracy of estimating the traffic in the divided regions. In addition, based on the accurate division result, the subsequent sales cost can be reduced, and the integration of grid selling and recommendation can be realized, which is beneficial to improve the convenience of subsequent advertisement selling and the recommendation and promotion effect of advertisements.

[0005] The embodiment of the application provides a kind of grid label classification method, comprising: obtaining the first heterogeneous graph corresponding to the job-seeking behavior information in the preset time period and the second heterogeneous graph corresponding to the recruitment behavior information in the preset time period, and determining the vector information of each node in the first heterogeneous graph and the vector information of each node in the second heterogeneous graph, wherein the first heterogeneous graph includes the first relationship information between job seeker node and position node, and the second heterogeneous graph includes the second relationship information between job seeker node and position node;Obtain recruitment information in the preset time period, and determine the grid label associated with each position node included in the first heterogeneous graph and the second heterogeneous graph according to the recruitment information;And, determine the grid label associated with each job seeker node included in the first heterogeneous graph and the second heterogeneous graph according to the job-seeking behavior information and the recruitment information;For each grid label, select N target position nodes and M target job seeker nodes in the nodes included in the first heterogeneous graph and the second heterogeneous graph, wherein the N target position nodes include the position node associated with the grid label, and the M target job seeker nodes include the job seeker node associated with the grid label;Determine the aggregation vector of each grid label according to the vector information of N target position nodes and the vector information of M target job seeker nodes corresponding to each grid label;Multiple aggregation vectors are clustered, and a target number of target cluster clusters are determined, wherein the target number is determined according to the classification fineness of business application requirement, and different target cluster clusters correspond to different grid categories.

[0006] Optionally, according to the method of the embodiment of the application, the job-seeking behavior information includes viewing position behavior and resume submission behavior;The first heterogeneous graph corresponding to the job-seeking behavior information in the preset time period is obtained, comprising: determining the first relationship information between the job seeker node and the position node according to the viewing position behavior and the resume submission behavior;According to the first relationship information between the job seeker node and the position node, the first heterogeneous graph is generated.

[0007] Optionally, according to the method of the embodiment of the application, the recruitment behavior information includes viewing resume behavior and recruitment communication behavior;The second heterogeneous graph corresponding to the recruitment behavior information in the preset time period is obtained, comprising: determining the second relationship information between the job seeker node and the position node according to the viewing resume behavior and the recruitment communication behavior;According to the second relationship information between the job seeker node and the position node, the second heterogeneous graph is generated.

[0008] Optionally, according to the method of the embodiment of the application, determining the vector information of each node in the first heterogeneous graph and the vector information of each node in the second heterogeneous graph comprises: generating a training sample set according to the first heterogeneous graph and the second heterogeneous graph;According to the training sample set, the pre-constructed vector generation model is trained, and the target vector generation model is obtained after reaching the preset training condition;The vector information of each node in the first heterogeneous graph and the vector information of each node in the second heterogeneous graph are determined by the target vector generation model.

[0009] Optionally, according to the method of the embodiment of the present application, the training sample set is generated according to the first heterogeneous graph and the second heterogeneous graph, including: walking the first heterogeneous graph according to the first meta path, and walking the second heterogeneous graph according to the second meta path to obtain a walking sequence set, the walking sequence set including a plurality of walking sequences corresponding to each meta path; and generating the training sample set according to the order of nodes in each walking sequence in the walking sequence set and the K nodes adjacent to the nodes, K being a positive integer.

[0010] Optionally, according to the method of the embodiment of the present application, for each grid label, N target position nodes and M target job seeker nodes are selected from the nodes included in the first heterogeneous graph and the second heterogeneous graph, including: for each grid label, obtaining the first position nodes associated with the grid label and the first job seeker nodes associated with the grid label in the first heterogeneous graph to obtain a first node set; and obtaining the second position nodes associated with the grid label and the second job seeker nodes associated with the grid label in the second heterogeneous graph to obtain a second node set; and for each grid label, selecting N target position nodes from the position nodes included in the first node set and the second node set, and selecting M target job seeker nodes from the job seeker nodes included in the first node set and the second node set.

[0011] Optionally, according to the method of the embodiment of the present application, the plurality of aggregated vectors are clustered and a target number of target cluster clusters are determined, including: randomly selecting a target number of first aggregated vectors in the plurality of aggregated vectors as initial cluster centers of a target number of cluster clusters; and iteratively updating each cluster cluster and the cluster center of the cluster cluster according to a preset clustering algorithm until a preset stop iteration condition is reached to obtain the target cluster cluster; wherein the second aggregated vector is an aggregated vector other than the first aggregated vector in the plurality of aggregated vectors.

[0012] The embodiment of the present application provides a grid label classification device, including: an acquisition module, configured to acquire a first heterogeneous graph corresponding to job-seeking behavior information in a preset time period and a second heterogeneous graph corresponding to recruitment behavior information in the preset time period, and determine vector information of each node in the first heterogeneous graph and vector information of each node in the second heterogeneous graph, wherein the first heterogeneous graph includes first relationship information between job seeker nodes and position nodes, and the second heterogeneous graph includes second relationship information between job seeker nodes and position nodes.

[0013] The acquisition module is further configured to acquire recruitment information in the preset time period, and determine, according to the recruitment information, a grid label associated with each position node included in the first heterogeneous graph and the second heterogeneous graph; and determine, according to the job-seeking behavior information and the recruitment information, a grid label associated with each job seeker node included in the first heterogeneous graph and the second heterogeneous graph.

[0014] The processing module is configured to, for each grid label, select N target position nodes and M target job seeker nodes from nodes included in the first heterogeneous graph and the second heterogeneous graph, respectively, wherein the N target position nodes include a position node associated with the grid label, and the M target job seeker nodes include a job seeker node associated with the grid label.

[0015] The processing module is further configured to determine an aggregated vector of each grid label according to vector information of the N target position nodes and vector information of the M target job seeker nodes corresponding to each grid label.

[0016] The processing module is further configured to perform clustering processing on the plurality of aggregated vectors and determine a target number of target clustering clusters, wherein the target number is determined according to classification fineness of a business application requirement, and different target clustering clusters correspond to different grid categories.

[0017] An electronic device is provided in an embodiment of the present application, and the electronic device includes a processor and a memory storing computer program instructions; the processor implements the steps of the method as above when executing the computer program instructions.

[0018] A computer readable storage medium is provided in an embodiment of the present application, and the computer readable storage medium stores computer program instructions; the computer program instructions are executed by a processor to implement the steps of the method as above.

[0019] A computer program product is provided in an embodiment of the present application, and the computer program product includes computer program instructions; the computer program instructions are executed by a processor to implement the steps of the method as above.

[0020] According to an embodiment of the present application, a heterogeneous graph is constructed based on the job-seeking behavior information and published recruitment information acquired within a preset time period. This graph conveniently represents the relationship between job seekers and various positions from the perspective of job seekers, and the relationship between various positions and job seekers from the perspective of recruiters. The published recruitment information includes the grids where the positions are located, thereby conveniently determining the grid labels associated with each position node in each heterogeneous graph. The existing job-seeking behavior information includes the positions that the job seekers have interacted with. Therefore, combined with the recruitment information, the grid labels corresponding to the job seeker nodes in each heterogeneous graph can be conveniently determined. Next, for each grid label, the associated N target position nodes and M target job seeker nodes are found. Combined with the vector information, an aggregation vector for each grid label is determined. Finally, the number of positions required to be covered within each grid category is determined, and the target number of clusters is flexibly determined. Multiple aggregation vectors are automatically processed into a target number of highly differentiated clusters of different categories, thereby improving the accuracy of traffic estimation within the divided area and the efficiency of grid classification processing, achieving more accurate grid classification and merging. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings in the embodiments of the present application.

[0022] Figure 1 It is a schematic diagram of the system architecture of an embodiment of the present application.

[0023] Figure 2 This is a flowchart of the grid label classification method according to an embodiment of the present application.

[0024] Figure 3 This is a schematic diagram of a first isomeric graph provided in an embodiment of the present application.

[0025] Figure 4 This is a schematic diagram of a second isomeric graph provided in an embodiment of the present application.

[0026] Figure 5 This is a schematic diagram of a walking heterogeneous graph provided in an embodiment of the present application.

[0027] Figure 6 This is a schematic diagram of another first isomeric graph provided in an embodiment of the present application.

[0028] Figure 7 This is a schematic diagram of another second isomeric graph provided in an embodiment of the present application.

[0029] Figure 8 This is a flow chart of another grid label classification method provided in an embodiment of the present application.

[0030] Figure 9 is a structural block diagram of a sorting device of a grid tag according to an embodiment of the present application.

[0031] Figure 10 is a schematic diagram of an electronic device for implementing a sorting method of a grid tag according to an embodiment of the present application. DETAILED DESCRIPTION

[0032] The principles and spirits of the present application will be described below with reference to several exemplary embodiments. It should be understood that the purpose of providing these embodiments is to make the principles and spirits of the present application clearer and more thorough, so that those skilled in the art can better understand and implement the principles and spirits of the present application. The exemplary embodiments provided herein are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments herein, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0033] In this document, terms such as first, second, third, etc. are used to distinguish one entity (or operation) from another entity (or operation), and do not imply or suggest any order or association between the entities (or operations).

[0034] Embodiments of the present application relate to terminal devices and / or servers. Those skilled in the art know that embodiments of the present application can be implemented as a system, device, apparatus, method, computer readable storage medium or computer program product. Therefore, the present disclosure can be embodied in at least one of the following forms: complete hardware, complete software, or a combination of hardware and software. According to embodiments of the present application, a sorting method, device, electronic device and storage medium of a grid tag are claimed. Figure 1 A schematic diagram of a system architecture according to an embodiment of the present application is shown. As shown in FIG. 1, the system architecture includes a terminal device 100 and a server 200. Figure 1As shown, the system includes a terminal device 102 and a server 104. The terminal device 102 can include at least one of a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart television, various wearable devices, an augmented reality (AR) device, a virtual reality (VR) device, and the like. The terminal device 102 can use a recruitment software. For example, the terminal device 102 can install a client, which can be a client (e.g., an application (app)) specially designed to perform a specific function, or a client embedded with multiple application programs (different functions), or a client logged in through a browser. A job seeker can perform operations on the terminal device 102, such as opening the terminal device 102, using the recruitment software, and viewing positions and posting resumes in the recruitment software. Alternatively, a recruiter can also perform operations on the terminal device, such as posting recruitment information, viewing resumes of job seekers, and communicating with job seekers in the recruitment software. The recruitment information can include a position, a job location, and a function of the position, wherein the grid label corresponding to the position can include the job location and the function.

[0035] The server 104 can provide storage and retrieval functions, such as storing resumes of job seekers and recruitment information of recruiters. After the terminal device 102 receives an instruction input by a job seeker or a recruiter, the terminal device 102 sends request information including the instruction to the server 104. After the server 104 receives the request information, the server 104 performs corresponding processing, and then returns processing result information to the terminal device 102. Through a series of data processing and information interaction, the user instruction is completed. Based on the instruction, the server 104 can obtain job-seeking behavior information and convenience of the recruitment software, and also facilitate the recruiter to find suitable employees.

[0036] Based on this, the server 104 can periodically obtain job-seeking behavior information, recruitment behavior information, and recruitment information, and analyze the correlation between the information, so as to classify and merge the grid labels accurately, and improve the flexibility and accuracy of the grid label classification process.

[0037] It should be noted that the acquisition, storage, use, processing, and the like of data in the embodiments of the present application comply with relevant provisions of national laws and regulations.

[0038] The classification method of the grid label provided by the embodiments of the present application will be described below with reference to the accompanying drawings, Figure 2 A flowchart of the classification method of the grid label of the embodiments of the present application is shown, which can include the following steps 201 to 205.

[0039] In step 201, a first heterogeneous graph corresponding to the job-seeking behavior information in a preset time period and a second heterogeneous graph corresponding to the recruitment behavior information in the preset time period are obtained, and vector information of each node in the first heterogeneous graph and vector information of each node in the second heterogeneous graph are determined, wherein the first heterogeneous graph includes first relationship information between job seeker nodes and position nodes, and the second heterogeneous graph includes second relationship information between the job seeker nodes and the position nodes.

[0040] In step 202, recruitment information in the preset time period is obtained, and grid labels associated with each position node included in the first heterogeneous graph and the second heterogeneous graph are determined according to the recruitment information; and grid labels associated with each job seeker node included in the first heterogeneous graph and the second heterogeneous graph are determined according to the job-seeking behavior information and the recruitment information.

[0041] In step 203, for each grid label, N target position nodes and M target job seeker nodes are selected from the nodes included in the first heterogeneous graph and the second heterogeneous graph, respectively, wherein the N target position nodes include a position node associated with the grid label, and the M target job seeker nodes include a job seeker node associated with the grid label.

[0042] In step 204, an aggregated vector of each grid label is determined according to the vector information of the N target position nodes and the vector information of the M target job seeker nodes corresponding to the grid label.

[0043] In step 205, a plurality of aggregated vectors are clustered, and a target number of target cluster groups are determined, wherein the target number is determined according to classification fineness of business application requirements, and different target cluster groups correspond to different grid categories.

[0044] The above steps will be described in detail in combination with specific embodiments as follows.

[0045] In the embodiments of the present application, the execution subject of the grid label classification method can be an electronic device with data processing function, for example, a server. Optionally, the server can be a physical server, a virtual server, or a cloud server, and the form of the server is not limited here.

[0046] In relation to step 201, the job-seeking behavior information and the recruitment behavior information are obtained, and the grid labels are classified by analyzing the job-seeking behavior information and the recruitment behavior information. Optionally, the preset time period can be a time interval for executing the grid label classification method each time, and the preset time period can be a time length set according to application requirements. The preset time period can also be a time period set according to actual application requirements, and the preset time period is not limited here.

[0047] For example, the job-seeking behavior information and the recruitment behavior information can be obtained from a service device of a recruitment software. For example, the job-seeking behavior information can include the following information: a job seeker views a position in the recruitment software, submits a resume, and the like. The recruitment behavior information can include the following information: a recruiter views a resume for a position, communicates with a job seeker, and the like.

[0048] The heterogeneous graph includes a plurality of nodes and a plurality of edges. The nodes having an association relationship are connected by the edges.

[0049] In one example, a job seeker submits a resume to a desired position in a job-seeking process. Thus, the job-seeking behavior information can be used to find the job seeker nodes and the positions having a job-seeking relationship, which is first relationship information. Then, the connection between different job seeker nodes and position nodes is established according to the job-seeking behavior information, and a first heterogeneous graph is established. Each resume can correspond to a job seeker node, and each position can correspond to a position node.

[0050] In another example, a recruiter views one or more resumes of job seekers when searching for suitable candidates for a position. The recruiter can also actively communicate with the job seekers, which includes, but is not limited to, asking for basic information of the job seekers, answering questions of the job seekers, sending an interview invitation, sending a job offer, and the like. Thus, the recruitment behavior information can be used to find the nodes and the positions having a recruitment relationship, which is second relationship information. Then, the connection between different position nodes and job seeker nodes is established according to the recruitment behavior, and a second heterogeneous graph is established.

[0051] The first heterogeneous graph can be used to conveniently represent the relationship between the job seekers and the positions from the perspective of the job seekers, and the second heterogeneous graph can be used to conveniently represent the relationship between the positions and the job seekers from the perspective of the recruiters. This provides a data basis for improving the accuracy of the generated grid tag classification processing.

[0052] In addition, in the embodiments of the present application, the vector information of each node in the first heterogeneous graph and the second heterogeneous graph can be determined, that is, the vector information of each job seeker node and the vector information of each position node in each heterogeneous graph can be determined. Optionally, a pre-trained vector generation model can be used to output the vector information of each node in each heterogeneous graph.

[0053] Next, referring to step 202, the grid tags associated with each position node in each heterogeneous graph and the grid tags associated with each job seeker node can be determined as grid tags to be classified. For example, the published recruitment information includes the position, the working area of the position, the function, and the like. The working area and the function of the position are extracted from the recruitment information, so that the grid tag corresponding to the position node can be determined.

[0054] Based on this, by obtaining the recruitment information, the grid tags associated with each position node in each heterogeneous graph can be determined conveniently and accurately.

[0055] In the existing job-seeking behavior information, the positions interacted by the job seekers are included. Therefore, in combination with the recruitment information, the grid tags corresponding to the job seeker nodes in each heterogeneous graph can be determined conveniently. For example, the positions interacted by each job seeker are determined, and the grid tags corresponding to the interacted positions are determined as the grid tags associated with the job seeker nodes corresponding to the job seekers.

[0056] Next, referring to step 203, for each grid tag, N target position nodes and M target job seeker nodes can be obtained from the nodes included in the first heterogeneous graph and the second heterogeneous graph.

[0057] For example, for each grid tag, the position nodes associated with the grid tag can be determined from the position nodes included in the first heterogeneous graph and the second heterogeneous graph, and all the position nodes associated with the grid tag are taken as a selection range, so that N position nodes are obtained as target position nodes from the selection range. N is a predetermined positive integer, for example, can be 200, 300, and the like. The value of N can be set according to actual application requirements, and the value of N is not specifically limited herein.

[0058] In another example, for each grid tag, all the job seeker nodes associated with the grid tag are taken as another selection range from the job seeker nodes included in the first heterogeneous graph and the second heterogeneous graph, so that M job seeker nodes are obtained as target job seeker nodes from the selection range. M is a predetermined positive integer, for example, can be 200, 300, and the like. The value of M can be set according to actual application requirements, and the value of M is not specifically limited herein.

[0059] Optionally, when the target position nodes and the target job seeker nodes are selected, a random selection manner can be used, or a manner of preferentially selecting nodes with higher degrees can be used, and the selection manner is not specifically limited herein.

[0060] After the N target position nodes and the M target job seeker nodes are selected, next, step 204 is involved, and the aggregated vector of each grid label can be determined according to the vector information of the N target position nodes and the vector information of the M target job seeker nodes corresponding to each grid label.

[0061] Illustratively, the aggregated vector refers to fusing a plurality of vectors into one vector. Alternatively, the aggregation manner can be selected according to specific data distribution, and the aggregation manner includes but is not limited to calculating an average value, calculating a weighted average value, calculating a maximum value or a minimum value, calculating a geometric mean, etc. Through the aggregated vector, the complexity of subsequent clustering processing of the grid label can be reduced.

[0062] Taking the calculation of the average value as an example, for each grid label, the average value of the vector information of the N target position nodes and the vector information of the M target job seeker nodes can be calculated, thereby obtaining the aggregated vector of the grid label.

[0063] After obtaining the aggregated vector of each grid label, that is, after obtaining the aggregated vectors corresponding to the plurality of grid labels to be classified, next, step 205 is involved, and the clustering processing can be performed on the plurality of aggregated vectors to generate a target number of target clustering clusters.

[0064] Illustratively, each target clustering cluster includes a plurality of aggregated vectors, each aggregated vector corresponds to a grid label in a one-to-one manner, and therefore, the aggregated vectors included in each target clustering cluster can correspond to a category, and the grid labels corresponding to the plurality of aggregated vectors included in the target clustering cluster all belong to the category, thereby facilitating accurate division of the category of the grid label.

[0065] Based on this, different target clustering clusters can be obtained, which is beneficial to accurately classify and merge the grid labels and can improve the distinguishability between grid labels of different categories.

[0066] Alternatively, the target number is determined according to the classification fineness of the business application requirement. For example, the more the number of positions that need to be covered in a category grid, the smaller the classification fineness, and accordingly, the value of the target number can be smaller. The fewer the number of positions that need to be covered in a category grid, the greater the classification fineness, and accordingly, the value of the target number can be greater.

[0067] According to the embodiments of the present application, the number of positions that need to be covered in each category of grid can be determined based on the accurately estimated required traffic, the target number of clustering clusters can be flexibly determined, and multiple aggregation vectors can be automatically processed into the target number of clustering clusters with high distinguishability. The clustering clusters of different categories not only improve the efficiency of classifying and processing the grid, but also can more accurately classify and merge the grid, which is beneficial to accurately estimate the traffic in the divided region. Based on this, in subsequent position promotion business, it is beneficial to accurately estimate the promotion effect of position promotion, and in the value estimation of position promotion business, it is beneficial to more accurately calculate the promotion value.

[0068] In some embodiments of the present application, the job-seeking behavior information includes viewing position behavior and resume submission behavior. Based on this, obtaining the first heterogeneous graph corresponding to the job-seeking behavior information in the preset time period includes: determining the first relationship information between the job seeker nodes and the position nodes according to the viewing position behavior and the resume submission behavior; and generating the first heterogeneous graph according to the first relationship information between the job seeker nodes and the position nodes.

[0069] For example, according to the viewing position behavior and the resume submission behavior, it can be determined which resumes interact with which positions, which not only reflects the delivery history of the job seeker, but also reflects the interest preferences of the job seeker when selecting positions. For example, the job seeker's interest in the position to which he submits his resume is higher than his interest in the position he only views.

[0070] As a specific example, Figure 3 is a schematic diagram of a first heterogeneous graph provided by an embodiment of the present application. The viewing position behavior and the resume submission behavior of the job seeker associate the job seeker with the position (J), and the first heterogeneous graph of the job seeker and the position is constructed with the job seeker (U) and the position (J) as nodes and the viewing and the submitting as edges. Continue to combine Figure 3 As shown, U1, U2, U3, U4, and U5 represent different job seeker nodes, and J1, J2, J3, and J4 represent different positions.

[0071] In the embodiments of the present application, the advantage of constructing a heterogeneous graph is that after the connection between the nodes with an association relationship is established, the relationship between the job seeker and the position can be conveniently expressed, so as to facilitate the extraction of relevant feature information of the grid label from the grid label, and then accurately classify and merge the grid label.

[0072] In some embodiments of the present application, the recruitment behavior information includes resume viewing behavior and recruitment communication behavior. Based on this, the second heterogeneous graph corresponding to the recruitment behavior information in a preset time period is obtained, including: determining the second relationship information between the job seeker nodes and the position nodes according to the resume viewing behavior and the recruitment communication behavior; and generating the second heterogeneous graph according to the second relationship information between the job seeker nodes and the position nodes.

[0073] For example, according to the resume viewing behavior and the recruitment communication behavior, it can be determined which resumes the recruiter viewed for the position, and which job seekers interacted with. This not only reflects the recruitment history of the recruiter, but also reflects the interest preferences of the recruiter when selecting job seekers. For example, the recruiter's interest in the job seekers he communicated with is higher than the interest in the job seekers whose resumes he only viewed. As a specific example, Figure 4 is a schematic diagram of a second heterogeneous graph provided by an embodiment of the present application. The job seeker's viewing position behavior and resume submission behavior associate the job seeker with the position (J). The first heterogeneous graph of the job seeker and the position is constructed with the position (J) and the job seeker (U) as nodes, and the viewing and the submission as edges. The second heterogeneous graph is continued to be constructed in combination with Figure 4 As shown, U1, U2, U3, and U4 represent different job seeker nodes, and J1, J2, J3, J4, and J5 represent different positions.

[0074] In an embodiment of the present application, the advantage of constructing a heterogeneous graph is that after connecting nodes with an association relationship, the relationship between the job seeker and the position can be conveniently expressed, so as to facilitate the extraction of relevant feature information of the grid label from the heterogeneous graph, and then accurately classify and merge the grid label.

[0075] In some embodiments of the present application, the steps of determining the vector information of each node in the first heterogeneous graph and the vector information of each node in the second heterogeneous graph can specifically refer to the following steps 301 to 303.

[0076] Step 301: generating a training sample set according to the first heterogeneous graph and the second heterogeneous graph.

[0077] Step 302: training a pre-constructed vector generation model according to the training sample set, and obtaining the target vector generation model after a preset training condition is reached.

[0078] Step 303: determining the vector information of each node in the first heterogeneous graph and the vector information of each node in the second heterogeneous graph through the target vector generation model.

[0079] In the embodiments of the present application, the training samples in the training sample set are determined according to the walk sequences obtained by performing walks on the first and second heterogeneous graphs according to different meta paths.

[0080] For example, a meta path refers to a specific path pattern between different types of nodes and edges in a heterogeneous graph. For example, a path from a job seeker node to a position node and then to a job seeker node can be a meta path. By setting different meta paths, multiple random walks can be performed in the heterogeneous graph, thereby generating walk sequences reflecting the complex relationships between nodes. Each walk sequence actually represents the node relationship under a certain association mode and can be used as training data for subsequent models. For example, the pre-constructed meta paths for the first heterogeneous graph are, for example: (1) job seeker→position→job seeker; (2) position→job seeker→position; and the pre-constructed meta paths for the second heterogeneous graph are, for example: (1) job seeker→position→job seeker; (2) position→job seeker→position. For each meta path, multiple walk sequences are found through walks.

[0081] In some embodiments, the pre-constructed vector generation model includes but is not limited to models constructed using various model frameworks such as Deep Graph Library (DGL), PyG, NeuGraph, etc. Optionally, in the embodiments of the present application, the pre-constructed vector generation model is based on the DGL model framework and performs graph embedding on the relationships between the nodes in the heterogeneous graph.

[0082] For example, the pre-constructed vector generation model can be a model developed based on a Skip-gram model, where the Skip-gram model is a prediction model in natural language processing and can be used to learn continuous vector representations of words.

[0083] The preset training condition can be that the function value of the loss function tends to a preset numerical value, or the training period (Epoch) of the model reaches a preset number of times, etc. The specific preset training condition can be set according to actual training requirements.

[0084] After the model training is completed, the target vector generation model is obtained. The target vector generation model can output the vector information of each node in the first heterogeneous graph and the vector information of each node in the second heterogeneous graph, i.e., the vector information of the job seeker nodes and the vector information of the position nodes in the first heterogeneous graph, and the vector information of the job seeker nodes and the vector information of the position nodes in the second heterogeneous graph.

[0085] According to the embodiment of the present application, the training sample is generated by using the heterogeneous graph generated based on the job-seeking behavior information and the recruitment information, the vector generation model is trained, the potential relationship between the job-seeking intention of the job seeker and the grid label is deep mined, and then the accuracy of the classification processing of the grid label corresponding to the position is improved.

[0086] In addition, by setting different meta-paths and combining with flexible adjustment of the walking strategy, the required training sample is generated, and the vector generation model is trained, so that the feature information can be conveniently and efficiently mined from the job-seeking behavior information and the recruitment information, and converted into reliable vector representation to support the subsequent classification processing of the grid label.

[0087] In some embodiments of the present application, according to the first heterogeneous graph and the second heterogeneous graph, the training sample set is generated, including: according to the first meta-path, walking the first heterogeneous graph, and according to the second meta-path, walking the second heterogeneous graph to obtain a walking sequence set, the walking sequence set including a plurality of walking sequences corresponding to each meta-path; generating the training sample set according to the order of the nodes in each walking sequence in the walking sequence set and the K nodes adjacent to the nodes, K being a positive integer.

[0088] For example, the meta-paths set in advance for the first heterogeneous graph are described as first meta-paths, for example, meta-path (1) and meta-path (2), and the meta-paths set in advance for the second heterogeneous graph are described as second meta-paths, for example, meta-path (3) and meta-path (4).

[0089] Figure 5 is a schematic diagram of walking the heterogeneous graph provided by the embodiment of the present application, which is shown in combination with Figure 5 (a). As shown in combination with Figure 5 (a), based on the meta-path (1) job seeker→position→job seeker, walking is performed, and one walking in the first heterogeneous graph can obtain a walking sequence: U1-J1-U3. As shown in combination with Figure 5 (b), based on the meta-path (3) position→job seeker→position, walking is performed, and one walking in the second heterogeneous graph can obtain a walking sequence: J1-U1-J3. Corresponding to each first meta-path, a plurality of walking sequences can be obtained after walking the first heterogeneous graph, and each walking sequence corresponding to each first meta-path is not listed one by one.

[0090] In some embodiments, the connection line between the nodes in the first heterogeneous graph and the second heterogeneous graph is an undirected edge. Based on this, the walking direction can be randomly selected at any node.

[0091] Optionally, in the process of walking each heterogeneous graph, the walking parameters can be set in advance, for example, the walking depth, the number of times each node in the walking process is used as the starting point of walking. Taking the first heterogeneous graph and the meta-path (1) as an example, if the walking depth is 1, the obtained walking sequence can be U1-J1-U3, and if the walking depth is 2, after obtaining U1-J1-U3, U3 is used as the starting point, and the walking in the first heterogeneous graph is continued according to the meta-path (1), and the final walking sequence obtained is, for example, U1-J1-U3-J4-U4. In the embodiments of the present application, optionally, the walking depth and the number of times each node in the walking process is used as the starting point of walking can be set according to application requirements, which is not limited here.

[0092] Optionally, the weight between the job seeker node and the position node can also be set in the embodiments of the present application, in combination with Figure 6 As shown in FIG. 6, V represents the viewing behavior of a job seeker, A represents the delivery behavior of a job seeker, and the number after the colon represents the weight of the edge. In combination with Figure 7 As shown in FIG. 7, H represents the viewing behavior of a recruiter, and C represents the delivery behavior of a recruiter. Among them, the greater the weight, the greater the probability of being selected in the walking process. The advantage of this is that it is beneficial to fully mine the intention information of the job seeker behavior and the intention information of the recruiter behavior, and to convert each node into a vector representation through a vector generation model, to support the subsequent step of classifying the grid label.

[0093] In each walking sequence, assuming that a node represents a job seeker node, the K neighbor nodes of the node can include the position nodes that have interacted with it. By taking the order relationship of these nodes and their neighbor nodes as training samples, the association between the resume and the position can be further analyzed, and training data can be provided for the subsequent vector generation model. The value of K can be adjusted according to actual requirements. Generally, the larger K is, the more neighbor nodes of the node can participate in training, thereby generating more abundant training samples.

[0094] According to the embodiments of the present application, the random walk based on the meta-path can capture the multi-dimensional association between different types of nodes in the heterogeneous graph, so that the training sample set is not limited to the specific relationship of a certain type of node, but can reflect the information of multiple dimensions such as resumes and positions. Based on this, it is beneficial to mine useful information from the job behavior information and the recruitment information to support the subsequent step of classifying the grid label.

[0095] In some embodiments of the present application, optionally, the training sample set includes a plurality of positive training samples and a plurality of negative training samples. Optionally, according to the order of nodes in each walk sequence and the K nodes of each node neighborhood, the training sample set can be generated, and the positive training samples and the negative training samples can be obtained by referring to the following steps. Specifically, for each walk sequence, the combination of the current node and each node in the K nodes of the current node neighborhood is obtained as a positive training sample, and the current node is any node in the walk sequence. In each positive training sample, the current node is the input sample node, and the other node is the output sample node.

[0096] For example, the walk sequence set includes a plurality of walk sequences corresponding to each first meta-path and a plurality of walk sequences corresponding to each second meta-path. For each walk sequence, the combination of the current node and each node in the K nodes of the current node neighborhood is obtained as a positive training sample.

[0097] For the input sample node in each positive training sample, L nodes are randomly obtained as output negative training sample nodes in the plurality of walk sequences included in the walk sequence set, and the L output negative training sample nodes are combined with the input sample node respectively to obtain L negative training samples corresponding to each positive training sample.

[0098] As a specific example, for example, in the walk sequence U1-J1-U3-J4-U4, U1, J1, U3, J4 and U4 will be respectively combined with each node in the K nodes of their neighborhood. For example, K=2, U1 is the center point, and the neighborhood nodes of U1 are J1 and U3. Thus, two positive training samples U1-J1 and U1-U3 can be obtained. Among them, R2 is the input sample node in the positive training sample.

[0099] J1 is the center point, and the neighborhood nodes of J1 are U1, U3 and J4. Thus, three positive training samples J1-U1, J1-U3 and J1-J4 can be obtained. Among them, J1 is the input sample node in the positive training sample.

[0100] U3 is the center point, and the neighborhood nodes of U3 are U1, J1, J4 and U4. Thus, four positive training samples U3-U1, U3-J1, U3-J4 and U3-U4 can be obtained. Among them, U3 is the input sample node in the positive training sample.

[0101] Based on the same positive training sample generation logic, three positive training samples can be obtained when J4 is the center point, and two positive training samples can be obtained when U4 is the center point, which are not listed here.

[0102] Generally, the number of nodes in the heterogeneous graph is large, in order to speed up the model training, the negative training samples can be generated according to the negative sampling manner, taking the meta path (1) as an example, U1~U5 are taken as the starting points respectively, according to the preset depth, the walk path is obtained, a total of 5 paths are obtained, among all the nodes included in the 5 paths, the nodes are selected to generate the negative training samples. For example, for the positive training sample U1-J1, 5 nodes are randomly selected from all U and J as negative training samples according to the probability: U1-F5, U1-U2, U1-U3, U1-J5 and U1-J4. It can be seen that in the process of selecting negative training samples, the frequency of the node as the probability of being selected as a negative training sample, that is, the higher the frequency of the node, the greater the probability of being selected as a negative training sample. The advantage of this setting is to strengthen the training of high-frequency nodes and reduce the risk of model bias caused by high-frequency nodes.

[0103] It can be understood that the number of negative training samples corresponding to each positive training sample is determined according to actual application requirements and performance settings, which is not specifically limited here.

[0104] According to the embodiments of the present application, the vectors of each node are generated based on the heterogeneous graph and the target vector generation model, which can effectively mine useful information from the job-seeking behavior information and the recruitment information and convert it into vector representation to support the subsequent step of classifying the grid labels.

[0105] In some embodiments of the present application, optionally, before selecting N target position nodes and M target job seeker nodes for each grid label, the grid label and the nodes can be subjected to screening processing. For example, the nodes with small degree in each heterogeneous graph are screened out, and the nodes included in the first heterogeneous graph and the second heterogeneous graph after the node screening processing are determined as the grid labels to be classified.

[0106] For example, the degree of a node refers to the number of edges directly connected to the node. In the embodiments of the present application, the edges in the first heterogeneous graph and the second heterogeneous graph are all undirected edges, and the degree of a node refers to the total number of edges connected to the node. For example, if a node is connected to three other nodes, then its degree is 3.

[0107] For example, the preset screening condition can be that nodes with a node degree less than and equal to 9 are screened out in each heterogeneous graph. Thus, the processing for the first heterogeneous graph is specifically to retain only nodes with a node degree greater than 10 and screen out nodes with a node degree less than and equal to 9. After the screening processing is completed, the grid tags associated with the remaining candidate nodes and the grid tags associated with the remaining position nodes in the first heterogeneous graph are determined as grid tags to be classified, and then target candidate nodes or target position nodes are selected for the grid tags. Based on the same processing manner, the grid tags corresponding to the remaining candidate nodes and the grid tags corresponding to the remaining position nodes in the second heterogeneous graph can be determined as grid tags to be classified, and then target candidate nodes or target position nodes are selected for the grid tags, which will not be described herein again.

[0108] Since the greater the node degree is, the more the association relationship between the node and other nodes (for example, candidate nodes or position nodes) is, thus, after the nodes with a small node degree are screened out, the relationship between the nodes in each heterogeneous graph is more referential, which is beneficial to mining more useful information therefrom. Therefore, based on the target candidate nodes and the target position nodes with a large node degree as the basis for classification of the grid tags, the accuracy of subsequent classification processing of the grid tags can be improved.

[0109] In the embodiments of the present application, for each grid tag, N target position nodes and M target candidate nodes are selected from the nodes included in the first heterogeneous graph and the second heterogeneous graph, which can specifically include the following steps: for each grid tag, a first position node associated with the grid tag and a first candidate node associated with the grid tag in the first heterogeneous graph are obtained to obtain a first node set; and a second position node associated with the grid tag and a second candidate node associated with the grid tag in the second heterogeneous graph are obtained to obtain a second node set; for each grid tag, N target position nodes are selected from the position nodes included in the first node set and the second node set, and M target candidate nodes are selected from the candidate nodes included in the first node set and the second node set.

[0110] For example, in the first heterogeneous graph and the second heterogeneous graph, each candidate node and each position node is associated with a grid tag, that is, each node is associated with a grid tag, wherein the grid tags associated with different nodes can be different or the same. Thus, for each grid tag, the candidate nodes and the position nodes associated with the grid tag can be determined conveniently, wherein the first position nodes associated with the grid tag and the first candidate nodes associated with the grid tag in the first heterogeneous graph constitute a first node set; and the second position nodes associated with the grid tag and the second candidate nodes associated with the grid tag in the second heterogeneous graph constitute a second node set.

[0111] Optionally, the number of target position nodes selected from the first node set and the second node set can be pre-configured, for example, N / 2 nodes are selected from each of the first node set and the second node set; for another example, the number of target position nodes selected from each of the first node set and the second node set can be determined according to the proportion of the number of position nodes included in each of the first node set and the second node set, for example, the more the number of position nodes, the greater the number of target position nodes that can be selected, which is not specifically limited herein. Optionally, the target candidate nodes can be selected based on the same selection logic as the target position nodes, which will not be described herein again.

[0112] In the embodiments of the present application, by selecting a certain number of target candidate nodes and target position nodes for each grid label to be classified for determining the aggregation vector of the grid label, it is helpful to reserve key information for grid classification processing and improve the clustering processing efficiency.

[0113] In some embodiments of the present application, after N target position nodes and M target candidate nodes are selected, next, the aggregation vector of each grid label can be determined according to the vector information of the N target position nodes and the vector information of the M target candidate nodes corresponding to each grid label. Compared with directly calculating using the vectors corresponding to the N target position nodes and the M target candidate nodes corresponding to the grid label respectively, using the aggregation vector to represent the grid label can reduce the calculation difficulty, and can also multi-dimensionally and more accurately summarize the features included in the grid label.

[0114] After obtaining the aggregation vector of each grid label, that is, obtaining the aggregation vectors corresponding to the plurality of grid labels to be classified respectively. Next, the plurality of aggregation vectors are clustered and a target number of target clustering clusters are determined, which can be specifically referred to the following steps: in the plurality of aggregation vectors, a target number of first aggregation vectors are randomly selected as initial clustering centers of a target number of clustering clusters respectively; according to a pre-configured clustering algorithm, each clustering cluster and the clustering center of the clustering cluster are iteratively updated until a pre-configured stop iteration condition is reached, and a target clustering cluster is obtained; wherein the second aggregation vector is an aggregation vector in the plurality of aggregation vectors except the first aggregation vector.

[0115] Optionally, the plurality of aggregation vectors can be clustered using a pre-configured clustering algorithm to generate a target number of target clustering clusters. For example, the pre-configured clustering algorithm can include but is not limited to any one of a plurality of clustering methods such as Mini Batch K-means clustering, hierarchical clustering, density clustering, etc.

[0116] Exemplarily, the clustering process is introduced taking the Mini Batch K-means clustering as an example and generating a target number of clustering clusters as an example.

[0117] The target number is determined according to the classification fineness of the business application requirement. Taking the target number Z as an example, first, Z aggregated vectors are randomly selected from the plurality of aggregated vectors, and the Z first aggregated vectors are determined as Z clustering centers respectively. At this time, the Z clustering centers are initial clustering centers.

[0118] For ease of description, the Z aggregated vectors randomly selected are described as first aggregated vectors, and the remaining aggregated vectors in the plurality of aggregated vectors are described as second aggregated vectors.

[0119] Next, in each iteration calculation, Y second aggregated vectors are randomly selected from all second aggregated vectors, and for each second aggregated vector in the Y second aggregated vectors, a clustering center associated with each second aggregated vector is determined, wherein the clustering center associated with each second aggregated vector is the clustering center closest to the second aggregated vector in the Z clustering centers. Then, the position of the clustering center associated with the second aggregated vector is updated according to the second aggregated vector, thereby completing one iteration calculation.

[0120] It can be understood that in the entire clustering process, each second aggregated vector is selected only once, and each iteration calculation selects only the second aggregated vector that has not been calculated. After the preset stopping iteration condition is reached, the Z target clustering centers are obtained.

[0121] Exemplarily, the preset stopping iteration condition can be that all second aggregated vectors participate in the iteration calculation, or the number of iteration calculations reaches a preset iteration number, or the position change of each clustering center is less than a threshold, etc.

[0122] Based on this, the grid labels can be accurately classified and merged to obtain Z different target clustering clusters with high discrimination, and different target clustering clusters correspond to different categories of grid labels, thereby flexibly coping with different classification requirements.

[0123] In the embodiments of the present application, the target number Z can be determined according to the business application requirement. Exemplarily, the target number can be determined according to the following business application requirements: (1) the proportion of the delivery of the grids in the cluster in the cluster; (2) the proportion of the existence of correlation between the grids in the cluster; (3) the balance degree of the traffic distribution between clusters.

[0124] Exemplarily, the business application requirement (1) can be calculated based on the perspective of the job seeker. Each job seeker node and each position node has an associated grid tag. Specifically, each position, when published, can include not only position information but also function information in the recruitment information. By analyzing the recruitment information, the association between the position and the grid tag can be found. For example, different positions can correspond to different grid tags or the same grid tag.

[0125] During the job-seeking process of the job seeker, the job seeker can deliver a resume to the desired position. Thus, the delivery relationship between the grid and the grid can be simplified. Based on this, the calculation method can be as follows: if the grid tag associated with the current job seeker and the grid tag associated with the position to which the job seeker delivers the resume are still within the current cluster, counting is performed. Finally, the count / all delivery numbers, i.e., the proportion of the delivery of the grid in the cluster by the position in the cluster, are obtained.

[0126] Exemplarily, the business application requirement (2) can be calculated based on the perspective of the job seeker.

[0127] The calculation formula can be as shown in formula (1).

[0128]

[0129] where C is all the grids in the current cluster, and G is a grid.

[0130] For each grid in the cluster, it is assumed that the cluster includes n grids. There are n*n combinations. Within a certain time range (such as 30 days), for each n*n combination in the cluster, there is a delivery relationship between the grid and the grid. The number of combinations that can exist / (n*n) is the proportion of the existence of the correlation between the grids in the cluster.

[0131] In yet another example, the business application requirement can also be calculated based on the perspective of the recruiter. For example, the proportion of the job seekers of the grid in the cluster to which the recruiter communicates is calculated. Each job seeker node and each position node has an associated grid tag and a grid tag. The recruiter has a desired resume for a certain position, and can actively communicate with the job seeker corresponding to the desired resume. Thus, based on the grid tag associated with the certain position and the grid tag associated with the job seeker communicated based on the position, the communication behavior of the recruiter can be simplified as the recruitment relationship between the grid and the grid. Based on this, the calculation method can be as follows: if the grid tag associated with the current position and the grid tag associated with the job seeker communicated based on the position are still within the current cluster, counting is performed. Finally, the count / all delivery numbers, i.e., the proportion of the job seekers of the grid in the cluster to which the recruiter communicates, are obtained.

[0132] In addition, the proportion of the chat correlation between the grids in the cluster can also be calculated from the perspective of the recruiter. The calculation manner is the same as the calculation logic of calculating the proportion of the correlation between the grids in the cluster from the perspective of the job seeker, which is not described herein again.

[0133] In some embodiments, when the business application requirement (3) requires that attention should be paid to the recommendation to the job seeker and the exposure of the grid should be increased for a certain grid, the grid size can be determined from the perspective of the job seeker; when the communication between the recruiter and the job seeker needs to be improved, the grid size can be determined from the perspective of the recruiter.

[0134] According to the embodiments of the present application, the number of positions that need to be covered in each category in the grid can be determined based on the accurately estimated required traffic, the target number of clustering clusters can be flexibly determined, and multiple aggregation vectors can be automatically processed into the target number of clustering clusters with high distinguishability. The clustering clusters of different categories not only improve the efficiency of classifying and processing the grid, but also more accurately classify and merge the grid, which is beneficial to accurately estimate the traffic in the divided region.

[0135] The implementation manners of the embodiments of the present application and the advantages brought by the implementation manners are described above through multiple embodiments. In order to more clearly introduce the embodiments of the present application, the specific processing process of the embodiments of the present application is described in detail below in combination with specific examples.

[0136] Figure 8 is a flowchart of another grid label classification method provided by the embodiments of the present application, and in combination with Figure 8 as shown, the scheme can include the following steps 801 to 810.

[0137] Step 801, obtain job behavior data and recruitment behavior information, and construct a first heterogeneous graph and a second heterogeneous graph.

[0138] Step 802, obtain a meta-path.

[0139] Step 803, traverse the job seeker nodes and the position nodes in the first heterogeneous graph and the second heterogeneous graph respectively, and respectively walk in the first heterogeneous graph and the second heterogeneous graph according to the preset meta-path, to generate multiple walk sequences corresponding to each meta-path.

[0140] Step 804, sample each walk sequence to generate multiple positive training samples and multiple negative training samples associated with each positive training sample.

[0141] Step 805, construct a vector generation model.

[0142] Step 806, train the vector generation model according to the positive training samples and the negative training samples.

[0143] At step 807, multiple cycles of training are performed. When the loss of the vector generation model no longer decreases, the training is stopped, and a target vector generation model is obtained.

[0144] After the training is stopped, the model parameters of the current model are saved, and a target vector generation model is obtained.

[0145] At step 808, nodes with a degree less than or equal to a preset threshold value are filtered out from the first and second heterogeneous graphs, and the nodes included in the first and second heterogeneous graphs after the node filtering are determined as grid labels to be classified.

[0146] At step 809, for each grid label, N target position nodes and M target job seeker nodes are obtained according to the nodes included in the first and second heterogeneous graphs.

[0147] At step 810, the aggregated vector of each grid label is determined according to the vector information of the N target position nodes and the vector information of the M target job seeker nodes corresponding to each grid label.

[0148] At step 811, the multiple aggregated vectors are clustered to generate a target number of target cluster groups.

[0149] Optionally, the target number is determined according to the classification fineness of business application requirements, and different target cluster groups correspond to different categories of grid labels.

[0150] For example, in the multiple aggregated vectors, a target number of first aggregated vectors are randomly selected as initial cluster centers of a target number of cluster groups; according to a preset clustering algorithm, each cluster group and the cluster center of the cluster group are iteratively updated until a preset stopping iteration condition is reached, and a target cluster group is obtained; wherein the second aggregated vector is an aggregated vector other than the first aggregated vector in the multiple aggregated vectors.

[0151] According to the embodiments of the present application, the number of positions that need to be covered in each category within the grid can be determined based on the required traffic that needs to be accurately estimated, the target number of cluster groups can be flexibly determined, and the multiple aggregated vectors can be automatically processed into a target number of cluster groups with high distinguishability and different categories. This not only improves the efficiency of classifying the grid, but also more accurately classifies and merges the grid, which is beneficial to accurately estimating the traffic in the divided area.

[0152] Corresponding to the method embodiments of the present application, the present application also provides a grid label classification device, as shown in Figure 9 The grid label classification device 900 includes an acquisition module 901 and a processing module 902.

[0153] The acquisition module 901 is configured to acquire a first heterogeneous graph corresponding to job-seeking behavior information in a preset time period and a second heterogeneous graph corresponding to recruitment behavior information in the preset time period, and determine vector information of each node in the first heterogeneous graph and vector information of each node in the second heterogeneous graph. The first heterogeneous graph includes first relationship information between a job seeker node and a position node, and the second heterogeneous graph includes second relationship information between the job seeker node and the position node.

[0154] The acquisition module 901 is further configured to acquire recruitment information in the preset time period, determine, according to the recruitment information, a grid label associated with each position node included in the first heterogeneous graph and the second heterogeneous graph, and determine, according to the job-seeking behavior information and the recruitment information, a grid label associated with each job seeker node included in the first heterogeneous graph and the second heterogeneous graph.

[0155] The processing module 902 is configured to, for each grid label, select N target position nodes and M target job seeker nodes from nodes included in the first heterogeneous graph and the second heterogeneous graph, wherein the N target position nodes include a position node associated with the grid label, and the M target job seeker nodes include a job seeker node associated with the grid label.

[0156] The processing module 902 is further configured to determine an aggregated vector of each grid label according to vector information of the N target position nodes and vector information of the M target job seeker nodes corresponding to the grid label.

[0157] The processing module 902 is further configured to perform clustering processing on the plurality of aggregated vectors, and determine a target number of target clustering clusters, wherein the target number is determined according to classification fineness of a business application requirement, and different target clustering clusters correspond to different grid categories.

[0158] In some embodiments of the present application, optionally, the job-seeking behavior information includes a position viewing behavior and a resume submission behavior.

[0159] The processing module 902 is further configured to determine the first relationship information between the job seeker node and the position node according to the position viewing behavior and the resume submission behavior.

[0160] The processing module 902 is further configured to generate the first heterogeneous graph according to the first relationship information between the job seeker node and the position node.

[0161] In some embodiments of the present application, optionally, the recruitment behavior information includes a resume viewing behavior and a recruitment communication behavior.

[0162] The processing module 902 is further configured to determine the second relationship information between the job seeker node and the position node according to the resume viewing behavior and the recruitment communication behavior.

[0163] The processing module 902 is further configured to generate a second heterogeneous graph according to the second relationship information between the job seeker nodes and the position nodes.

[0164] In some embodiments of the present application, the processing module 902 is further configured to generate a training sample set according to the first heterogeneous graph and the second heterogeneous graph.

[0165] The processing module 902 is further configured to train a pre-constructed vector generation model according to the training sample set, and obtain a target vector generation model after a preset training condition is reached.

[0166] The processing module 902 is further configured to determine the vector information of each node in the first heterogeneous graph and the vector information of each node in the second heterogeneous graph by using the target vector generation model.

[0167] In some embodiments of the present application, the processing module 902 is further configured to walk the first heterogeneous graph according to the first meta-path, and walk the second heterogeneous graph according to the second meta-path to obtain a walking sequence set, the walking sequence set including a plurality of walking sequences corresponding to each meta-path.

[0168] The processing module 902 is further configured to generate a training sample set according to the order of the nodes in each walking sequence in the walking sequence set and the K nodes adjacent to the nodes, K being a positive integer.

[0169] In some embodiments of the present application, the acquisition module 901 is further configured to, for each grid label, acquire a first position node associated with the grid label and a first job seeker node associated with the grid label in the first heterogeneous graph to obtain a first node set, and acquire a second position node associated with the grid label and a second job seeker node associated with the grid label in the second heterogeneous graph to obtain a second node set.

[0170] The processing module 902 is further configured to, for each grid label, select N target position nodes from the position nodes included in the first node set and the second node set, and select M target job seeker nodes from the job seeker nodes included in the first node set and the second node set.

[0171] In some embodiments of the present application, the processing module 902 is further configured to randomly select a target number of first aggregated vectors as initial cluster centers of a target number of cluster clusters from the plurality of aggregated vectors.

[0172] The processing module 902 is further configured to iteratively update each cluster cluster and the cluster center of the cluster cluster according to a preset clustering algorithm until a preset stop iteration condition is reached to obtain a target cluster cluster.

[0173] The second aggregated vector is an aggregated vector other than the first aggregated vector in the plurality of aggregated vectors.

[0174] It can be understood that the grid tag classification device of the embodiments of the present application can correspond to the execution subject of the grid tag classification method provided by the embodiments of the present application. The specific details of the operation and / or function of each module / unit of the grid tag classification device can be referred to the description of the corresponding part in the grid tag classification method provided by the embodiments of the present application. For brevity, it will not be repeated here.

[0175] The electronic device in the embodiments of the present application can be a user terminal device, can be a server, can also be other computing devices, and can also be a cloud server. Figure 10 A hardware structure schematic diagram of the electronic device of the embodiments of the present application is shown. The electronic device can include a processor 1001 and a memory 1002 storing computer program instructions. The processor 1001 executes the computer program instructions to implement the flow or function of the method of any of the above embodiments.

[0176] Specifically, the processor 1001 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or can be configured to implement one or more integrated circuits of the embodiments of the present application. The memory 1002 can include a mass storage for data or instructions. For example, the memory 1002 can be at least one of a hard disk drive (HDD), a read-only memory (ROM), a random access memory (RAM), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, a universal serial bus (USB) drive, or other physical / tangible memory storage devices. For another example, the memory 1002 can include removable or non-removable (or fixed) media. For another example, the memory 1002 can be inside or outside the integrated gateway disaster recovery device. The memory 1002 can be a non-volatile solid-state memory. In other words, the memory 1002 generally includes a tangible (non-transitory) computer-readable storage medium (such as a memory device) encoded with computer-executable instructions and, when the software is executed (such as by one or more processors), can perform the operations described in the method of the embodiments of the present application. The processor 1001 implements the flow or function of any of the above methods by reading and executing the computer program instructions stored in the memory 1002.

[0177] In one example, Figure 10The electronic device shown can also include a communication interface 1003 and a bus 1010. Among them, the processor 1001, the memory 1002, the communication interface 1003 are connected through the bus 1010 and complete the communication between each other. The communication interface 1003 is mainly used to realize the communication between the modules, devices, units and / or equipment in the embodiments of the application. The bus 1010 includes hardware, software or both, which can couple the components of the online data flow billing device to each other. For example, the bus can include at least one of the following: an accelerated graphics port (AGP) or other graphics bus, an enhanced industry standard architecture (EISA) bus, a front side bus (FSB), a hyper transport (HT) interconnect, an industry standard architecture (ISA) bus, an infiniband interconnect, a low pin count (LPC) bus, a memory bus, a micro channel architecture (MCA) bus, a peripheral component interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a serial advanced technology attachment (SATA) bus, a video electronics standards association local (VLB) bus or other suitable bus. The bus 1010 can include one or more buses. Although the embodiments of the application describe or show a specific bus, any suitable bus or interconnection method can be considered by the embodiments of the application.

[0178] In combination with the method in the above embodiments, the embodiments of the present application further provide a computer readable storage medium, which has stored thereon computer program instructions, and the computer program instructions are executed by a processor to implement the flow or function of any of the methods in the above embodiments.

[0179] In addition, the embodiments of the present application also provide a computer program product, which has stored thereon computer program instructions, and the computer program instructions are executed by a processor to implement the flow or function of any of the methods in the above embodiments.

[0180] The flowcharts and / or block diagrams of the methods, devices, systems and computer program products of the embodiments of the present application are described above as examples, and the related aspects are described. It should be understood that each block in the flowchart and / or block diagram can be implemented by computer program instructions, or by special hardware that performs specified functions or actions, or by a combination of special hardware and computer instructions. For example, these computer program instructions can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, to form a machine, so that the instructions executed by the processor enable the implementation of the functions / actions specified in each block or combination of blocks in the flowchart and / or block diagram. Such a processor can be a general purpose processor, a special purpose processor, a special application processor, or a field programmable logic circuit.

[0181] The functional blocks shown in the structural block diagram of the embodiments of the present application can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and the like; when implemented in software, it is a program or a code segment used to perform the required tasks. The program or code segment can be stored in a memory or transmitted through a data signal carried in a carrier wave over a transmission medium or a communication link. The code segment can be downloaded via a computer network, such as the Internet, an intranet, and the like.

[0182] It should be noted that the present application is not limited to the specific configurations and processes described above or shown in the drawings. The above is merely a specific implementation of the present application, and those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the described systems, devices, modules or units can refer to the corresponding processes in the method embodiments, which need not be described again. It should be understood that the protection scope of the present application is not limited thereto, and any skilled person in the art can think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be covered within the protection scope of the present application.

Claims

1. A grid label classification method, characterized in that: include: Obtaining a first heterogeneous graph corresponding to the job-seeking behavior information within a preset time period and a second heterogeneous graph corresponding to the recruitment behavior information within the preset time period, and determining vector information of each node in the first heterogeneous graph and vector information of each node in the second heterogeneous graph, wherein the first heterogeneous graph includes first relationship information between job seeker nodes and position nodes, and the second heterogeneous graph includes second relationship information between job seeker nodes and position nodes; Obtaining recruitment information within the preset time period, and determining, based on the recruitment information, a grid label associated with each position node included in the first heterogeneous graph and the second heterogeneous graph; and determining, based on the job-seeking behavior information and the recruitment information, a grid label associated with each job seeker node included in the first heterogeneous graph and the second heterogeneous graph; For each of the grid labels, selecting N target position nodes and M target job seeker nodes from the nodes respectively included in the first heterogeneous graph and the second heterogeneous graph, wherein the N target position nodes include the position node associated with the grid label, and the M target job seeker nodes include the job seeker node associated with the grid label; Determine an aggregate vector for each grid label according to vector information of the N target job nodes and vector information of the M target job seeker nodes corresponding to each grid label; Clustering is performed on the plurality of aggregation vectors, and a target number of target clusters is determined, wherein the target number is determined according to the classification fineness required by the business application, and different target clusters correspond to different grid categories.

2. The method according to claim 1, characterized in that The job-seeking behavior information includes job search behavior and resume submission behavior; The obtaining of the first heterogeneous graph corresponding to the job-seeking behavior information within a preset time period includes: Determining first relationship information between the job seeker node and the job position node according to the job viewing behavior and the resume submission behavior; The first heterogeneous graph is generated according to the first relationship information between the job seeker nodes and the position nodes.

3. The method according to claim 1, characterized in that The recruitment behavior information includes resume viewing behavior and recruitment communication behavior; Obtaining the second heterogeneous graph corresponding to the recruitment behavior information within the preset time period includes: Determining second relationship information between the job seeker node and the position node based on the resume viewing behavior and the recruitment communication behavior; The second heterogeneous graph is generated according to the second relationship information between the job seeker nodes and the position nodes.

4. The method according to claim 1, wherein The determining of the vector information of each node in the first heterogeneous graph and the vector information of each node in the second heterogeneous graph includes: Generating a training sample set according to the first heterogeneous graph and the second heterogeneous graph; Training the pre-built vector generation model according to the training sample set, and obtaining a target vector generation model after meeting preset training conditions; The target vector generation model is used to determine vector information of each node in the first heterogeneous graph and vector information of each node in the second heterogeneous graph.

5. The method according to claim 4, characterized in that Generating a training sample set according to the first heterogeneous graph and the second heterogeneous graph includes: Walk the first heterogeneous graph according to the first meta-path, and walk the second heterogeneous graph according to the second meta-path to obtain a walk sequence set, wherein the walk sequence set includes multiple walk sequences corresponding to each meta-path; The training sample set is generated according to the order of nodes in each walking sequence in the walking sequence set and K nodes adjacent to the node, where K is a positive integer.

6. The method according to claim 1, characterized in that For each of the grid labels, N target position nodes and M target job seeker nodes are selected from the nodes respectively included in the first heterogeneous graph and the second heterogeneous graph, including: For each of the grid labels, obtaining a first position node and a first job seeker node associated with the grid label in the first heterogeneous graph to obtain a first node set; and obtaining a second position node and a second job seeker node associated with the grid label in the second heterogeneous graph to obtain a second node set; For each grid label, N target job nodes are selected from the job nodes included in the first node set and the second node set, and M target job seeker nodes are selected from the job seeker nodes included in the first node set and the second node set.

7. The method according to claim 1, characterized in that The clustering process is performed on the plurality of aggregation vectors, and determining a target number of target clusters, including: Randomly selecting a target number of first aggregation vectors from the multiple aggregation vectors as initial cluster centers of the target number of clusters; According to a preset clustering algorithm, each cluster and the cluster center of the cluster are iteratively updated until a preset stopping iteration condition is reached to obtain the target cluster; The second aggregation vector is an aggregation vector among the multiple aggregation vectors except the first aggregation vector.

8. A grid label classification device, characterized in that: include: an acquisition module, configured to acquire a first heterogeneous graph corresponding to job-seeking behavior information within a preset time period and a second heterogeneous graph corresponding to recruitment behavior information within the preset time period, and determine vector information of each node in the first heterogeneous graph and vector information of each node in the second heterogeneous graph, wherein the first heterogeneous graph includes first relationship information between job seeker nodes and position nodes, and the second heterogeneous graph includes second relationship information between job seeker nodes and position nodes; The acquisition module is further configured to acquire recruitment information within the preset time period, and determine, based on the recruitment information, a grid label associated with each position node included in the first heterogeneous graph and the second heterogeneous graph; and, based on the job-seeking behavior information and the recruitment information, determine a grid label associated with each job seeker node included in the first heterogeneous graph and the second heterogeneous graph; a processing module configured to select, for each of the grid labels, N target position nodes and M target job seeker nodes from the nodes respectively included in the first heterogeneous graph and the second heterogeneous graph, wherein the N target position nodes include the position node associated with the grid label, and the M target job seeker nodes include the job seeker node associated with the grid label; The processing module is further configured to determine an aggregate vector for each grid label based on the vector information of the N target job nodes and the vector information of the M target job seeker nodes corresponding to each grid label; The processing module is further configured to perform clustering processing on the plurality of aggregated vectors and determine a target number of target clusters, wherein the target number is determined according to the classification fineness required by the business application, and different target clusters correspond to different grid categories.

9. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; and when the electronic device executes the computer program instructions, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 7.

11. A computer program product, characterized in that The method comprises computer program instructions, which implement the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Online recruitment bidirectional reciprocity recommendation system and method based on multi-behavior modeling

    CN116662676A

  • Encoding a job posting as an embedding using a graph neural network

    US20230125711A1