A method and system for identifying false job recruitments

By extracting multiple indicator data from recruitment interview information and using machine learning models to evaluate it, the problem of narrow sources of false recruitment identification data and single identification dimensions in the existing technology is solved, and multi-dimensional false recruitment identification and risk prediction are achieved, which improves the defense ability of job seekers before interview.

CN114626810BActive Publication Date: 2025-07-29QIAN JIN NETWORK INFORMATION TECH SHANGHAI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210232584.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-09
Publication Date
2025-07-29
Estimated Expiration
2042-03-09

AI Technical Summary

Technical Problem

When identifying false recruitment information, the data source range is narrow and the recognition dimension is single, resulting in insufficient accuracy of false recruitment testing, especially in the interview notification process that fails to effectively detect false recruitment.

Method used

The target enterprise information, job information and interview information are extracted from the recruitment interview information provided by users, and the data is crawled in the network through multiple indicators at the enterprise level, position level and interview level, and processed it into the corresponding level indicator data. The machine learning model is used to evaluate the risks at the enterprise level, position level and interview level.

Benefits of technology

It improves the accuracy of false recruitment identification, and actively searches for multiple information, increases the accuracy of judgment, and provides the most sufficient data assessment in the job seeker interview notification process to help job seekers avoid risks to the greatest extent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114626810B_ABST
    Figure CN114626810B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for identifying false recruitment. The method includes the following steps: extracting target recruitment enterprise information, target position information, and interview information from the recruitment interview information provided by a user; based on the extracted target recruitment enterprise information, target position information, and interview information, respectively crawling information in the network according to multiple indicators in enterprise level, position level, and interview level, and processing the crawled information into index data of corresponding levels; and evaluating according to the multiple index data of each level to obtain the enterprise-level risk, position-level risk, and interview-level risk of the target enterprise for false recruitment. The present invention identifies false recruitment from multiple aspects and dimensions, with high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for identifying fake job recruitments. Background Art

[0002] Fake job recruitments generally refer to fraudulent recruitments where the recruitment information posted on the Internet or in the talent market does not match the actual employment situation. It can refer to fake recruitments carried out by regular companies or illegal recruitments by fake companies. Some regular companies, for reasons such as collecting talent information and promoting company popularity, create some false positions and post recruitment information on the Internet and / or in the talent market. Such fake recruitment behaviors not only delay the job-seeking opportunities of job seekers, endanger personal information security, but also occupy and waste the job-seeking public platform, damaging the credibility of the recruitment platform. Illegal recruitments by fake companies can not only cause property and personal damage to job seekers, such as charging fees randomly, pyramid schemes, and involvement in porn, drugs, and gambling, but also infringe on the interests of third parties, such as regular companies whose names are misappropriated and recruitment platforms that post recruitment information for them.

[0003] To identify these fake job recruitment information, the industry has made a lot of efforts. The Chinese patent application with the publication number CN113704409 A and the invention name "A method for detecting fake job recruitment information based on cascade forest" provides a method for detecting fake job recruitment information. Based on the cascade forest algorithm of decision trees, a model is established using the job data posted on the online recruitment platform for fake job recruitment prediction. The Chinese patent application with the publication number CN 113506084 A and the invention name "A method for detecting fake job positions based on deep learning" provides a method for detecting fake job positions. Fake job recruitment information is collected from online recruitment platforms or recruitment APPs, and the fake job recruitment information is processed into sample data for training the model. A detection model is obtained through training, and the detection model is used to detect fake job recruitment information on the online recruitment platform or recruitment APP.

[0004] From the above-provided technical solutions, it can be seen that currently, when identifying fake job recruitment information, the training data for training the identification model comes from fake job information and / or fake job recruitment information collected from online recruitment platforms. The source range of the data is narrow, and the identification dimension is single. According to the foregoing description of fake job recruitments, fraudulent fake job recruitments have multiple aspects. If only limited to identifying from the position aspect, a large number of fake job recruitments are bound to be missed. Therefore, there is an urgent need for a solution to identify fake job recruitments from multiple aspects and dimensions. Summary of the Invention

[0005] In view of the technical problems existing in the prior art, the present invention proposes a method and system for identifying fake job recruitments, which can identify fake job recruitments from multiple aspects and dimensions.

[0006] To solve the above technical problems, according to one aspect of the present invention, the present invention provides a method for identifying false recruitment, including the following steps:

[0007] Extract the target recruitment enterprise information, target position information, and interview information from the recruitment interview information provided by the user;

[0008] Based on the extracted target recruitment enterprise information, target position information, and interview information, crawl information in the network according to multiple indicators in enterprise level, position level, and interview level respectively, and process the crawled information into indicator data of the corresponding level; and

[0009] Evaluate according to the multiple indicator data at each level to obtain the enterprise-level risk, position-level risk, and interview-level risk of false recruitment by the target enterprise.

[0010] According to one aspect of the present invention, the present invention provides a false recruitment identification system, which includes a user information acquisition module, a data collection module, an indicator data generation module, and a risk assessment module. Among them, the user information acquisition module is configured to receive the recruitment interview information provided by the user, and extract the target recruitment enterprise information, target position information, and interview information from the user recruitment interview information; the data collection module is connected to the Internet and is connected to the user information acquisition module, and is configured to crawl information in the network according to multiple indicators in enterprise level, position level, and interview level respectively based on the extracted target recruitment enterprise information, target position information, and interview information; the indicator data generation module is connected to the data collection module, and processes the crawled enterprise-level information, position-level information, and interview-level information to obtain multiple indicator data of the corresponding level; the risk assessment module is connected to the indicator data generation module, and is configured to evaluate according to the multiple indicator data at each level to obtain the enterprise-level risk, position-level risk, and interview-level risk of false recruitment by the target enterprise.

[0011] As can be seen from the method and system provided by the present invention above, the present invention makes up for the lack of predecessors in actively searching for information and risk discrimination: traditional algorithms only use information related to job descriptions to determine whether a recruitment is false, while the present invention actively searches for other relevant information, such as enterprise information, related enterprise information, target job information, related job information, etc., thus increasing the accuracy of judgment. The present invention makes up for the lack of predecessors in predicting risks in the interview notice link: existing false recruitment detection schemes are earlier than the interview notice link and cannot effectively detect false recruitments that show risks only at the interview notice stage. However, after a job seeker obtains an interview notice from the target enterprise, the present invention analyzes and evaluates based on the interview information given by the job seeker. This timing is not only the last line of defense before the job seeker goes to the interview but also the best timing to obtain the most sufficient data to predict risks, thereby being able to help job seekers avoid risks directionally to the greatest extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Next, the preferred embodiments of the present invention will be further described in detail with reference to the drawings, where:

[0013] Figure 1 is a flowchart of a method for identifying false recruitments according to an embodiment of the present invention;

[0014] Figure 2 is a flowchart of risk assessment according to an embodiment of the present invention;

[0015] Figure 3 is a flowchart of risk assessment according to another embodiment of the present invention;

[0016] Figure 4 is a flowchart of risk assessment according to yet another embodiment of the present invention;

[0017] Figure 5 is a flowchart of assessing risks according to another embodiment of the present invention;

[0018] Figure 6 is a flowchart of a method for processing test sample data according to an embodiment of the present invention;

[0019] Figure 7 is a flowchart of a false recruitment early warning method according to an embodiment of the present invention;

[0020] Figure 8 is a block diagram of the principle of a false recruitment identification system provided according to an embodiment of the present invention;

[0021] Figure 9 is a partial block diagram of the principle of a false recruitment identification system provided according to an embodiment of the present invention;

[0022] Figure 10It is a partial principle block diagram of a false recruitment identification system provided according to an embodiment of the present invention;

[0023] Figure 11 It is a principle block diagram of a user information acquisition module according to an embodiment of the present invention;

[0024] Figure 12 It is a principle block diagram of a data processing system according to an embodiment of the present invention;

[0025] Figure 13 It is a principle block diagram of a data collection module according to an embodiment of the present invention;

[0026] Figure 14 It is a principle block diagram of an index data generation module according to an embodiment of the present invention;

[0027] Figure 15 It is a principle block diagram of a model training module according to an embodiment of the present invention;

[0028] Figure 16 It is a principle block diagram of a false recruitment warning system according to an embodiment of the present invention;

[0029] Figure 17 It is a principle block diagram of a warning module according to an embodiment of the present invention;

[0030] Figure 18 It is a principle block diagram of a warning module according to another embodiment of the present invention; and

[0031] Figure 19 It is a principle block diagram of a warning module according to yet another embodiment of the present invention. Detailed implementation manners

[0032] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] In the following detailed description, reference may be made to the accompanying drawings that form a part hereof and in which are shown, by way of illustration, specific embodiments in which the application may be practiced. In the drawings, like reference numerals describe substantially similar components in different figures. The various specific embodiments of the application have been described in sufficient detail below to enable those of ordinary skill in the relevant art and technology to practice the technical solutions of the application. It should be understood that other embodiments may be utilized or structural, logical or electrical changes may be made to the embodiments of the application.

[0034] Figure 1 It is a flowchart of a fake recruitment identification method according to an embodiment of the present invention. In this embodiment, the method includes the following steps:

[0035] Step S1a, determine whether recruitment interview information provided by the user is received. If the recruitment interview information provided by the user is received, execute step S2a; if not, repeat this step. In one embodiment, when the method is applied to an online recruitment platform, such as a recruitment website, a recruitment APP, etc., the online recruitment platform opens an interface to the user to facilitate the user to input recruitment interview information through this interface. The interface can be implemented as a user interface, for example, through which the user inputs recruitment interview information. In some embodiments, the user interface is further provided with multiple information segments for inputting target recruitment enterprise information, target position information, and interview information respectively. The target recruitment enterprise information includes, for example, information such as enterprise name and address. The target position information includes, for example, information such as position name and work location. The interview information includes, for example, information such as interview location, interview time, interview notification method, and interview form, such as oral test, written test, etc.

[0036] Step S2a, extract the target recruitment enterprise information, target position information, and interview information from the recruitment interview information provided by the user. In one embodiment, when the user inputs various information through the information segments in the user interface, the information in the field is read from the information segment, and the information is processed to obtain one or more keywords. The processing includes, for example, stop word removal, word segmentation, etc.

[0037] Step S3a, based on the extracted target recruitment enterprise information, target position information, and interview information, perform information crawling in the network according to multiple indicators in enterprise level, position level, and interview level respectively.

[0038] Among them, in one embodiment, enterprise-level indicators include static indicators and dynamic indicators. The static indicators include, but are not limited to, one or more of the number of employees, number of branch companies, registered capital, financing information, enterprise age, annual turnover, business scope, and enterprise type of the target recruitment enterprise. The dynamic indicators include, but are not limited to, one or more of the enterprise legal litigation events, major investments received, social news events occurred, and enterprise evaluations on social media of the target recruitment enterprise. The position-level indicators include, but are not limited to, one or more of the position name, work location, work department, monthly salary range, job description, welfare benefits, work type, experience requirements, educational requirements, school type, work industry risk, and number of repeated recruitments for the same position. The interview-level indicators include, but are not limited to, one or more of the interview location, interview time, interview notification method, number of interviews conducted, whether there is a written test, and interview form. Corresponding information is obtained from the network according to these indicators. For example, according to the information such as the name and address of the target recruitment enterprise provided by the user, the official website, social media public account, public news reports, etc. of the target recruitment enterprise can be searched in the network using its name and / or address as keywords.

[0039] In one embodiment, the number of employees information, information including "xx branch company", business scope description information, information containing enterprise categories, financing-related information, enterprise age description information, annual turnover description information, registered capital information, etc. are searched and obtained on the official website of the target recruitment enterprise. If a certain piece of the above information is not obtained from the official website of the target recruitment enterprise, continue to search in the network. For example, if the registered capital, enterprise type and other information are not obtained from the official website of the target recruitment enterprise, the registration information of the target recruitment enterprise can be obtained from the industrial and commercial information network, and the registered capital and enterprise type can be obtained therefrom. Another example, if the financing information is not obtained from the official website of the target recruitment enterprise, the financing-related information can be obtained by querying the social media public account or public news reports of the target recruitment enterprise. The comment information, litigation information, news information about received investments, etc. of the target recruitment enterprise in the recent five years are identified on the web page. Among them, in order to obtain the data of the position-level indicators, the position information of the same position and other positions published by the target recruitment enterprise is searched on multiple recruitment platforms, and at the same time, the position information of the same position and other positions published by enterprises of the same type as the target recruitment enterprise is also searched. The information related to the position is obtained from the official website of the target recruitment enterprise. The job posting information released by the target recruitment enterprise within a certain period of time is obtained from the network.

[0040] Step S4a: Process the crawled information into corresponding level index data. Process the crawled corresponding information according to the processing strategy of specific indexes, and standardize the corresponding information into specific index data according to various indexes. The index data can be floating-point numbers, string vectors, or a certain encoding.

[0041] For example, the processing of enterprise-level static indexes includes: Regarding the number of employees in the enterprise, process the data of the number of employees of the target recruitment enterprise obtained by crawling into a floating-point value. Regarding the number of branch companies, calculate the number according to the words containing "**Branch Company" identified from the company introduction on the official website, and process it into a floating-point value; Regarding the registered capital, according to the words containing "Registered Capital**" identified from the company registration information introduction, process the value therein into a floating-point value. Regarding the business scope, identify the keywords in the business scope description and process them into a string vector. Regarding the enterprise type, process the enterprise category information on the official website into a string vector, and match it among "State-owned enterprise", "Foreign-funded enterprise", "Joint venture", and "Private enterprise". The above options are processed as 3, 2, 1, and 0 respectively, and are recorded. Regarding the financing information, process the net value in the latest financial report of the enterprise obtained into a floating-point value. Regarding the enterprise age, subtract the establishment time of the enterprise obtained from the current year to obtain the number of years of establishment and process it into a floating-point value. Regarding the annual turnover, process the annual turnover obtained from the latest financial report of the enterprise into a floating-point value.

[0042] For the processing of enterprise-level dynamic indexes includes: Regarding the enterprise legal litigation events, first extract the litigation information of the target recruitment enterprise in the recent period (such as five years) from the obtained information; then use keywords such as "Labor" and "Defendant** (name of the target recruitment enterprise)" to perform regular matching on each piece of information. If the match is successful, record it as 1, and if the match fails, record it as 0; When the match is successful, obtain the litigation time information from the litigation information and encode it. For example, when the litigation time is less than 1 year, encode it as 1; when the litigation time is less than 2 years and greater than 1 year, encode it as 0.8; when the litigation time is less than 3 years and greater than 2 years, encode it as 0.6; when the litigation time is less than 4 years and greater than 3 years, encode it as 0.5; when the litigation time is less than 5 years and greater than 4 years, encode it as 0.3; then multiply all the encoded values of successful matches by the encoded value of the litigation time to obtain the litigation code; Finally, add up all the litigation codes to obtain the litigation event score of the enterprise.

[0043] Regarding the major investments received by the enterprise, first extract the news information related to investments received in the recent period (such as five years) from the crawled information; then process the amount of investment received into a floating-point value; Then add up all the floating-point values to obtain the total amount of major investments received by the target recruitment enterprise in the recent period (such as five years); If not, the default value is 0.

[0044] Regarding the evaluation of social media enterprises, first extract multiple comment information from the captured information for a recent period (such as five years); then process each comment information according to the following process: remove HTML symbols in the comment through a parsing method; remove punctuation through regular expressions; use a stop word library to filter and remove all stop words; label the data text after the above cleaning; use a sentiment analysis model to analyze the text, and output each comment as a sentiment encoding of -1 (negative) or 1 (positive) for the comment text; then obtain the comment time information, and encode it as 1 when the comment time is less than 1 year from the current, encode it as 0.8 when the comment time is less than 2 years and greater than 1 year from the current, encode it as 0.6 when the comment time is less than 3 years and greater than 2 years from the current; encode it as 0.5 when the comment time is less than 4 years and greater than 3 years from the current; encode it as 0.3 when the comment time is less than 5 years and greater than 4 years from the current; then multiply the sentiment encoding of all comment texts by the time encoding to obtain the sentiment score of each comment. For example, if an enterprise has received a negative comment, and the time of occurrence is less than 4 years and greater than 3 years from the current, then the sentiment score of this comment is -1 * 0.5 = -0.5; the enterprise has also received a positive comment, and the time of occurrence is less than 3 years and greater than 2 years from the current, then the sentiment score of this comment is 1 * 0.6 = 0.6; then add up the sentiment scores of all comments of the enterprise to obtain the social media evaluation score of the enterprise. The social media evaluation score of the enterprise in the above example is -0.5 + 0.6 = 0.1.

[0045] The processing of indicator data at the position level includes: for the position name, it is processed into a string vector; for the work location, first, the extracted work location information is processed into a string vector; the corresponding geographical location information is found through a map search engine; then, the name of the owner / tenant of the location building is obtained and matched with the company name; if the target recruitment company name can be matched, it is encoded as 0, otherwise it is encoded as 1. For the work department, first, the extracted work department information is processed into a string vector; then it is matched with the information on the company's official website, and if the corresponding string can be matched on the official website, it is encoded as 0, otherwise it is encoded as 1. For the monthly salary range, it is processed into a floating-point value; for the job description, first, HTML symbols, spaces, and punctuation marks are removed through parsing methods and regular expressions; then stop words are identified and removed through a stop word library; finally, the remaining text is processed into a string vector. Regarding the welfare benefits, first, HTML symbols, spaces, and punctuation marks are removed through parsing methods and regular expressions; then stop words are identified and removed through a stop word library; finally, the remaining text is processed into a string vector. Regarding the work type, check the crawled relevant information to see if there are characters such as "full-time", "part-time", "labor dispatch", "hourly wage", etc. If there is "full-time", it is recorded as 3, if there is "part-time", it is recorded as 2, if there is "labor dispatch", it is recorded as 1, if there is "hourly wage", it is recorded as 0, and if none of them exist, it is recorded as a null value. Regarding the experience requirement, first, check if the crawled position-related information contains the string "experience". If it does, identify the numerical value before the string "year" to obtain the required number of working years and process it into a floating-point value, and record the minimum value among them; if not, check if there is the string "fresh graduate". If there is, record it as 0, and if not, record it as a null value. Regarding the educational requirement, first, identify the description of the educational requirement from the crawled position-related information and match it among the strings "master or above", "bachelor's degree", "college degree", "technical secondary school or below", "no requirement". The above options are processed and recorded as 4, 3, 2, 1, 0 respectively. For the school type, identify the description of the graduating school from the crawled position-related information and match it among the strings "985", "211", etc. If "985" is identified, it is recorded as 2, if "211" is identified, it is recorded as 1, if both "985" and "211" are identified, it is also recorded as 1, and if no matching field is identified, it is recorded as 0. For the work industry risk, first, match the crawled job description information in the "Classification of National Economy Sectors of the People's Republic of China" to obtain the affiliated industry; then search for relevant news using the affiliated industry as the keyword. If keywords such as "layoffs", "transformation", "market value evaporation", "store closures" are searched, it is defined as a high risk and recorded as 1, otherwise it is defined as a low risk and recorded as 0.Regarding the number of times of repeated recruitment for the same position, first, all job positions published by the target recruitment enterprises crawled in the recent period (such as three years) are captured; then, the duration of the same position published online is identified and converted into a floating-point value on an annual basis; if a position goes offline and then goes online again, the online durations are added separately and converted into a floating-point value.

[0046] The job title, work location, work department, monthly salary range, job description, welfare benefits, work type, experience requirements, educational requirements, school type, work industry risk, the number of times of repeated recruitment for the same position, etc. of the aforementioned target position are deep indicators for a target position, and the collection after their standardized processing can be used as a prediction sample set for the target position. In another embodiment, while obtaining the deep indicator data of the target position, it further includes the step of establishing an enterprise online recruitment information matrix according to the position level information, and obtaining relevant indicators in terms of breadth for the recruitment platform and the target position from the enterprise online recruitment information matrix, hereinafter referred to as breadth indicators.

[0047] Among them, the enterprise online recruitment information matrix includes one or more job information released by the target enterprise and the recruitment platform qualification information for releasing these job information; job information released by one or more similar enterprises similar to the target enterprise and the recruitment platform qualification information for releasing these job information. Among them, the job positions released by the target enterprise include the target position and a second position different from the target position; the recruitment platforms include one or more target recruitment platforms for releasing the target position provided by the target enterprise and one or more second recruitment platforms for releasing the second job information different from the target position.

[0048] The steps for processing the corresponding breadth indicator data based on the enterprise online recruitment information matrix include:

[0049] Calculating the qualification level vector of each recruitment platform based on the recruitment platform qualification information;

[0050] Counting the number of job positions released by each recruitment platform from the enterprise online recruitment information matrix;

[0051] Determining the input degree of each recruitment platform based on the number of job positions; for example, taking the number of job positions as the input degree, or normalizing the number of job positions and taking the normalized value as the input degree.

[0052] Taking the input degree of each recruitment platform as the weight of the qualification level vector of the recruitment platform, and calculating the weighted vector value of the qualification level of each recruitment platform. Among them, the weighted vector value of the qualification level of the recruitment platform is a kind of breadth indicator.

[0053] The steps for processing the corresponding breadth indicator data based on the enterprise online recruitment information matrix further include:

[0054] Query multiple second target positions that are the same as the target position and provided by similar enterprises;

[0055] Calculate the position vectors of the target position and each second target position based on the position information;

[0056] Perform clustering operations on the multiple second target positions to obtain the second target position with the largest clustering result;

[0057] Calculate the first vector difference between the position vector of the target position and the position vector of the second target position with the largest clustering result, and use the first vector difference as the external consistency coefficient of the target position. Among them, the external consistency coefficient of the target position is another breadth indicator.

[0058] The steps of processing the enterprise network recruitment information matrix to obtain the corresponding breadth indicator data further include:

[0059] Calculate the second vector difference between the position vector of the target position and the position vector of each second position provided by the target enterprise, and calculate the average value of the second vector differences, and use the average value of the second vector differences as the internal consistency coefficient of the target position. Among them, the internal consistency coefficient of the target position is another breadth indicator.

[0060] The steps of processing the enterprise network recruitment information matrix to obtain the corresponding breadth indicator data further include:

[0061] Determine the weights of the first vector difference and the average value of the second vector differences according to their respective closeness relationships with the target position vector, and calculate the weighted average of the two as the consistency coefficient of the target position. Among them, the consistency coefficient of the target position is another breadth indicator.

[0062] In one embodiment, the obtained depth indicator data and breadth indicator data are combined together as the prediction sample set of the target position.

[0063] The processing of interview-level indicator data includes: regarding the interview location, first process the interview location information into a string vector; then find the corresponding geographical location information through a map search engine; then obtain the name of the owner / tenant of the location building and match it with the name of the target recruiting enterprise; if the name of the target recruiting enterprise is matched, it is encoded as 0, otherwise it is 1. For the interview time, if the interview time information is between 8:00 and 18:00, it is recorded as 0, otherwise it is recorded as 1. For the interview notification method, if it is by phone contact, it is encoded as 1, if it is by text message, it is encoded as 2, if it is by email, it is encoded as 3, if it is by recruitment application (recruitment App), it is encoded as 4, if it is by ordinary social software, it is encoded as 5, and multiple contact methods are allowed to be filled in. Regarding whether there is a written test, if there is a written test, it is encoded as 0, otherwise it is encoded as 1.

[0064] Step S5a, evaluate the enterprise-level risk, position-level risk, and interview-level risk of false recruitment of the target enterprise according to multiple indicator data at each level. In this embodiment, when evaluating risks at each level, a trained model is used for evaluation respectively. The model can adopt algorithms such as decision tree, naive Bayes, multi-layer, k-nearest neighbor, random forest, neural network, etc.

[0065] In this embodiment, the machine learning models at each level are respectively trained by using labeled samples at each level. Taking the training of the enterprise-level machine learning model as an example, the process of training the machine learning model is described.

[0066] The training data of the enterprise-level machine learning model includes a certain number of enterprise sample data. Each enterprise sample data includes multiple indicator data labeled as a false recruitment enterprise or a real recruitment enterprise. Each indicator data is used as a feature data of the enterprise sample. For example, factors such as "number of employees in the enterprise", "number of branch companies", "registered capital", "business scope", "enterprise type", "financing information", "enterprise age", "annual turnover", "enterprise legal litigation events", "major investments received", "enterprise evaluations on social media", etc. are used as independent variables (i.e., one-dimensional features of a sample), and whether there is an enterprise-level risk for each sample is manually marked. For example, a sample with risk is marked as "negative" / "positive", and a sample without risk is marked as "positive" / "negative", so as to obtain a sample set. The sample set should meet some requirements for training the model, such as the number balance of positive and negative samples, the same feature dimension of each sample, etc. 80% of the samples in the sample set are used as training samples, and 20% of the samples are used as verification samples, so as to form a training set and a verification set respectively.

[0067] Then, a model is built using any one of the algorithms such as the aforementioned decision tree, naive Bayes, multi-layer, k-nearest neighbor, random forest, neural network, etc. The model performs supervised learning based on the samples in the training set according to the model algorithm to obtain the probability of enterprise-level risk for each sample.

[0068] Since only two result types, "at risk" and "not at risk", are set in this embodiment, the model uses binary classification output. Assuming the number of independent variables of the sample (i.e., the dimension of the sample features) is n, the mathematical representation of the model is:

[0069] X = {x1, x2,..., x n}

[0070] Y = {y0, y1}

[0071] Where X is the feature set of a sample input to the model, and each x i is a one-dimensional feature. For example: x1 is the number of employees in the enterprise, x2 is the number of branch companies, x3 is the registered capital,......, x n is the annual turnover. Y is the classification set of the model, and y i ∈ {0, 1}. In this embodiment, y0 represents "not at risk" and y1 represents "at risk". The model calculates, based on a given sample X i , the probability p(c j (the label corresponding to each classification y. For example, when y0 takes the value 0, the corresponding label c0 is "no enterprise-level risk")) given the sample X j . Then the output of the model is: i ) is:

[0072]

[0073] Where represents the classification prediction value, and p(c j |X i ) represents the probability of each classification given the sample. In one embodiment, a comparison threshold is set, such as set to 0.6 - 0.9. In this embodiment, a threshold of 0.8 is used for illustration. When determining whether there is enterprise-level risk, if the predicted probability of y1 output by the algorithm model exceeds 0.8, it can be confirmed that the enterprise corresponding to the sample has enterprise-level risk. Among them, the threshold can be an empirical value obtained by repeatedly operating the model based on practical data, or further adjusted as the model iterates. The trained model is verified using the validation set samples. When the model meets the evaluation criteria, it can be used for online risk assessment.

[0074] The training processes of the position-level machine learning model and the interview-level machine learning model are similar to the foregoing process, and will not be elaborated herein.

[0075] Among them, in step S4a, in order to make the normalized data meet the input requirements of the model, that is, a sample consists of multi-dimensional features and the number of features is the same. After specifically normalizing the data of each level index into floating-point numbers, string vectors or codes, according to the input requirements of each level model, these data are combined into the prediction samples of the model, and then in step S5a, the prediction samples of each level are input into the corresponding model, so as to obtain the risks corresponding to each level.

[0076] In the foregoing embodiment, the output of the machine learning model at each level is "at risk" or "not at risk". In one embodiment, as Figure 2 shown, in steps S510a, 511a and 512a respectively, each prediction sample is input into the machine learning model at the corresponding level, and then in step S513a, the results output by each level of machine learning model, that is, the enterprise-level risk, position-level risk and interview-level risk, are combined as the evaluation result, and the evaluation result is provided to the user in step S6a, so that the user can know whether there is a risk from the three aspects of enterprise, position and interview.

[0077] In another embodiment, as Figure 3 shown, taking "at risk" or "not at risk" output by the model as two levels, represented by 1 and 0 respectively, so that the risk levels of enterprise level, position level and interview level can be encoded to obtain risk codes. After obtaining the risks at the corresponding levels by inputting each prediction sample into the machine learning model at the corresponding level in steps S510a, 511a and 512a, the enterprise-level risk, position-level risk and interview-level risk are encoded in step S523a. For example, the risk code 001 represents that there is a risk only during the interview, while the risk code 111 represents that there are risks at the enterprise, position and interview. In order to enable the user to feel the magnitude of the risk, different risk levels are set in this embodiment, and the risk codes correspond to the risk levels, as shown in Table 1 below:

[0078] Table 1:

[0079] Risk Code Risk Level 111 High Risk 110、101 Higher Risk 100 Medium Risk 010、011 Lower Risk 000、001 Low Risk

[0080] In step S524a, according to the obtained risk code, query the corresponding table of risk code and risk level, as shown in Table 1, then the risk level corresponding to the obtained risk code can be obtained and determined as the final risk level, and then in step S6a, the final risk level is provided to the user as the evaluation result.

[0081] In another embodiment, as Figure 4 shown, after encoding the obtained enterprise-level, position-level and interview-level risk levels in step S533a, it further includes:

[0082] Step S534a, obtain the respective weights of the enterprise level, position level, and interview level.

[0083] Step S535a, perform weighted calculation according to the values on each digit in the risk code and their respective weights to obtain the weighted sum of the risk code.

[0084] Step S536a, query the correspondence table between the weighted sum of the risk code and the risk level according to the current weighted sum of the risk code, as shown in Table 2, so as to obtain the risk level matching the current weighted sum of the risk code.

[0085] Among them, in one embodiment, according to the weights of the enterprise level, position level, and interview level being 5, 4, and 1 respectively, calculate the weighted sums of the risk codes 000 - 111 respectively, and obtain 0, 1, 4, 5, 5, 6, 9, 10. Then, according to the gap between adjacent values, divide the following 8 values into 3 groups, namely (0, 1), (4, 5, 6), (9, 10), thus obtaining Table 2:

[0086] Table 2:

[0087] Weighted Sum of Risk Codes Risk Level 9、10 High Risk 4,5、6 Medium Risk 0、1 Low Risk

[0088] For example, in one embodiment, when the following risks are obtained according to the models at each level: enterprise level: "There is a risk"; position level: "No risk"; interview level: "No risk", the risk code is 100, and the corresponding weighted sum of the risk code is 5. Query Table 2, so as to match the corresponding risk level of "Medium risk".

[0089] Finally, in step S6a, provide the risk level queried from Table 2 to the user as the evaluation result.

[0090] In the above embodiments, the models at each level adopt binary classification output, that is, two risk levels of "There is a risk" and "No risk". Of course, it is also possible to train the models at each level to adopt multi-classification output. For example, when using 5 levels of "High risk", "Higher risk", "Medium risk", "Lower risk", and "Low risk", assuming the number of sample independent variables (i.e., the dimension of sample features) is n, the mathematical representation of the model is:

[0091] X = {x1, x2,..., x n}

[0092] Y = {y0, y1, y2, y3, y4}

[0093] Wherein the X is the feature set of a sample input to the model, and each x i is a one-dimensional feature. For example, x1 is the number of enterprise employees, x2 is the number of branch companies, x3 is the registered capital,..., xn is the annual turnover. Y is the classification set of the model, and y i ∈ {0, 1, 2, 3, 4}. In this embodiment, y0 represents "low risk", y1 represents "relatively low risk", y2 represents "medium risk", y3 represents "relatively high risk", and y4 represents "high risk". The model is based on a given sample X i , and calculates the probability p(c j (the label corresponding to each classification y. For example, when y0 takes the value 0, the corresponding label c0 is "risk-free") given X j |X i ). Then the output of the model is:

[0094]

[0095] where represents the classification prediction value, and p(c j |x i ) represents the probability of each classification under the given sample. In one embodiment, a comparison threshold is set, generally set to 0.6 - 0.9. In this embodiment, the threshold of 0.8 is used for illustration. For example, when judging the enterprise-level risk category, if the predicted probability of y1 output by the algorithm model exceeds 0.8, it can be confirmed that the enterprise corresponding to the sample has a relatively low enterprise-level risk. Among them, the threshold can be an empirical value obtained by repeatedly operating the model with practical data, or further adjusted with the iteration of the model. The trained model is verified using the validation set samples, and when the model meets the evaluation criteria, it can be used for online risk assessment.

[0096] The training processes of the position-level machine learning model and the interview-level machine learning model are similar to the foregoing processes, and will not be elaborated here.

[0097] After obtaining the risk level of each level, the foregoing embodiment is used for encoding or calculating the weighted sum of the encodings to determine the final risk level, which will not be elaborated here.

[0098] Figure 5It is a flowchart for evaluating risks in another embodiment of the present invention. In this embodiment, there is one enterprise-level machine learning model, and multiple position-level machine learning models and interview-level machine learning models. The output categories of the enterprise-level machine learning model, the position-level machine learning models, and the interview-level machine learning models are more than two. The number of output categories of these three models can be the same or different. In the order from the enterprise level, the position level to the interview level, the levels are set from top to bottom in sequence. The number of machine learning models at the lower level is the same as the number of output categories of the machine learning model at the upper level, and they are respectively trained with data having risks of the output categories of the machine learning model at the upper level. For example, multiple position-level machine learning models are respectively trained with data having corresponding enterprise-level risk levels, and respectively correspond to the risk levels output by the corresponding enterprise-level machine learning models; multiple interview-level machine learning models are respectively trained with data having corresponding enterprise-level risk levels and corresponding position-level risk levels, and respectively correspond to the risk levels output by the corresponding position-level machine learning models.

[0099] For example, in one embodiment, the outputs of the enterprise-level machine learning model and the position-level machine learning model are respectively binary classification outputs, which are respectively defined as two risk levels of "at risk" and "not at risk". The interview-level machine learning model is a multi-classification output. For example, it is respectively defined as 5 levels of "high risk", "relatively high risk", "medium risk", "low risk" and "not at risk". The enterprise-level machine learning model is denoted as M1, and there are two position-level machine learning models, which are respectively denoted as model M2.1 and model M2.2. Among them, the position samples in the training set of model M2.1 are all samples corresponding to the enterprise level without risk, while the position samples in the training set of model M2.2 are all samples corresponding to the enterprise level at risk. There are 4 interview-level machine learning models, which are respectively denoted as model M3.1, M3.2, M3.3 and model M3.4. The position samples in the training set of model M3.1 are all samples corresponding to the enterprise level without risk and the position level without risk. The position samples in the training set of model M3.2 are all samples corresponding to the enterprise level without risk and the position level at risk. The position samples in the training set of model M3.3 are all samples corresponding to the enterprise level at risk and the position level without risk. The position samples in the training set of model M3.4 are all samples corresponding to the enterprise level at risk and the position level at risk.

[0100] From the training samples of the machine learning models, it can be seen that the enterprise-level machine learning model is the first level, the position-level machine learning models are its subordinate models, and the interview-level machine learning models are the subordinate models of the position-level machine learning models. The selection of the subordinate models should correspond to the risk and other outputs of the superior models.

[0101] Select the next-level machine learning model according to the risk level output by the previous-level machine learning model in the order from top to bottom of enterprise level, position level, and interview level, where the risk and level output by the interview-level machine learning model are the final risk and its level. The specific risk assessment process is as Figure 5 shown and includes the following steps:

[0102] Step S51a: Input the enterprise-level prediction sample into the enterprise-level machine learning model M1 for prediction. In one embodiment, the enterprise-level machine learning model is denoted as M1, and the enterprise-level prediction sample is input into this model M1.

[0103] Step S52a: Determine whether the output of the enterprise-level machine learning model M1 is "at risk". If it is "at risk", execute Step S53a; if it is "not at risk", execute Step S57a.

[0104] Step S53a: Select the position-level machine learning model M2.2, and input the position-level prediction sample into the position-level machine learning model M2.2 for prediction.

[0105] Step S54a: Determine whether the output of the position-level machine learning model M2.2 is "at risk". If it is "at risk", execute Step S55a; if it is "not at risk", execute Step S56a.

[0106] Step S55a: Select the interview-level machine learning model M3.4, and input the interview-level prediction sample into the interview-level machine learning model M3.4 for prediction. Take the output of inputting the interview-level prediction sample into model M3.4 as the risk assessment result, and then end the risk assessment process.

[0107] Step S56a: Select the interview-level machine learning model M3.3, and input the interview-level prediction sample into the interview-level machine learning model M3.3 for prediction. Take the output of the interview-level machine learning model M3.3 as the risk assessment result, and then end the risk assessment process.

[0108] Step S57a: Select the position-level machine learning model M2.1, and input the position-level prediction sample into the position-level machine learning model M2.1 for prediction.

[0109] Step S58a: Determine whether the output of the position-level machine learning model M2.1 is "at risk". If the output of the position-level machine learning model M2.1 is "at risk", execute Step S59a; if the output of the position-level machine learning model M2.1 is "not at risk", execute Step S510a.

[0110] Step S59a, select the interview-level machine learning model M3.2, input the interview-level prediction samples into the interview-level machine learning model M3.2 for prediction, use the output of the interview-level machine learning model M3.2 as the risk assessment result, and then end the risk assessment process.

[0111] Step S510a, select the interview-level machine learning model M3.1, input the interview-level prediction samples into the interview-level machine learning model M3.1 for prediction, use the output of the interview-level machine learning model M3.1 as the risk assessment result, and then end the risk assessment process.

[0112] In this embodiment, all information is divided according to the overall enterprise level, position level, and interview level. The enterprise-level assessment results subsume the position-level assessment results, and the position-level assessment results subsume the interview-level assessment results, so as to make a comprehensive and accurate risk judgment.

[0113] In this embodiment, the interview-level machine learning model outputs multiple classifications. In one embodiment, the interview-level machine learning model calculates the respective classification probabilities for the input prediction samples. For example, it obtains the probabilities of "high risk", "relatively high risk", "medium risk", "low risk", and "no risk" respectively; then compares each classification probability with the corresponding second classification threshold; if one of the classification probabilities is greater than or equal to the corresponding second classification threshold, it determines that the model output is the risk level defined by the classification.

[0114] In another embodiment, the output of the interview-level machine learning model is binary classification, defined as two risk levels of "at risk" and "not at risk"; when calculating the prediction samples, the interview-level machine learning model calculates the probability of "at risk"; and compares the probability of "at risk" with multiple third classification thresholds, where the multiple third classification thresholds form multiple risk probability intervals from low to high, corresponding to multiple risk levels from low to high; determines the corresponding risk level according to the risk probability interval where the probability of "at risk" output by the interview-level machine learning model is located.

[0115] In another embodiment, the outputs of the enterprise-level machine learning model and the position-level machine learning model can also be the same as the output of the interview-level machine learning model, that is, multiple classifications. Each position-level machine learning model corresponds to one type of output of the enterprise-level machine learning model. Therefore, the number of position-level machine learning models is the same as the number of risk output categories of the enterprise-level machine learning model. Correspondingly, the number of interview-level machine learning models is the product of the number of position-level machine learning models and the number of risk output categories of the position-level machine learning models.

[0116] In step S6a, the determined risk is provided to the user as an identification result, and the providing method includes displaying it in the terminal interface, sending an email, sending a short message through a mobile communication network, or sending a social message to the user's social media account.

[0117] Figure 6 It is a flowchart of a method for processing prediction sample data according to an embodiment of the present invention. In this embodiment, the following steps are included:

[0118] Step S1b, monitoring the feedback information of the user on each identification result. After evaluating the recruitment interview information provided by the user to identify whether the current recruitment faced by the user is a false recruitment and providing the identification result to the user, monitor the feedback information of the user on the identification result. In one embodiment, the feedback information may include confirmation information of "at risk" or "risk-free" in the identification result, or evaluation information of "correct" / "wrong" on the identification result, etc.

[0119] Step S2b, extracting the confirmation information of the user on the identification result. According to the information fed back in step S1b, extract the confirmation information of "at risk" or "risk-free", or the evaluation information of "correct" / "wrong".

[0120] Step S3b, setting corresponding labels for the corresponding prediction samples according to the confirmation information of the user. For example, when the user confirms that the recruitment interview input by him is a real recruitment, set the label of real recruitment for the prediction samples at all levels evaluated for its recruitment. If the user confirms that the recruitment interview input by him is a false recruitment, set the label of false recruitment for the prediction samples at all levels evaluated for its recruitment. In a better embodiment, the feedback information of the user also includes more information fields. For example, when the user confirms it is a false recruitment, the reason that the user needs to fill in. The present invention analyzes the reason filled in by the user to evaluate the root cause of the falsehood from two aspects of "enterprise" and "position", and sets labels for the current prediction samples at all levels.

[0121] Step S4b, storing the prediction samples with labels into the training set, that is, the system stores the prediction samples with labels into the dataset used for training the model, thereby enriching the training data.

[0122] Step S5b, determining whether the model update condition is met. The update condition is, for example, a preset update period, such as updating the model once a week / month, or counting the number of new training samples in the model training dataset and determining whether the number of new training samples reaches the threshold. If the update period is reached or the number of new training samples reaches the threshold, it is determined that the model update condition is met, and step S6b is executed. Otherwise, return to execute step S1b.

[0123] Step S6b, optimize and update the currently used machine learning model using the training dataset.

[0124] The present invention can continuously accumulate training data, without the need for manual setting of labels for the training data. The optimization of the model can be automatically performed without manual intervention, thus saving a large amount of manpower and time. Moreover, with the optimization of the model, the accuracy of model evaluation is gradually improved, enabling more accurate identification of false job recruitments, thereby better safeguarding the interests of users.

[0125] In one embodiment, the present invention further provides a false job recruitment warning method. Refer to Figure 7 , Figure 7 which is a flowchart of the false job recruitment warning method according to an embodiment of the present invention, and includes:

[0126] Step S1c, determine whether recruitment interview information provided by the user is received. If the recruitment interview information provided by the user is received, execute Step S2c; otherwise, repeat this step.

[0127] Step S2c, extract target recruitment enterprise information, target position information, and interview information from the recruitment interview information provided by the user.

[0128] Step S3c, perform information crawling in the network. Based on the extracted target recruitment enterprise information, target position information, and interview information, perform information crawling in the network according to multiple indicators such as enterprise level, position level, and interview level.

[0129] Step S4c, process the crawled information into corresponding level of indicator data.

[0130] Step S5c, perform risk assessment according to the multiple indicator data at each level. For example, perform risk assessment according to the process shown in Figures 2 - 5 any one of them.

[0131] Step S6c, based on the evaluated risk, trace back and analyze the indicator data corresponding to the risk to obtain warning information and provide it to the user. Among them, the warning information includes the problem indicator data and risk content that cause the risk.

[0132] The step of tracing back and analyzing the indicator data corresponding to the risk includes:

[0133] First, according to the monitored risk type, traverse and evaluate the indicator data used when evaluating the risk to calculate the contribution degree of each indicator data to the evaluated risk.

[0134] Then, sort the multiple indicators according to the contribution degree to the risk, and determine the abnormal indicators as the multiple indicators with the highest ranking or the indicators with a contribution degree greater than the threshold.

[0135] Finally, risk content is generated based on the content of the abnormal indicator and / or the correlation relationship of the content of multiple abnormal indicators.

[0136] Among them, the feature importance determination method in feature engineering can be used to determine the contribution degree of indicator data to the evaluated risk. The feature importance determination method mentioned above can be, for example, the expert meeting method, in which experts in the field specify and identify the importance of each indicator used in the present invention. Or, methods such as rough set theory and information entropy can be used to identify the importance of each indicator used in the present invention. Or, the importance of the indicator can be determined by mining the correlation relationship between the indicator and the output through data mining technology. Based on the importance of the indicator determined by the above methods, when tracing back the indicator, the contribution degree can be determined by querying the importance identifier corresponding to each indicator.

[0137] There are also some methods for determining feature importance different from the above. For example, when using a machine learning model of a neural network, since the weights of the hidden layer nodes in the neural network represent the importance of the features corresponding to the input layer nodes, when the machine learning model uses a neural network model, the weights of the hidden layer nodes are read, and the contribution degree of the indicator can be determined according to the weights of the hidden layer nodes.

[0138] For another example, for a machine learning model using a decision tree or random forest algorithm, according to the decision tree generation principle, the order of the features selected during the decision tree division process can be used as the importance ranking of the features, and this ranking can be obtained through the feature_importances_ attribute in sklearn. Therefore, when tracing back the indicators in the present invention, the importance of each indicator used in risk assessment, that is, the contribution degree, can be obtained by calling the feature_importances_ attribute.

[0139] After determining the contribution degree of the indicator, several indicators with the highest contribution degree ranking (such as the top three indicators) can be used as abnormal indicators. After determining the abnormal indicators, the content of the abnormal indicators is read, so as to locate the points where risks may occur and the specific risk content. Analyze whether there is a correlation relationship among the contents of the three abnormal indicators. If there is, the associated content is also used as risk content.

[0140] The present invention pre-sets risk events that may be generated for various risk contents and corresponding solutions. Therefore, after determining the risk content, the database is queried according to the risk content to determine the corresponding solution.

[0141] Taking the evaluation of enterprise risks as an example, if it is found after retrospective analysis that "enterprise legal litigation events" are the main cause of enterprise risks, the corresponding solutions provided by the system include: (1) During the interview, it is appropriate to ask the HR to verify the situation of enterprise labor dispute litigation; (2) Before the interview or before employment, please collect relevant information online to understand the details of relevant litigation and strengthen the comprehensive understanding of this issue, etc.

[0142] Add the above-mentioned risk content and the corresponding solutions to the warning information and present them to the user together. In addition, in order to enable the user to better understand the content identified this time or facilitate the user's future query, the present invention can also generate a risk report from the content in the warning information and store or provide it to the user. For example, generate a risk report including the evaluated risk level, risk content and corresponding solutions and send it to the user or store it in the user account.

[0143] Figure 8 FIG. is a schematic block diagram of a false recruitment identification system according to an embodiment of the present invention. In this embodiment, the false recruitment identification system includes a user information acquisition module 1, a data collection module 2, an index data generation module 3, and a risk assessment module 4. Among them, the user information acquisition module 1 is connected to the data collection module 2. The user information acquisition module 1 receives the recruitment interview information provided by the user and extracts the target recruitment enterprise information, target position information, and interview information from the user recruitment interview information. The data collection module 2 is connected to the Internet and is connected to the user information acquisition module 1. Based on the extracted target recruitment enterprise information, target position information, and interview information, it crawls information in the network according to multiple indexes such as enterprise level, position level, and interview level to obtain relevant information such as the recruitment enterprise and target position. The index data generation module 3 is connected to the data collection module 2, and processes the crawled enterprise level, position level, and interview level information to obtain multiple index data corresponding to each level. Among them, after being processed by the index data generation module 3, based on the obtained information at each level, it is standardized into specific index data, such as floating-point numbers, string vectors, or a certain encoding. The risk assessment module 4 is connected to the index data generation module 3 and is configured to evaluate according to multiple index data at each level to obtain the enterprise-level risk, position-level risk, and interview-level risk of false recruitment of the target enterprise.

[0144] Figure 9 FIG. is a partial schematic block diagram of a false recruitment identification system according to another embodiment of the present invention. In this embodiment, a machine learning model is used to evaluate risks. Therefore, in this embodiment, in addition to including Figure 8In addition to the modules described above, it further includes a prediction sample generation module 5, which is connected to the metric data generation module 3. Based on the samples required by machine learning models at all levels, each metric data is normalized into feature data, and the feature data corresponding to the metric data at all levels are combined together to form prediction samples for machine learning models at all levels. In this embodiment, the risk assessment module 4 includes an enterprise-level risk assessment unit 41, a position-level risk assessment unit 42, an interview-level risk assessment unit 43, a coding unit 44, and a risk level query unit 45. The prediction sample generation module 5 inputs the prediction samples at all levels to the enterprise-level risk assessment unit 41, the position-level risk assessment unit 42, and the interview-level risk assessment unit 43 respectively. The enterprise-level risk assessment unit 41 obtains the enterprise-level risk probability by inputting the enterprise-level prediction sample to the trained enterprise-level machine learning model. The position-level risk assessment unit 42 obtains the position-level risk probability by inputting the position-level prediction sample to the trained position-level machine learning model. The interview-level risk assessment unit 43 obtains the interview-level risk probability by inputting the interview-level prediction sample to the trained interview-level machine learning model. In one embodiment, the risk probabilities obtained by these three units can be directly provided to the user as the recognition result. In this embodiment, the risk probabilities obtained by these three units are output to the coding unit 44. The coding unit 44 encodes the obtained enterprise-level, position-level, and interview-level risk levels to obtain a risk code, and outputs the risk code to the risk level query unit 45. The risk level query unit 45 queries the correspondence table between the code and the risk level according to the risk code, such as Table 1 involved when the method was described above, and determines the risk level corresponding to the risk code as the final risk level. Of course, after obtaining the risk code, the coding unit 44 can also calculate the weighted sum of the risk codes according to the weights of risks at all levels. The risk level query unit 45 queries the correspondence table between the weighted sum of the codes and the risk level according to the weighted sum of the risk codes, such as Table 2 involved when the method was described above, and determines the risk level corresponding to the weighted sum of the risk codes as the final risk level.

[0145] Figure 10 is a principle block diagram of a false recruitment recognition system provided according to another embodiment of the present invention. This embodiment is the same as Figure 9Compared with the embodiments shown, the risk assessment module 4 in this embodiment further includes a selection unit 46 in addition to the enterprise-level risk assessment unit 41, the position-level risk assessment unit 42, and the interview-level risk assessment unit 43. In this embodiment, the machine learning models used by the three assessment units are related to each other. Among them, in the order from top to bottom of the enterprise level, the position level, and the interview level, the lower-level machine learning model is trained according to the training data corresponding to the risk level output by its upper-level machine learning model. Therefore, the number of lower-level machine learning models is the same as the number of risk levels of the upper-level machine learning model, and the risk assessment is also carried out step by step from top to bottom, and it is necessary to select the lower-level machine learning model according to the risk level of the upper-level machine learning model.

[0146] Specifically, the enterprise-level risk assessment unit 41 receives the enterprise-level prediction samples obtained by the prediction sample generation module 5 and notifies the selection unit 46. The selection unit 46 selects an enterprise-level machine learning model from the model library and sends it to the enterprise-level risk assessment unit 41. The enterprise-level risk assessment unit 41 inputs the enterprise-level prediction samples into the enterprise-level machine learning model, and obtains the enterprise-level risk level through the evaluation of the enterprise-level machine learning model. While sending the enterprise-level risk level to the selection unit 46, it notifies the position-level risk assessment unit 42. The selection unit 46 selects a suitable position-level machine learning model according to the enterprise-level risk level and sends it to the position-level risk assessment unit 42. After receiving the notification from the enterprise-level risk assessment unit 41 and the position-level machine learning model sent by the selection unit 46, the position-level risk assessment unit 42 inputs the position-level prediction samples received from the prediction sample generation module 5 into the position-level machine learning model, obtains the position-level risk level through the evaluation of the position-level machine learning model, outputs the position-level risk level to the selection unit 46, and simultaneously outputs a notification to the interview-level risk assessment unit 43. The selection unit 46 selects the corresponding interview-level machine learning model according to the position-level risk level and sends it to the interview-level risk assessment unit 43. After receiving the notification from the position-level risk assessment unit 42 and the interview-level machine learning model sent by the selection unit 46, the interview-level risk assessment unit 43 inputs the interview-level prediction samples received from the prediction sample generation module 5 into the interview-level machine learning model, and obtains the final risk level through the evaluation of the interview-level machine learning model.

[0147] The outputs of the enterprise-level machine learning model and the position-level machine learning model in the foregoing embodiments are respectively binary classifications, which are respectively defined as two risk levels of "at risk" and "not at risk". The output of the interview-level machine learning model is binary classification or multi-classification. When it is binary classification, it is respectively defined as two risk levels of "at risk" and "not at risk". When it is multi-classification, it is respectively defined as multiple risk levels of different degrees from none to some.

[0148] Figure 11 It is a schematic block diagram of a user information acquisition module according to an embodiment of the present invention. In this embodiment, the user information acquisition module includes a message transceiver unit 11 and an information extraction unit 12. The message transceiver unit 11 serves as an interaction interface between the system and the user. On the one hand, it is connected to the risk assessment module 4 and sends the final recognition result to the user. On the other hand, it receives the recruitment interview information provided by the user and outputs it to the information extraction unit 12. The information extraction unit 12 extracts the target recruitment enterprise information, target position information, and interview information from the recruitment interview information provided by the user.

[0149] The message transceiver unit 11 includes one or more of the following units: an application terminal user interaction unit 110, an email processing unit 111, a mobile short message processing unit 112, and a social media message processing unit 113. The application terminal user interaction unit 110 includes at least an input interface through which the recruitment interview information provided by the user via the application terminal can be obtained. In addition, the application terminal user interaction unit 110 may also include a display interface for displaying messages such as the identified risk level, risk content, or corresponding solutions. The email processing unit 111 can identify the recruitment interview information provided by the user via email according to the email address or subject, and can also send messages to the user via email. The mobile short message processing unit 112 identifies the recruitment interview information provided by the user via short message according to the message sending number and subject in the short message received based on the mobile communication network, and can also send messages to the user. The social media message processing unit 113 identifies the recruitment interview information provided by the user via social media from the social media messages or sends messages to the user according to the message sender.

[0150] Figure 12 It is a schematic block diagram of a data processing system according to an embodiment of the present invention. In this embodiment, the data processing system includes a user information acquisition module 1, a data collection module 2, an index data generation module 3, and a prediction sample generation module 5, that is Figure 9 or Figure 10 Some modules in the false recruitment identification system in constitute a data processing system, which performs information extraction, information crawling, data normalization, etc. based on the recruitment interview information provided by the user, so as to obtain prediction samples for evaluating risks by a machine learning model. Specifically, Figure 13 It is a schematic block diagram of a data collection module according to an embodiment of the present invention. In this embodiment, the data collection module includes an index acquisition unit 21, an index analysis unit 22, an information crawling unit 23, and an information matrix construction unit 24.

[0151] The index acquisition unit 21 is configured to read a plurality of indexes applied to the enterprise level, the position level, and the interview level respectively. In this embodiment, indexes for evaluating risks at each level are stored in the system. The index acquisition unit 21 reads these indexes from the system database and sends them to the index analysis unit 22. The index analysis unit 22 analyzes each index, determines the index reference content required to obtain the index data, and sends the index reference content to the information crawling unit 23. The information crawling unit 23 is connected to the index analysis unit 22 and crawls the corresponding information from the Internet according to the determined index reference content. In one embodiment, one or more retrieval keywords corresponding to the indexes are stored in the system. For example, when the index acquisition unit 21 reads the index of "number of enterprise employees", the index analysis unit 22 queries its retrieval keyword to obtain the index reference content of "number of enterprise employees / quantity", and the information crawling unit 23 searches and obtains the information of the number of enterprise employees on the official website of the target recruitment enterprise according to the index reference content of "number of enterprise employees / quantity", so as to obtain the information that meets the index of "number of enterprise employees". Another example is that when the index acquisition unit 21 reads the enterprise-level index of "enterprise legal litigation events", the index analysis unit 22 queries its retrieval keyword to obtain the index reference content of "labor / defendant", etc., and the information crawling unit 23 queries the relevant litigation information of the target recruitment enterprise in the recent period (such as five years) according to these contents.

[0152] In this embodiment, the information matrix construction unit 24 is connected to the information crawling unit 23 and establishes an information matrix according to the association between the crawled information. The information matrix includes one or more position information released by the target enterprise and the recruitment platform qualification information for releasing these position information; one or more position information released by one or more peer enterprises of the target enterprise and the recruitment platform qualification information for releasing these position information. Among them, the positions released by the target enterprise include the target position and a second position different from the target position; the recruitment platforms include one or more target recruitment platforms for releasing the target positions provided by the target enterprise and one or more second recruitment platforms for releasing the second position information different from the target position.

[0153] Figure 14Principle block diagram of the indicator data generation module according to an embodiment of the present invention. In this embodiment, the indicator data generation module includes a data cleaning unit 31, a single indicator extraction unit 32, and a composite indicator calculation unit 33. Among them, the data cleaning unit 31 performs data cleaning on the crawled original information, including removing some network symbols, punctuation marks, querying the stop word list to remove stop words, etc. The single indicator extraction unit 32 is connected to the data cleaning unit 31, and extracts and normalizes single indicator data from the cleaned data. The single indicators are, for example, some indicators that do not require complex calculations and processing, such as "number of employees in the enterprise", "number of branch companies", "registered capital", "job title", "educational requirements", etc. For these indicators, the indicator data can be extracted by identifying keywords, and according to the specific indicators, they can be processed into string vectors, floating-point values, or codes. The composite indicator calculation unit 33 is connected to the single indicator extraction unit 32 and is configured to calculate more than one single indicator data according to the composite indicator calculation rule to obtain composite indicator data. The composite indicators are, for example, indicators such as "investment degree of the recruitment platform", "qualification level of the recruitment platform", "external consistency coefficient of the target position", etc. The calculation method is as described in the foregoing description of the method part and will not be elaborated here.

[0154] Figure 15It is a principle block diagram of a model training module in a data processing system according to another embodiment of the present invention. In this embodiment, the model training module 6 includes a training data set unit 61, a model training unit 62, a user feedback monitoring unit 63, a sample annotation unit 64, and a model update unit 65. The training data set unit 61 is used to provide a data set for training a model, and includes an enterprise-level training data subset, a position-level training data subset, and an interview-level training data subset according to the type of the trained model. Further, each training data subset further includes a training subset and a validation subset. The model training unit 62 trains a model with the data in the corresponding type of training data subset according to the type of the trained model, and respectively obtains an enterprise-level machine learning model, a position-level machine learning model, and an interview-level machine learning model. In one embodiment, the model training unit 62 trains the model with mutually independent training sets respectively, so as to obtain three independent machine learning models. In another embodiment, according to the risk level of the enterprise-level machine learning model, the position-level training data subset respectively includes different subsets composed of training data with corresponding enterprise risks, and different position-level machine learning models are obtained according to different training subsets. Similarly, corresponding to the interview-level training data subset, which includes multiple training subsets composed of training data with specific corresponding enterprise risk levels and position-level risk levels, different interview-level machine learning models are obtained according to different training subsets. As an example involved in the description of the method of the present invention as described above, an enterprise-level machine learning model M1, position-level machine learning models M2.1, M, and interview-level machine learning models M3.1, M3.2, M3.3, and M3.4 are obtained through different training subsets. In order to expand the training set, the user feedback monitoring unit 63 monitors the feedback information of the user on the false recruitment recognition result, and the feedback information at least includes confirmation information of being risky or risk-free. The sample annotation unit 64 is connected to the user feedback monitoring unit 63, and based on the user's feedback information, risks are annotated for the data for obtaining the recognition result and added to the corresponding training data subset, so as to achieve the purpose of enriching the training data. The model update unit 65 is connected to the training data set unit 61, and is used to monitor the update condition, and send an update notice to the model training unit 62 when the model update condition is met. The model training unit 62 trains, optimizes, and updates the model with the training data. Among them, the update condition is, for example, reaching a preset update period, such as updating the model once a week / month, or the number of newly added training samples reaching a threshold. Therefore, the model update unit 65 times after each model update, and sends an update notice to the model training unit 62 when the timing period is reached. Or the model update unit 65 counts the newly added training data in the training data set, and when the number of newly added training data reaches the threshold, notifies the model training unit to optimize and update the original machine learning model with the current training data.

[0155] Figure 16 It is a block diagram of the principle of a false recruitment warning system according to an embodiment of the present invention. In this embodiment, the warning system includes a user information acquisition module 1, a data collection module 2, an index data generation module 3, a risk assessment module 4, and a warning module 7. Among them, the user information acquisition module 1, the data collection module 2, the index data generation module 3, and the risk assessment module 4 are the same as the modules in the false recruitment identification system in the foregoing embodiment, and will not be elaborated here.

[0156] The warning module 7 is connected to the risk assessment module 4, and in response to the evaluated risk, traces back and analyzes the index data corresponding to the risk to obtain warning information and provides it to the user.

[0157] Figure 17 It is a block diagram of the principle of the warning module according to an embodiment of the present invention. In this embodiment, the warning module includes a risk monitoring unit 71, a risk content determination unit 72, a warning information generation unit 73, and a warning information sending unit 74.

[0158] The risk monitoring unit 71 is connected to the risk assessment module 4 to monitor whether the risk assessment module 4 evaluates a risk. When it detects that the risk assessment module 4 evaluates a risk, it sends a notification to the risk content determination unit 72. The risk content determination unit 72 traverses the index data for evaluating the risk based on the evaluated risk type to obtain abnormal index data, and determines the risk content based on the content or its association relationship of the one or more abnormal index data. Among them, any one of the foregoing methods can be used to determine the abnormal index data, and will not be elaborated here. The warning information generation unit 73 is connected to the risk content determination unit 72 to generate warning information based on the risk type and the risk content. The warning information sending unit 74 is connected to the warning information generation unit 73 and is used to provide the warning information to the user. Among them, the warning information sending unit 74 can adopt one or more of an interactive information sending unit, an email processing unit, a mobile short message processing unit, a mobile short message processing unit, and a social media message processing unit. That is to say, the warning information sending unit 74 can be combined with the user information acquisition module into one module, and its module is implemented as a two-way function with both message receiving and sending functions. For specific reference, see the foregoing embodiment, and it will not be elaborated here.

[0159] Figure 18 It is a block diagram of the principle of the warning module according to another embodiment of the present invention. In this embodiment, Figure 17Compared with the embodiment of [], except for adding the prediction unit 75, the functions of other modules are the same and will not be elaborated here. The prediction unit 75 is connected to the risk content determination unit 72 and is configured to predict risk events and risk response plans according to the risk content, and add the risk events and their response plans to the warning information. For example, the prediction unit 75 queries various suggestions and corresponding plans preset in the system database according to the risk content, so as to obtain the matching suggestions, and combines multiple suggestions or corresponding plans and adds them to the warning information. Another example is that various possible risk events and corresponding response plans corresponding to the risk content are preset in the system database, and these information can be added to the warning information and sent to the user together.

[0160] Figure 19 is a principle block diagram of the warning module according to another embodiment of the present invention. In this embodiment, compared with Figure 18 the embodiment of [], except for adding the risk report generation unit 76, the functions of other modules are the same and will not be elaborated here. The risk report generation unit 76 is connected to the warning information generation unit 73 and is configured to generate a risk report according to the content in the warning information. Correspondingly, the warning information sending unit stores the risk report in a predetermined position or sends it to the user.

[0161] The above embodiments are only for illustrating the present invention and are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can make various changes and modifications without departing from the scope of the present invention. Therefore, all equivalent technical solutions should also fall within the scope of the disclosure of the present invention.

Claims

1. A method for identifying false job recruitment, characterized in that, Including: Extracting target recruitment enterprise information, target position information, and interview information from the recruitment interview information provided by the user; Based on the extracted target recruitment enterprise information, target position information, and interview information, crawling information in the network according to multiple indicators among enterprise level, position level, and interview level respectively, and processing the crawled information into index data of corresponding levels to obtain enterprise level index data, position level index data, and interview level index data; Inputting the target recruitment enterprise information and the crawled enterprise level index data into a trained enterprise level machine learning model, and the model outputs an enterprise level risk assessment result for the target recruitment enterprise; Among them, the enterprise level risk assessment result includes: the target recruitment enterprise belongs to an enterprise with enterprise level risks, or the target recruitment enterprise belongs to an enterprise without enterprise level risks; If the target recruitment enterprise belongs to an enterprise without enterprise level risks, inputting the information of the target recruitment enterprise and the crawled position level index data into a trained position level machine learning model M2.1, and the model outputs a position level risk assessment result for the target recruitment enterprise; If the target recruitment enterprise belongs to an enterprise with enterprise level risks, inputting the information of the target recruitment enterprise and the crawled position level index data into a trained position level machine learning model M2.2, and the model outputs a position level risk assessment result for the target recruitment enterprise; Among them, the position level risk assessment result includes: the target recruitment enterprise belongs to an enterprise with position level risks, or the target recruitment enterprise belongs to an enterprise without position level risks; If the target recruitment enterprise belongs to an enterprise without enterprise level risks and without position level risks, inputting the information of the target recruitment enterprise and the crawled interview level index data into a trained interview level machine learning model M3.1, and the model outputs an interview level risk assessment result for the target recruitment enterprise; If the target recruitment enterprise belongs to an enterprise without enterprise level risks but with position level risks, inputting the information of the target recruitment enterprise and the crawled interview level index data into a trained interview level machine learning model M3.2, and the model outputs an interview level risk assessment result for the target recruitment enterprise; If the target recruitment enterprise belongs to an enterprise with enterprise level risks but without position level risks, inputting the information of the target recruitment enterprise and the crawled interview level index data into a trained interview level machine learning model M3.3, and the model outputs an interview level risk assessment result for the target recruitment enterprise; If the target recruitment enterprise belongs to an enterprise with enterprise level risks and with position level risks, inputting the information of the target recruitment enterprise and the crawled interview level index data into a trained interview level machine learning model M3.4, and the model outputs an interview level risk assessment result for the target recruitment enterprise; Among them, the interview level risk assessment result includes: the target recruitment enterprise belongs to an enterprise with interview level risks, or the target recruitment enterprise belongs to an enterprise without interview level risks.

2. The method according to claim 1, characterized in that, The risk and level output by the interview level machine learning model are the final risk and level.

3. The method according to claim 2, wherein Further including: Encoding the obtained enterprise level, position level, and interview level risk levels to obtain risk codes; Query the correspondence table between risk codes and risk levels according to the said risk code; And Determine the risk level corresponding to the said risk code as the final risk level.

4. The method according to claim 3, wherein Further include: Provide the determined risk type and risk level to the user as the recognition result, and the providing methods include displaying on the user interaction interface, sending an email, sending a short message through the mobile communication network, or sending a social message to the user's social media account.

5. The method according to claim 1, characterized in that, The enterprise-level index data obtained from the crawled enterprise-level information includes one or more static index data among the number of employees, number of branch companies, registered capital, financing information, enterprise age, annual turnover, business scope, and enterprise type of the target recruitment enterprise.

6. The method according to claim 1, wherein The enterprise-level index data obtained from the crawled enterprise-level information further includes one or more dynamic index data among the enterprise legal litigation events, major investments received, social news events occurred, and social media enterprise evaluations of the target recruitment enterprise.

7. The method according to claim 1, wherein The said position-level index data includes the following in-depth index data obtained from the position information: one or more of position name, work location, work department, monthly salary range, job description, welfare benefits, work type, experience requirements, educational requirements, school type, work industry risk, and number of repeated recruitments for the same position.

8. The method according to claim 1, characterized in that, The said interview-level index data includes one or more of interview location, interview time, interview notification method, number of interviews conducted, and interview form.

9. A false recruitment identification system, characterized in that, Include: A user information acquisition module, configured to receive the recruitment interview information provided by the user, and extract the target recruitment enterprise information, target position information, and interview information from the said user recruitment interview information; A data collection module, which is connected to the Internet and connected to the said user information acquisition module, and is configured to crawl information on the network respectively according to multiple indexes among enterprise level, position level, and interview level based on the extracted target recruitment enterprise information, target position information, and interview information; An index data generation module, which is connected to the said data collection module, processes the crawled enterprise-level information, position-level information, and interview-level information to obtain multiple index data of the corresponding levels, and obtains enterprise-level index data, position-level index data, and interview-level index data; And An enterprise-level machine learning model, configured to receive the target recruitment enterprise information and the crawled enterprise-level index data, and output the enterprise-level risk assessment result of the target recruitment enterprise; Among them, the enterprise-level risk assessment result includes: the target recruitment enterprise belongs to an enterprise-level risky enterprise, or the target recruitment enterprise belongs to an enterprise-level risk-free enterprise; A position-level machine learning model M2.1, configured to receive the information of the enterprise-level risk-free target recruitment enterprise and the crawled position-level index data, and output the position-level risk assessment result of the said enterprise-level risk-free target recruitment enterprise; A position-level machine learning model M2.2, configured to receive the information of the enterprise-level risky target recruitment enterprise and the crawled position-level index data, and output the position-level risk assessment result of the said enterprise-level risky target recruitment enterprise; Among them, the position-level risk assessment results include: the target recruitment enterprise belongs to an enterprise with position-level risks, or the target recruitment enterprise belongs to an enterprise without position-level risks; The interview-level machine learning model M3.1 is configured to receive information of a target recruitment enterprise that is risk-free at the enterprise level and risk-free at the occupational level, as well as the crawled interview-level indicator data, and output the interview-level risk assessment result of the target recruitment enterprise that is risk-free at the enterprise level and risk-free at the occupational level; The interview-level machine learning model M3.2 is configured to receive information of a target recruitment enterprise that is risk-free at the enterprise level and has position-level risks, as well as the crawled interview-level indicator data, and output the interview-level risk assessment result of the target recruitment enterprise that is risk-free at the enterprise level and has position-level risks; The interview-level machine learning model M3.3 is configured to receive information of a target recruitment enterprise that is risky at the enterprise level and risk-free at the position level, as well as the crawled interview-level indicator data, and output the interview-level risk assessment result of the target recruitment enterprise that is risky at the enterprise level and risk-free at the position level; The interview-level machine learning model M3.4 is configured to receive information of a target recruitment enterprise that is risky at the enterprise level and has position-level risks, as well as the crawled interview-level indicator data, and output the interview-level risk assessment result of the target recruitment enterprise that is risky at the enterprise level and has position-level risks; Among them, the interview-level risk assessment results include: the target recruitment enterprise belongs to an enterprise with interview-level risks, or the target recruitment enterprise belongs to an enterprise without interview-level risks.

Citation Information

Patent Citations

  • False recruitment information detection method based on cascade forest

    CN113704409A

  • False recruitment position detection method based on deep learning

    CN113506084A

  • Recruitment method and system based on block chain

    CN113673234A