A false recruitment warning method and system

By extracting data from the recruitment interview information provided by users, multi-dimensional risk assessments of enterprises, positions and interview levels are carried out, and the problems of narrow data sources and single identification dimensions in the existing technology are solved, achieving a more accurate and comprehensive false recruitment warning.

CN114612062BActive Publication Date: 2025-07-29QIAN JIN NETWORK INFORMATION TECH SHANGHAI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210222605.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-09
Publication Date
2025-07-29
Estimated Expiration
2042-03-09

AI Technical Summary

Technical Problem

When identifying false recruitment information, the data source range is narrow and the recognition dimension is single, resulting in a large number of false recruitment being missed and the risk is not effectively detected during the interview notification process.

Method used

Extract company, position and interview information from the recruitment interview information provided by users, crawl data from the network through multiple indicators, conduct risk assessments at the enterprise level, position level and interview level, and provide early warning information.

Benefits of technology

It increases the accuracy of judgment of false recruitment, makes up for the lack of traditional algorithms, and can provide the most complete data analysis in the interview notification process, accurately locate and identify risk sources, and provide more comprehensive risk assessment conclusions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114612062B_ABST
    Figure CN114612062B_ABST
Patent Text Reader

Abstract

The present invention relates to a false recruitment warning method and system, wherein the method comprises: extracting target recruitment enterprise information, target position information and interview information from the recruitment interview information provided by a user; based on the extracted target recruitment enterprise information, target position information and interview information, respectively performing information crawling in the network according to multiple indicators in enterprise level, position level and interview level, and processing the crawled information into corresponding index data; according to a risk assessment strategy, respectively performing risk assessment from enterprise level, position level and interview level according to the index data of each level; and in response to the assessed risk, tracing back and analyzing the index data corresponding to the risk to obtain warning information and providing the warning information to the user. The present invention identifies false recruitment from multiple aspects and dimensions and timely warns the user to avoid the user being deceived.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a false recruitment warning method and system. Background Art

[0002] False recruitment generally refers to fraudulent recruitment where the recruitment information posted on the Internet or in the talent market does not match the actual employment situation. It can refer to either false recruitment carried out by a regular company or illegal recruitment by a false company. Some regular companies create some fake positions for reasons such as collecting talent information and promoting the company's popularity, and post recruitment information on the Internet and / or in the talent market. Such false recruitment behavior not only delays the job-seeking opportunities of job seekers, endangers personal information security, but also occupies and wastes the job-seeking public platform, damaging the credibility of the recruitment platform. Illegal recruitment by false companies can not only cause consequences such as property and personal damage to job seekers, such as charging arbitrary fees, pyramid schemes, and involvement in porn, drugs, etc., but may also infringe on the interests of third parties, such as regular companies whose names are misappropriated and recruitment platforms that post recruitment information for them.

[0003] In order to identify such false recruitment information, the industry has made a great deal of efforts. The Chinese patent application with the publication number CN113704409A and the invention name "A false recruitment information detection method based on cascade forest" provides a false recruitment information detection method. Based on the cascade forest algorithm of decision trees, a model is established using the job data posted on the online recruitment platform for false recruitment prediction. The Chinese patent application with the publication number CN113506084A and the invention name "A false recruitment position detection method based on deep learning" provides a false recruitment position detection method. False recruitment information is collected from online recruitment platforms or recruitment APPs, and the false recruitment information is processed into sample data for training the model. A detection model is obtained through training, and the detection model is used to detect false recruitment information on the online recruitment platform or recruitment APP.

[0004] From the technical solutions provided above, it can be seen that currently, when identifying false recruitment information, the training data for training the identification model comes from false job information and / or false recruitment information collected from online recruitment platforms. The source range of the data is narrow and the identification dimension is single. According to the foregoing description of false recruitment, fraudulent false recruitment has multiple aspects. If only limited to identifying from the job aspect, a large number of false recruitments are bound to be missed. Therefore, there is an urgent need for a solution to identify false recruitment from multiple aspects and dimensions. Summary of the Invention

[0005] Aiming at the technical problems existing in the prior art, the present invention proposes a false recruitment warning method and system for identifying false recruitment from multiple aspects and dimensions and timely warning users to avoid users being deceived.

[0006] To solve the above technical problems, according to one aspect of the present invention, the present invention provides a method for false recruitment warning, including the following steps:

[0007] Extract target recruitment enterprise information, target position information, and interview information from the recruitment interview information provided by the user;

[0008] Based on the extracted target recruitment enterprise information, target position information, and interview information, crawl information in the network respectively according to multiple indicators in enterprise level, position level, and interview level, and process the crawled information into corresponding indicator data;

[0009] According to the risk assessment strategy, conduct risk assessment respectively from enterprise level, position level, and interview level according to the indicator data of each level; and

[0010] In response to the evaluated risk, trace back and analyze the indicator data corresponding to the risk to obtain warning information and provide it to the user.

[0011] According to one aspect of the present invention, the present invention provides a false recruitment warning system. The system includes a user information acquisition module, a data collection module, an indicator data generation module, a risk assessment module, and a warning module. Among them, the user information acquisition module is configured to receive the recruitment interview information provided by the user, and extract target recruitment enterprise information, target position information, and interview information from the user recruitment interview information; the data collection module is connected to the Internet and is connected to the user information acquisition module, and is configured to crawl information in the network respectively according to multiple indicators in enterprise level, position level, and interview level based on the extracted target recruitment enterprise information, target position information, and interview information; the indicator data generation module is connected to the data collection module, and processes the crawled enterprise level, position level, and interview level information to obtain multiple indicator data corresponding to each level; the risk assessment module is connected to the indicator data generation module, and is configured to conduct risk assessment respectively from enterprise level, position level, and interview level according to the indicator data of each level according to the risk assessment strategy; the warning module is connected to the risk assessment module, and is configured to, in response to the evaluated risk, trace back and analyze the indicator data corresponding to the risk to obtain warning information and provide it to the user.

[0012] As can be seen from the method and system provided by the present invention above, the present invention makes up for the lack of the predecessors in actively searching for information and risk discrimination: traditional algorithms only use information related to job descriptions to judge whether a recruitment is fake, while the present invention actively searches for other relevant information, such as enterprise information, related enterprise information, target job information, related job information, etc., thus increasing the accuracy of judgment. The present invention makes up for the lack of the predecessors in predicting risks in the interview notice link: existing fake recruitment detection schemes are earlier than the interview notice link and cannot effectively detect fake recruitments that show risks only at the interview notice stage. However, after a job seeker obtains an interview notice from the target enterprise, the present invention analyzes and evaluates based on the interview information given by the job seeker. This timing is not only the last line of defense before the job seeker goes to the interview but also the best timing to obtain the most sufficient data to predict risks, thereby being able to help job seekers avoid risks in a targeted manner to the greatest extent. The present invention also makes up for the lack of the predecessors in accurately positioning and identifying the risk sources: based on the relevant data of this interview, after evaluating the risks, the present invention can trace back upward to explore problems at the overall enterprise level and job level, sort out the essential reasons for the existence of risks, obtain the full picture of possible recruitment fraud problems behind the risks, and obtain a more general evaluation conclusion and a more detailed analysis report. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Next, the preferred embodiments of the present invention will be further described in detail with reference to the drawings, where:

[0014] Figure 1 is a flowchart of a method for identifying fake recruitment according to an embodiment of the present invention;

[0015] Figure 2 is a flowchart of risk assessment according to an embodiment of the present invention;

[0016] Figure 3 is a flowchart of risk assessment according to another embodiment of the present invention;

[0017] Figure 4 is a flowchart of risk assessment according to yet another embodiment of the present invention;

[0018] Figure 5 is a flowchart of assessing risks according to another embodiment of the present invention;

[0019] Figure 6 is a flowchart of a method for processing test sample data according to an embodiment of the present invention;

[0020] Figure 7 is a flowchart of a method for warning against fake recruitment according to an embodiment of the present invention;

[0021] Figure 8It is a block diagram of the principle of a false recruitment identification system provided according to an embodiment of the present invention;

[0022] Figure 9 It is a partial block diagram of the principle of a false recruitment identification system provided according to an embodiment of the present invention;

[0023] Figure 10 It is a partial block diagram of the principle of a false recruitment identification system provided according to an embodiment of the present invention;

[0024] Figure 11 It is a block diagram of the principle of a user information acquisition module according to an embodiment of the present invention;

[0025] Figure 12 It is a block diagram of the principle of a data processing system according to an embodiment of the present invention;

[0026] Figure 13 It is a block diagram of the principle of a data collection module according to an embodiment of the present invention;

[0027] Figure 14 It is a block diagram of the principle of an index data generation module according to an embodiment of the present invention;

[0028] Figure 15 It is a block diagram of the principle of a model training module according to an embodiment of the present invention;

[0029] Figure 16 It is a block diagram of the principle of a false recruitment warning system according to an embodiment of the present invention;

[0030] Figure 17 It is a block diagram of the principle of a warning module according to an embodiment of the present invention;

[0031] Figure 18 It is a block diagram of the principle of a warning module according to another embodiment of the present invention; and

[0032] Figure 19 It is a block diagram of the principle of a warning module according to yet another embodiment of the present invention. Detailed implementation manners

[0033] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] In the following detailed description, reference is made to the various specification drawings that form a part of the present application and illustrate specific embodiments of the present application. In the drawings, like reference numerals generally describe substantially similar components in different figures. The specific embodiments of the present application are described in sufficient detail below so that those of ordinary skill in the relevant art and technology can implement the technical solutions of the present application. It should be understood that other embodiments may also be utilized, or structural, logical, or electrical changes may be made to the embodiments of the present application.

[0035] Figure 1 It is a flowchart of a method for identifying false job recruitments according to an embodiment of the present invention. In this embodiment, the method includes the following steps:

[0036] Step S1a: Determine whether recruitment interview information provided by the user is received. If the recruitment interview information provided by the user is received, then execute step S2a; if not, repeat this step. In one embodiment, when the method is applied to an online recruitment platform, such as a recruitment website, a recruitment APP, etc., the online recruitment platform opens an interface to the user to facilitate the user to input recruitment interview information through this interface. The interface can be implemented, for example, as a user interface through which the user inputs recruitment interview information. In some embodiments, the user interface is also provided with multiple information segments for inputting target recruitment enterprise information, target position information, and interview information respectively. The target recruitment enterprise information includes, for example, information such as the enterprise name and address. The target position information includes, for example, information such as the position name and the work location. The interview information includes, for example, information such as the interview location, interview time, interview notification method, and interview form, such as oral examination, written examination, etc.

[0037] Step S2a: Extract the target recruitment enterprise information, target position information, and interview information from the recruitment interview information provided by the user. In one embodiment, when the user inputs various information through the information segments in the user interface, the information in the fields is read from the information segments and processed to obtain one or more keywords. The processing includes, for example, stop word removal, word segmentation, etc.

[0038] Step S3a: Based on the extracted target recruitment enterprise information, target position information, and interview information, perform information crawling in the network according to multiple indicators among enterprise level, position level, and interview level respectively.

[0039] Among them, in one embodiment, enterprise-level indicators include static indicators and dynamic indicators. The static indicators include, but are not limited to, one or more of the number of employees, number of branch companies, registered capital, financing information, enterprise age, annual turnover, business scope, and enterprise type of the target recruitment enterprise. The dynamic indicators include, but are not limited to, one or more of the enterprise legal litigation events, major investments received, social news events occurred, and enterprise evaluations on social media of the target recruitment enterprise. The position-level indicators include, but are not limited to, one or more of the position name, work location, work department, monthly salary range, job description, welfare benefits, work type, experience requirements, educational requirements, school type, work industry risk, and number of repeated recruitments for the same position. The interview-level indicators include, but are not limited to, one or more of the interview location, interview time, interview notice method, number of interviews conducted, whether there is a written test, and interview form. Corresponding information is obtained from the network according to these indicators. For example, according to the information such as the name and address of the target recruitment enterprise provided by the user, the official website, social media public account, and public news reports of the target recruitment enterprise can be searched in the network using its name and / or address as keywords.

[0040] In one embodiment, enterprise number information, information including "xx branch company", business scope description information, information including enterprise category, financing-related information, description information about enterprise age, description information about annual turnover, and information about registered capital are searched and obtained on the official website of the target recruitment enterprise. If a certain piece of the above information is not obtained from the official website of the target recruitment enterprise, continue to search in the network. For example, if the registered capital, enterprise type, etc. are not obtained from the official website of the target recruitment enterprise, the registration information of the target recruitment enterprise can be obtained from the industrial and commercial information network, and the registered capital and enterprise type can be obtained therefrom. Another example is that if the financing information is not obtained from the official website of the target recruitment enterprise, the financing-related information can be obtained by querying the social media public account or public news reports of the target recruitment enterprise. Comment information, litigation information, news information about received investments, etc. of the target recruitment enterprise in the recent five years are identified on the web page. Among them, in order to obtain the data of the position-level indicators, the position information of the same position and other positions published by the target recruitment enterprise is searched on multiple recruitment platforms, and at the same time, the position information of the same position and other positions published by enterprises of the same type as the target recruitment enterprise is also searched. The information related to the position is obtained from the official website of the target recruitment enterprise. The job posting information released by the target recruitment enterprise within a certain period of time is obtained from the network.

[0041] Step S4a: Process the crawled information into corresponding-level indicator data. Process the crawled corresponding information according to the processing strategies of specific indicators. Standardize the corresponding information into specific indicator data according to various indicators. The indicator data can be floating-point numbers, string vectors, or a certain encoding.

[0042] For example, the processing of enterprise-level static indicators includes: Regarding the number of employees in the enterprise, process the crawled data of the number of employees of the target recruitment enterprise into a floating-point value. Regarding the number of branch companies, calculate the number according to the words containing "**Branch Company" identified from the company introduction on the official website and process it into a floating-point value; Regarding the registered capital, process the value in the words containing "Registered Capital**" identified from the company registration information introduction into a floating-point value. Regarding the business scope, identify the keywords in the business scope description and process them into a string vector. Regarding the enterprise type, process the enterprise category information on the official website into a string vector, match it among "state-owned enterprise", "foreign-funded enterprise", "joint venture", and "private enterprise", and process the above options into 3, 2, 1, and 0 respectively and record them. Regarding the financing information, process the net value in the latest financial report of the enterprise obtained into a floating-point value. Regarding the enterprise age, subtract the establishment time of the enterprise obtained from the current year to obtain the number of years of establishment and process it into a floating-point value. Regarding the annual turnover, process the annual turnover obtained from the latest financial report of the enterprise into a floating-point value.

[0043] For the processing of enterprise-level dynamic indicators: Regarding the enterprise legal litigation events, first extract the litigation information of the target recruitment enterprise in the recent period (such as five years) from the obtained information; then use keywords such as "labor service" and "defendant**(name of the target recruitment enterprise)" to perform regular matching on each piece of information. If the match is successful, record it as 1, and if the match fails, record it as 0; when the match is successful, obtain the litigation time information from the litigation information and encode it. For example, when the litigation time is less than 1 year, encode it as 1; when the litigation time is less than 2 years and greater than 1 year, encode it as 0.8; when the litigation time is less than 3 years and greater than 2 years, encode it as 0.6; when the litigation time is less than 4 years and greater than 3 years, encode it as 0.5; when the litigation time is less than 5 years and greater than 4 years, encode it as 0.3; then multiply all the encoded values of the successful matches by the encoded value of the litigation time to obtain the litigation code; finally, add up all the litigation codes to obtain the enterprise's litigation event score.

[0044] Regarding the major investments received by the enterprise, first extract the news information related to received investments in the recent period (such as five years) from the crawled information; then process the amount of received investments into a floating-point value; then add up all the floating-point values to obtain the total amount of major investments received by the target recruitment enterprise in the recent period (such as five years); if not, the default value is 0.

[0045] Regarding the evaluation of social media enterprises, first extract multiple comment information from the captured information for a recent period (such as five years); then process each comment information according to the following process: remove HTML symbols in the comment through a parsing method; remove punctuation marks through regular expressions; use a stop word library to filter and remove all stop words; tokenize the data text after the above cleaning; use a sentiment analysis model to analyze the text, and output each comment as a sentiment encoding of -1 (negative) or 1 (positive) for the comment text; then obtain the comment time information. When the comment time is less than 1 year from the current, it is encoded as 1. When the comment time is less than 2 years and greater than 1 year from the current, it is encoded as 0.8. When the comment time is less than 3 years and greater than 2 years from the current, it is encoded as 0.6; when the comment time is less than 4 years and greater than 3 years from the current, it is encoded as 0.5; when the comment time is less than 5 years and greater than 4 years from the current, it is encoded as 0.3; then multiply the sentiment encoding of all comment texts by the time encoding to obtain the sentiment score of each comment. For example, if an enterprise has received a negative comment, and the time when it appears is less than 4 years and greater than 3 years from the current, then the sentiment score of this comment is -1 * 0.5 = -0.5; the enterprise has also received a positive comment, and the time when it appears is less than 3 years and greater than 2 years from the current, then the sentiment score of this comment is 1 * 0.6 = 0.6; then add up the sentiment scores of all comments of the enterprise to obtain the social media evaluation score of the enterprise. The social media evaluation score of the enterprise in the above example is -0.5 + 0.6 = 0.1.

[0046] The processing of indicator data at the position level includes: for the position name, process it into a string vector; for the work location, first process the extracted work location information into a string vector; find the corresponding geographical location information through a map search engine; then obtain the name of the owner / tenant of the location building and match it with the enterprise name; if the target recruitment enterprise name can be matched, encode it as 0, otherwise encode it as 1. For the work department, first process the extracted work department information into a string vector; then match it with the information on the company's official website, and if the corresponding string can be matched on the official website, encode it as 0, otherwise encode it as 1. For the monthly salary range, process it into a floating-point value; for the job description, first remove HTML symbols, spaces, and punctuation marks through parsing methods and regular expressions; then identify and remove stop words through a stop word library; finally, process the remaining text into a string vector. Regarding the welfare benefits, first remove HTML symbols, spaces, and punctuation marks through parsing methods and regular expressions; then identify and remove stop words through a stop word library; finally, process the remaining text into a string vector. Regarding the work type, check the crawled relevant information to see if there are characters such as "full-time", "part-time", "labor dispatch", "hourly wage", etc. If there is "full-time", record it as 3, if there is "part-time", record it as 2, if there is "labor dispatch", record it as 1, if there is "hourly wage", record it as 0, and if none of them exist, record it as a null value. Regarding the experience requirement, first identify whether the crawled position-related information contains the string "experience". If it does, identify the numerical value before the string "year" to obtain the required number of working years and process it into a floating-point value, and record the minimum value among them; if not, identify whether there is the string "fresh graduate", and if there is, record it as 0, if not, record it as a null value. Regarding the educational requirement, first identify the description of the educational requirement from the crawled position-related information and match it among the strings "master or above", "bachelor's degree", "college degree", "secondary school or below", "no requirement". The above options are processed as 4, 3, 2, 1, 0 for recording respectively. For the school type, identify the description of the graduating school from the crawled position-related information and match it among the strings "985", "211", etc. If "985" is identified, record it as 2, if "211" is identified, record it as 1, if both "985" and "211" are identified, also record it as 1, and if no matching field is identified, record it as 0. For the work industry risk, first match the crawled job description information in the "Classification of National Economy Industries of the People's Republic of China" to obtain the affiliated industry; then search for relevant news using the affiliated industry as the keyword. If keywords such as "layoffs", "transformation", "market value evaporation", "store closure" are searched, it is defined as high risk and recorded as 1, otherwise it is defined as low risk and recorded as 0.Regarding the number of times of repeated recruitment for the same position, first, all the job positions published by the target recruitment enterprises crawled in the recent period (such as three years) are captured; then, the duration of the same position published online is identified and converted into a floating-point value on an annual basis; if a position goes offline and then goes online again, the online durations are added separately and converted into a floating-point value.

[0047] The job title, work location, work department, monthly salary range, job description, welfare benefits, work type, experience requirements, educational requirements, school type, work industry risk, number of times of repeated recruitment for the same position, etc. of the aforementioned target position are deep indicators for a target position, and the collection after their standardized processing can be used as a prediction sample set for the target position. In another embodiment, while obtaining the target position deep indicator data, it also includes the step of establishing an enterprise online recruitment information matrix according to the position level information, and obtaining the relevant indicators in terms of breadth between the recruitment platform and the target position from the enterprise online recruitment information matrix, hereinafter referred to as breadth indicators.

[0048] Among them, the enterprise online recruitment information matrix includes one or more job information released by the target enterprise and the recruitment platform qualification information for releasing these job information; job information released by one or more similar enterprises of the target enterprise and the recruitment platform qualification information for releasing these job information. Among them, the job positions released by the target enterprise include the target position and a second position different from the target position; the recruitment platform includes one or more target recruitment platforms for releasing the target position provided by the target enterprise and one or more second recruitment platforms for releasing the second job information different from the target position.

[0049] The steps for processing the corresponding breadth indicator data based on the enterprise online recruitment information matrix include:

[0050] Calculate the qualification level vector of each recruitment platform based on the recruitment platform qualification information;

[0051] Count the number of job positions released by each recruitment platform from the enterprise online recruitment information matrix;

[0052] Determine the input degree of each recruitment platform based on the number of job positions; for example, use the number of job positions as the input degree, or normalize the number of job positions and use the normalized value as the input degree.

[0053] Use the input degree of each recruitment platform as the weight of the qualification level vector of the recruitment platform, and calculate the weighted vector value of the qualification level of each recruitment platform. Among them, the weighted vector value of the qualification level of the recruitment platform is a kind of breadth indicator.

[0054] The steps for processing the corresponding breadth indicator data based on the enterprise online recruitment information matrix also include:

[0055] Query for multiple second target positions that are the same as the target position and provided by similar enterprises;

[0056] Calculate the position vectors of the target position and each second target position based on the position information;

[0057] Perform clustering operations on the multiple second target positions to obtain the second target position with the largest clustering result;

[0058] Calculate the first vector difference between the position vector of the target position and the position vector of the second target position with the largest clustering result, and use the first vector difference as the external consistency coefficient of the target position. Among them, the external consistency coefficient of the target position is another breadth metric.

[0059] The steps of processing the enterprise network recruitment information matrix to obtain the corresponding breadth metric data further include:

[0060] Calculate the second vector difference between the position vector of the target position and the position vector of each second position provided by the target enterprise, and calculate the average value of the second vector differences, and use the average value of the second vector differences as the internal consistency coefficient of the target position. Among them, the internal consistency coefficient of the target position is another breadth metric.

[0061] The steps of processing the enterprise network recruitment information matrix to obtain the corresponding breadth metric data further include:

[0062] Determine the weights of the first vector difference and the average value of the second vector differences according to their respective closeness relationships with the target position vector, and calculate the weighted average of the two as the consistency coefficient of the target position. Among them, the consistency coefficient of the target position is another breadth metric.

[0063] In one embodiment, the above-obtained depth metric data and breadth metric data are combined together as the prediction sample set of the target position.

[0064] The processing of interview-level indicator data includes: regarding the interview location, first process the interview location information into a string vector; then find the corresponding geographical location information through a map search engine; then obtain the name of the owner / tenant of the location building and match it with the name of the target recruitment enterprise; if the name of the target recruitment enterprise is matched, it is encoded as 0, otherwise it is 1. For the interview time, if the interview time information is between 8:00 and 18:00, it is recorded as 0, otherwise it is recorded as 1. For the interview notification method, if it is by phone call, it is encoded as 1, if it is by text message, it is encoded as 2, if it is by email, it is encoded as 3, if it is by recruitment application (recruitment App), it is encoded as 4, if it is by ordinary social software, it is encoded as 5, and multiple contact methods are allowed to be filled in. Regarding whether there is a written test, if there is a written test, it is encoded as 0, otherwise it is encoded as 1.

[0065] Step S5a, evaluate the enterprise-level risk, position-level risk, and interview-level risk of false recruitment by the target enterprise according to multiple indicator data at each level. In this embodiment, when evaluating risks at each level, a trained model is used for evaluation respectively. The models can adopt algorithms such as decision tree, naive Bayes, multi-layer, k-nearest neighbor, random forest, neural network, etc.

[0066] In this embodiment, the machine learning models at each level are trained respectively using the labeled samples at each level. Taking the training of the enterprise-level machine learning model as an example, the process of training the machine learning model is described.

[0067] The training data of the enterprise-level machine learning model includes a certain number of enterprise sample data. Each enterprise sample data includes multiple indicator data labeled as a false recruitment enterprise or a real recruitment enterprise. Each indicator data is used as a feature data of the enterprise sample. For example, factors such as "number of employees in the enterprise", "number of branch companies", "registered capital", "business scope", "enterprise type", "financing information", "enterprise age", "annual turnover", "enterprise legal litigation events", "major investments received", "enterprise evaluations on social media", etc. are used as independent variables (i.e., one-dimensional features of a sample), and each sample is manually marked for whether there is enterprise-level risk. For example, a sample with risk is marked as "negative" / "positive", and a sample without risk is marked as "positive" / "negative", so as to obtain a sample set. The sample set should meet some requirements for training the model, such as the balance of the number of positive and negative samples, the same feature dimension for each sample, etc. 80% of the samples in the sample set are used as training samples, and 20% of the samples are used as verification samples, so as to form a training set and a verification set respectively.

[0068] Then, model building is performed using any one of the algorithms such as the aforementioned decision tree, Naive Bayes, multi-layer, k-nearest neighbor, random forest, neural network, etc. The model performs supervised learning according to the model algorithm based on the samples in the training set to obtain the probability of enterprise-level risk for each sample.

[0069] Since only two result types, namely "risky" and "risk-free", are set in this embodiment, the model uses binary classification output. Assuming the number of sample independent variables (i.e., the dimension of sample features) is n, the mathematical representation of the model is:

[0070] X = {x1, x2,..., x n}

[0071] Y = {y0, y1}

[0072] Wherein X is the feature set of a sample input to the model, and each x i is a one-dimensional feature. For example: x1 is the number of employees in the enterprise, x2 is the number of branch companies, x3 is the registered capital,..., x n is the annual turnover. Y is the classification set of the model, and y i ∈ {0, 1}. In this embodiment, y0 represents "risk-free" and y1 represents "risky". The model calculates, based on a given sample X i , the probability p(c j (the label corresponding to each classification y. For example, when y0 takes the value 0, the corresponding label c0 is "no enterprise-level risk")) given X j . Then the output of the model is: i ) is:

[0073]

[0074] Where represents the classification prediction value, and p(c j |X i ) represents the probability of each classification given the sample. In one embodiment, a comparison threshold is set, such as set to 0.6 - 0.9. In this embodiment, the threshold of 0.8 is used for illustration. When determining whether there is enterprise-level risk, if the predicted probability of y1 output by the algorithm model exceeds 0.8, it can be confirmed that the enterprise corresponding to the sample has enterprise-level risk. Among them, the threshold can be an empirical value obtained by repeatedly operating the model based on practical data, or further adjusted as the model iterates. The trained model is verified using the verification set samples, and when the model meets the evaluation criteria, it can be used for online risk assessment.

[0075] The training processes of the position-level machine learning model and the interview-level machine learning model are similar to the aforementioned process, and will not be elaborated here.

[0076] Among them, in step S4a, in order to make the normalized data meet the input requirements of the model, that is, a sample consists of multi-dimensional features and the number of features is the same. After specifically normalizing the data of each level index into floating-point numbers, string vectors or codes, according to the input requirements of each level model, these data are combined into the prediction samples of the model, and then in step S5a, the prediction samples of each level are input into the corresponding model, so as to obtain the risks corresponding to each level.

[0077] In the foregoing embodiment, the output of the machine learning model at each level is "at risk" or "not at risk". In one embodiment, as Figure 2 shown, in steps S510a, 511a and 512a respectively, each prediction sample is input into the machine learning model at the corresponding level, and then in step S513a, the results output by each level of machine learning model, that is, the enterprise-level risk, the position-level risk and the interview-level risk, are combined as the evaluation result, and the evaluation result is provided to the user in step S6a, so that the user can know whether there is a risk from three aspects: enterprise, position and interview.

[0078] In another embodiment, as Figure 3 shown, taking "at risk" or "not at risk" output by the model as two grades, represented by 1 and 0 respectively, so that the risk grades at the enterprise level, the position level and the interview level can be encoded to obtain risk codes. After obtaining the risks at the corresponding levels by inputting each prediction sample into the machine learning model at the corresponding level in steps S510a, 511a and 512a, the enterprise-level risk, the position-level risk and the interview-level risk are encoded in step S523a. For example, the risk code 001 represents that there is a risk only during the interview, while the risk code 111 represents that there are risks at the enterprise, position and interview levels. In order to enable the user to feel the magnitude of the risk, different risk levels are set in this embodiment, and the risk codes correspond to the risk levels, as shown in Table 1 below:

[0079] Table 1:

[0080] Risk Code Risk Level 111 High Risk 110、101 Relatively High Risk 100 Medium Risk 010、011 Relatively Low Risk 000、001 Low Risk

[0081] In step S524a, according to the obtained risk code, query the corresponding table of the risk code and the risk level, as shown in Table 1, then the risk level corresponding to the obtained risk code can be obtained and determined as the final risk level, and then in step S6a, the final risk level is provided to the user as the evaluation result.

[0082] In another embodiment, as Figure 4 shown, after encoding the obtained enterprise-level, position-level and interview-level risk levels in step S533a, it further includes:

[0083] Step S534a, obtain the respective weights of the enterprise level, position level, and interview level.

[0084] Step S535a, perform weighted calculation based on the values on each digit in the risk code and their respective weights to obtain the weighted sum of the risk code.

[0085] Step S536a, query the correspondence table between the weighted sum of the risk code and the risk level according to the current weighted sum of the risk code, as shown in Table 2, so as to obtain the risk level matching the current weighted sum of the risk code.

[0086] Among them, in one embodiment, according to the weights of the enterprise level, position level, and interview level being 5, 4, and 1 respectively, calculate the weighted sums of the risk codes 000 - 111 respectively, obtaining 0, 1, 4, 5, 5, 6, 9, 10. Then, according to the gap between adjacent values, divide the following 8 values into 3 groups, namely (0, 1), (4, 5, 6), (9, 10), thus obtaining Table 2:

[0087] Table 2:

[0088]

[0089]

[0090] For example, in one embodiment, when the following risks are obtained according to the models at each level: enterprise level: "There is a risk"; position level: "No risk"; interview level: "No risk", the risk code is 100, and the corresponding weighted sum of the risk code is 5. Query Table 2, so as to match the corresponding risk level of "Medium risk".

[0091] Finally, in step S6a, provide the risk level queried from Table 2 to the user as the evaluation result.

[0092] In the above embodiments, the models at each level adopt binary classification output, that is, two risk levels of "There is a risk" and "No risk". Of course, it is also possible to train the models at each level to adopt multi-classification output. For example, when using 5 levels of "High risk", "Higher risk", "Medium risk", "Lower risk", and "Low risk", assuming the number of sample independent variables (i.e., the dimension of sample features) is n, the mathematical representation of the model is:

[0093] X = {x1, x2,..., x n}

[0094] Y = {y0, y1, y2, y3, y4}

[0095] Among them, X is the feature set of a sample input to the model, and each x iis a one-dimensional feature. For example, x1 is the number of employees in the enterprise, x2 is the number of branch companies, x3 is the registered capital, …, x n is the annual turnover. Y is the classification set of the model, and y i ∈{0,1,2,3,4}. In this embodiment, y0 represents "low risk", y1 represents "relatively low risk", y2 represents "medium risk", y3 represents "relatively high risk", and y4 represents "high risk". The model is based on a given sample X i , and calculates the probability p(c j (the label corresponding to each classification y. For example, when y0 takes the value 0, the corresponding label c0 is "risk-free") given X j |X i ). Then the output of the model is:

[0096]

[0097] where represents the classification prediction value, and p(c j |x i ) represents the probability of each classification under the given sample. In one embodiment, a comparison threshold is set, generally set to 0.6 - 0.9. In this embodiment, the threshold of 0.8 is used for illustration. For example, when judging the enterprise-level risk category, if the predicted probability of y1 output by the algorithm model exceeds 0.8, it can be confirmed that the enterprise corresponding to the sample has a relatively low enterprise-level risk. Among them, the threshold can be an empirical value obtained by repeatedly operating the model with practical data, or further adjusted as the model iterates. The trained model is verified using the validation set samples, and when the model meets the evaluation criteria, it can be used for online risk assessment.

[0098] The training processes of the job level machine learning model and the interview level machine learning model are similar to the foregoing processes, and will not be elaborated here.

[0099] After obtaining the risk level of each level, the foregoing embodiment is used to encode or calculate the weighted sum of the encodings to determine the final risk level, which will not be elaborated here.

[0100] Figure 5It is a flowchart for risk assessment according to another embodiment of the present invention. In this embodiment, there is one enterprise-level machine learning model, and multiple position-level and interview-level machine learning models. The output categories of the enterprise-level, position-level, and interview-level machine learning models are more than two, and the number of output categories of these three models can be the same or different. In the order from enterprise level, position level to interview level, the levels are set from top to bottom. The number of machine learning models at the next lower level is the same as the number of output categories of the machine learning model at the previous higher level, and they are respectively trained with data having risks of the output categories of the machine learning model at the previous higher level. For example, multiple position-level machine learning models are respectively trained with data having corresponding enterprise-level risk levels, and respectively correspond to the risk levels output by the corresponding enterprise-level machine learning models; multiple interview-level machine learning models are respectively trained with data having corresponding enterprise-level risk levels and corresponding position-level risk levels, and respectively correspond to the risk levels output by the corresponding position-level machine learning models.

[0101] For example, in one embodiment, the outputs of the enterprise-level machine learning model and the position-level machine learning model are respectively binary classification outputs, which are respectively defined as two risk levels of "at risk" and "not at risk". The interview-level machine learning model is a multi-classification output. For example, it is respectively defined as 5 levels of "high risk", "relatively high risk", "medium risk", "low risk", and "not at risk". The enterprise-level machine learning model is denoted as M1, and there are two position-level machine learning models, which are respectively denoted as model M2.1 and model M2.2. Among them, the position samples in the training set of model M2.1 are all samples corresponding to no risk at the enterprise level, while the position samples in the training set of model M2.2 are all samples corresponding to risk at the enterprise level. There are 4 interview-level machine learning models, which are respectively denoted as model M3.1, M3.2, M3.3, and model M3.4. The position samples in the training set of model M3.1 are all samples corresponding to no risk at the enterprise level and no risk at the position level. The position samples in the training set of model M3.2 are all samples corresponding to no risk at the enterprise level and risk at the position level. The position samples in the training set of model M3.3 are all samples corresponding to risk at the enterprise level and no risk at the position level. The position samples in the training set of model M3.4 are all samples corresponding to risk at the enterprise level and risk at the position level.

[0102] It can be seen from the training samples of the machine learning model that the enterprise-level machine learning model is the first level, the position-level machine learning model is its subordinate model, and the interview-level machine learning model is the subordinate model of the position-level machine learning model. The selection of the subordinate model should correspond to the risk and other outputs of the superior model.

[0103] Select the next-level machine learning model according to the risk level output by the previous-level machine learning model in the order from top to bottom of enterprise level, position level, and interview level, where the risk and level output by the interview-level machine learning model are the final risk and its level. The specific risk assessment process is as Figure 5 shown and includes the following steps:

[0104] Step S51a: Input the enterprise-level prediction sample into the enterprise-level machine learning model M1 for prediction. In one embodiment, the enterprise-level machine learning model is denoted as M1, and the enterprise-level prediction sample is input into this model M1.

[0105] Step S52a: Determine whether the output of the enterprise-level machine learning model M1 is "at risk". If it is "at risk", execute Step S53a; if it is "not at risk", execute Step S57a.

[0106] Step S53a: Select the position-level machine learning model M2.2, and input the position-level prediction sample into the position-level machine learning model M2.2 for prediction.

[0107] Step S54a: Determine whether the output of the position-level machine learning model M2.2 is "at risk". If it is "at risk", execute Step S55a; if it is "not at risk", execute Step S56a.

[0108] Step S55a: Select the interview-level machine learning model M3.4, and input the interview-level prediction sample into the interview-level machine learning model M3.4 for prediction. Take the output of inputting the interview-level prediction sample into model M3.4 as the risk assessment result, and then end the risk assessment process.

[0109] Step S56a: Select the interview-level machine learning model M3.3, and input the interview-level prediction sample into the interview-level machine learning model M3.3 for prediction. Take the output of the interview-level machine learning model M3.3 as the risk assessment result, and then end the risk assessment process.

[0110] Step S57a: Select the position-level machine learning model M2.1, and input the position-level prediction sample into the position-level machine learning model M2.1 for prediction.

[0111] Step S58a: Determine whether the output of the position-level machine learning model M2.1 is "at risk". If the output of the position-level machine learning model M2.1 is "at risk", execute Step S59a; if the output of the position-level machine learning model M2.1 is "not at risk", execute Step S510a.

[0112] Step S59a, select the interview-level machine learning model M3.2, input the interview-level prediction samples into the interview-level machine learning model M3.2 for prediction, use the output of the interview-level machine learning model M3.2 as the risk assessment result, and then end the risk assessment process.

[0113] Step S510a, select the interview-level machine learning model M3.1, input the interview-level prediction samples into the interview-level machine learning model M3.1 for prediction, use the output of the interview-level machine learning model M3.1 as the risk assessment result, and then end the risk assessment process.

[0114] In this embodiment, all information is divided according to the overall enterprise level, position level, and interview level. The enterprise-level assessment results govern the position-level assessment results, and the position-level assessment results govern the interview-level assessment results, so as to make a comprehensive and accurate risk judgment.

[0115] In this embodiment, the interview-level machine learning model outputs multiple classifications. In one embodiment, the interview-level machine learning model calculates the respective classification probabilities for the input prediction samples. For example, it obtains the probabilities of "high risk", "relatively high risk", "medium risk", "low risk", and "no risk" respectively; then compares each classification probability with the corresponding second classification threshold; if one of the classification probabilities is greater than or equal to the corresponding second classification threshold, it determines that the model output is the risk level defined by the classification.

[0116] In another embodiment, the output of the interview-level machine learning model is binary classification, defined as two risk levels of "at risk" and "not at risk"; when calculating the prediction samples, the interview-level machine learning model calculates the probability of "at risk"; and compares the probability of "at risk" with multiple third classification thresholds, where the multiple third classification thresholds form multiple risk probability intervals from low to high, corresponding to multiple risk levels from low to high; determines the corresponding risk level according to the risk probability interval where the probability of "at risk" output by the interview-level machine learning model is located.

[0117] In another embodiment, the outputs of the enterprise-level machine learning model and the position-level machine learning model can also be the same as the output of the interview-level machine learning model, that is, multiple classifications. Each position-level machine learning model corresponds to one type of output of the enterprise-level machine learning model, so the number of position-level machine learning models is the same as the number of risk output categories of the enterprise-level machine learning model. Correspondingly, the number of interview-level machine learning models is the product of the number of position-level machine learning models and the number of risk output categories of the position-level machine learning models.

[0118] In step S6a, the determined risk is provided to the user as an identification result, and the providing method includes displaying it in the terminal interface, sending an email, sending a short message through a mobile communication network, or sending a social message to the user's social media account.

[0119] Figure 6 It is a flowchart of a method for processing prediction sample data according to an embodiment of the present invention. In this embodiment, the following steps are included:

[0120] Step S1b, monitoring the feedback information of the user on each identification result. After evaluating the recruitment interview information provided by the user to identify whether the current recruitment faced by the user is a false recruitment and providing the identification result to the user, monitor the feedback information of the user on the identification result. In one embodiment, the feedback information may include confirmation information of "risky" or "risk-free" in the identification result, or evaluation information of "correct" / "wrong" on the identification result, etc.

[0121] Step S2b, extracting the confirmation information of the user on the identification result. According to the information fed back in step S1b, extract the confirmation information of "risky" or "risk-free", or the evaluation information of "correct" / "wrong" from it.

[0122] Step S3b, setting corresponding labels for the corresponding prediction samples according to the confirmation information of the user. For example, when the user confirms that the recruitment interview input by him is a real recruitment, set the label of real recruitment for the prediction samples at all levels evaluated for its recruitment. If the user confirms that the recruitment interview input by him is a false recruitment, set the label of false recruitment for the prediction samples at all levels evaluated for its recruitment. In a better embodiment, the feedback information of the user also includes more information fields. For example, when the user confirms it is a false recruitment, the reason that the user needs to fill in. The present invention analyzes the reason filled in by the user to evaluate the root cause of the false from two aspects of "enterprise" and "position", and sets labels for the current prediction samples at all levels.

[0123] Step S4b, storing the prediction samples with labels set into the training set, that is, the system stores the prediction samples with labels set into the dataset for training the model, thereby enriching the training data.

[0124] Step S5b, determining whether the model update condition is reached. The update condition is, for example, a preset update period, such as updating the model once a week / month, or counting the number of new training samples in the model training dataset, and determining whether the number of new training samples reaches the threshold. If the update period is reached or the number of new training samples reaches the threshold, it is determined that the model update condition is met, and step S6b is executed. Otherwise, return to execute step S1b.

[0125] Step S6b, optimize and update the currently used machine learning model using the training dataset.

[0126] The present invention can continuously accumulate training data, without the need for manual setting of labels for the training data. The optimization of the model can be automatically performed without manual intervention, thus saving a large amount of manpower and time. Moreover, with the optimization of the model, the accuracy of model evaluation is gradually improved, enabling more accurate identification of false job recruitment, and thus better safeguarding the interests of users.

[0127] In one embodiment, the present invention further provides a false job recruitment warning method. Refer to Figure 7 , Figure 7 which is a flowchart of the false job recruitment warning method according to an embodiment of the present invention, and includes:

[0128] Step S1c, determine whether recruitment interview information provided by the user is received. If the recruitment interview information provided by the user is received, then execute Step S2c; if not, repeat this step.

[0129] Step S2c, extract target recruitment enterprise information, target position information, and interview information from the recruitment interview information provided by the user.

[0130] Step S3c, perform information crawling in the network. Based on the extracted target recruitment enterprise information, target position information, and interview information, perform information crawling in the network according to multiple indicators among enterprise level, position level, and interview level respectively.

[0131] Step S4c, process the crawled information into corresponding level of indicator data.

[0132] Step S5c, perform risk assessment according to multiple indicator data of each level. For example, perform risk assessment according to the process shown in Figure 2-5 any one.

[0133] Step S6c, based on the evaluated risk, trace back and analyze the indicator data corresponding to the risk to obtain warning information and provide it to the user. Among them, the warning information includes problem indicator data causing the risk and risk content.

[0134] The step of tracing back and analyzing the indicator data corresponding to the risk includes:

[0135] First, according to the monitored risk type, traverse and evaluate the indicator data used for evaluating the risk to calculate the contribution degree of each indicator data to the evaluated risk.

[0136] Then, sort the multiple indicators according to the contribution degree to the risk, and determine the abnormal indicators as the multiple indicators with the highest sorting or the indicators with a contribution degree greater than the threshold.

[0137] Finally, risk content is generated based on the content of the abnormal indicator and / or the correlation relationship of the content of multiple abnormal indicators.

[0138] Among them, a feature importance determination method in feature engineering can be used to determine the contribution degree of indicator data to the evaluated risk. The feature importance determination method mentioned above can be, for example, the expert meeting method, in which experts in the field specify and identify the importance of each indicator used in the present invention. Or, methods such as rough set theory and information entropy can be used to identify the importance of each indicator used in the present invention. Or, the importance of an indicator can be determined by mining the correlation relationship between the indicator and the output through data mining technology. Based on the importance of the indicators determined by the above methods, when tracing back the indicators, the contribution degree is determined by querying the importance identifier corresponding to each indicator.

[0139] There are also some methods for determining feature importance that are different from the above. For example, when using a machine learning model based on a neural network, since the weights of the hidden layer nodes in the neural network represent the importance of the features corresponding to the input layer nodes, when the machine learning model uses a neural network model, the weights of the hidden layer nodes are read, and the contribution degree of the indicator can be determined according to the weights of the hidden layer nodes.

[0140] For another example, for a machine learning model using a decision tree or random forest algorithm, according to the decision tree generation principle, the order of the features selected during the decision tree division process can be used as the importance ranking of the features, and this ranking can be obtained through the feature_importances_ attribute in sklearn. Therefore, when tracing back the indicators in the present invention, the importance of each indicator used when evaluating the risk, that is, the contribution degree, can be obtained by calling the feature_importances_ attribute.

[0141] After determining the contribution degree of the indicators, several indicators with the highest contribution degree ranking (such as the top three indicators) can be used as abnormal indicators. After determining the abnormal indicators, the content of the abnormal indicators is read, thereby locating the points where risks may occur and the specific risk content. Analyze whether there is a correlation relationship among the contents of the three abnormal indicators. If so, the associated content is also used as risk content.

[0142] The present invention pre-sets risk events that may occur corresponding to various risk contents and their countermeasures. Therefore, after determining the risk content, the database is queried according to the risk content, so as to determine the matching countermeasure.

[0143] Taking the assessment of enterprise risks as an example, if it is found after retrospective analysis that "enterprise legal litigation events" are the main causes of enterprise risks, the corresponding solutions provided by the system include: (1) During the interview, it is appropriate to ask and verify with the HR about the enterprise's labor dispute litigation situation; (2) Before the interview or employment, collect relevant information online to understand the details of relevant litigations and strengthen the overall understanding of this issue, etc.

[0144] Add the above-mentioned risk content and corresponding solutions to the warning information and present them to the user together. In addition, in order to enable the user to better understand the content identified this time or facilitate the user's future query, the present invention can also generate a risk report from the content in the warning information and store it or provide it to the user. For example, generate a risk report including the evaluated risk level, risk content, and corresponding solutions and send it to the user or store it in the user account.

[0145] Figure 8 It is a principle block diagram of a false recruitment identification system according to an embodiment of the present invention. In this embodiment, the false recruitment identification system includes a user information acquisition module 1, a data collection module 2, an index data generation module 3, and a risk assessment module 4. Among them, the user information acquisition module 1 is connected to the data collection module 2. The user information acquisition module 1 receives the recruitment interview information provided by the user and extracts the target recruitment enterprise information, target position information, and interview information from the user's recruitment interview information. The data collection module 2 is connected to the Internet and is connected to the user information acquisition module 1. Based on the extracted target recruitment enterprise information, target position information, and interview information, it crawls information in the network according to multiple indicators such as enterprise level, position level, and interview level to obtain relevant information such as the recruitment enterprise and the target position. The index data generation module 3 is connected to the data collection module 2 and processes the crawled enterprise level, position level, and interview level information to obtain multiple index data corresponding to each level. Among them, after being processed by the index data generation module 3, based on the obtained information at each level, it is standardized into specific index data, such as floating-point numbers, string vectors, or a certain encoding. The risk assessment module 4 is connected to the index data generation module 3 and is configured to evaluate according to multiple index data at each level to obtain the enterprise-level risk, position-level risk, and interview-level risk of the target enterprise for false recruitment.

[0146] Figure 9 It is a partial principle block diagram of a false recruitment identification system according to another embodiment of the present invention. In this embodiment, a machine learning model is used to evaluate risks. Therefore, in this embodiment, in addition to including Figure 8In addition to the modules in [description], it also includes a prediction sample generation module 5, which is connected to the indicator data generation module 3. Based on the samples required by each level of machine learning model, each piece of indicator data is normalized into feature data, and the feature data corresponding to the indicator data at each level is combined together to form the prediction samples for each level of machine learning model. In this embodiment, the risk assessment module 4 includes an enterprise-level risk assessment unit 41, a position-level risk assessment unit 42, an interview-level risk assessment unit 43, a coding unit 44, and a risk level query unit 45. The prediction sample generation module 5 inputs the prediction samples at each level to the enterprise-level risk assessment unit 41, the position-level risk assessment unit 42, and the interview-level risk assessment unit 43 respectively. The enterprise-level risk assessment unit 41 obtains the enterprise-level risk probability by inputting the enterprise-level prediction sample to the trained enterprise-level machine learning model. The position-level risk assessment unit 42 obtains the position-level risk probability by inputting the position-level prediction sample to the trained position-level machine learning model. The interview-level risk assessment unit 43 obtains the interview-level risk probability by inputting the interview-level prediction sample to the trained interview-level machine learning model. In one embodiment, the risk probabilities obtained by these three units can be directly provided to the user as the recognition result. In this embodiment, the risk probabilities obtained by these three units are output to the coding unit 44. The coding unit 44 encodes the obtained enterprise-level, position-level, and interview-level risk levels to obtain a risk code, and outputs the risk code to the risk level query unit 45. The risk level query unit 45 queries the corresponding table of the code and the risk level according to the risk code, such as Table 1 involved in the description of the method above, and determines the risk level corresponding to the risk code as the final risk level. Of course, after obtaining the risk code, the coding unit 44 can also calculate the weighted sum of the risk codes according to the weights of each level of risk. The risk level query unit 45 queries the corresponding table of the weighted sum of the code and the risk level according to the weighted sum of the risk codes, such as Table 2 involved in the description of the method above, and determines the risk level corresponding to the weighted sum of the risk codes as the final risk level.

[0147] Figure 10 is a schematic block diagram of a false recruitment identification system according to another embodiment of the present invention. This embodiment is the same as Figure 9Compared with the embodiments shown, in addition to the enterprise-level risk assessment unit 41, the position-level risk assessment unit 42, and the interview-level risk assessment unit 43, the risk assessment module 4 in this embodiment further includes a selection unit 46. In this embodiment, the machine learning models used by the three assessment units are related to each other. Among them, in the order from top to bottom of the enterprise level, the position level, and the interview level, the lower-level machine learning model is trained according to the training data corresponding to the risk level output by its upper-level machine learning model. Therefore, the number of lower-level machine learning models is the same as the number of risk levels of the upper-level machine learning model, and the risk assessment is also carried out step by step from top to bottom, and it is necessary to select the lower-level machine learning model according to the risk level of the upper-level machine learning model.

[0148] Specifically, the enterprise-level risk assessment unit 41 receives the enterprise-level prediction samples obtained by the prediction sample generation module 5 and notifies the selection unit 46. The selection unit 46 selects an enterprise-level machine learning model from the model library and sends it to the enterprise-level risk assessment unit 41. The enterprise-level risk assessment unit 41 inputs the enterprise-level prediction samples into the enterprise-level machine learning model, and after being evaluated by the enterprise-level machine learning model, obtains the enterprise-level risk level, and while sending the enterprise-level risk level to the selection unit 46, notifies the position-level risk assessment unit 42. The selection unit 46 selects a suitable position-level machine learning model according to the enterprise-level risk level and sends it to the position-level risk assessment unit 42. After receiving the notification from the enterprise-level risk assessment unit 41 and the position-level machine learning model sent by the selection unit 46, the position-level risk assessment unit 42 inputs the position-level prediction samples received from the prediction sample generation module 5 into the position-level machine learning model, and after being evaluated by the position-level machine learning model, obtains the position-level risk level, and outputs the position-level risk level to the selection unit 46, and at the same time outputs a notification to the interview-level risk assessment unit 43. The selection unit 46 selects the corresponding interview-level machine learning model according to the position-level risk level and sends it to the interview-level risk assessment unit 43. After receiving the notification from the position-level risk assessment unit 42 and the interview-level machine learning model sent by the selection unit 46, the interview-level risk assessment unit 43 inputs the interview-level prediction samples received from the prediction sample generation module 5 into the interview-level machine learning model, and after being evaluated by the interview-level machine learning model, obtains the final risk level.

[0149] The outputs of the enterprise-level machine learning model and the position-level machine learning model in the foregoing embodiments are respectively binary classifications, which are respectively defined as two risk levels of "at risk" and "not at risk". The output of the interview-level machine learning model is a binary classification or a multi-classification. When it is a binary classification, it is respectively defined as two risk levels of "at risk" and "not at risk". When it is a multi-classification, it is respectively defined as multiple risk levels of different degrees from none to having.

[0150] Figure 11 It is a schematic block diagram of a user information acquisition module according to an embodiment of the present invention. In this embodiment, the user information acquisition module includes a message sending and receiving unit 11 and an information extraction unit 12. The message sending and receiving unit 11 serves as an interaction interface between the system and the user. On the one hand, it is connected to the risk assessment module 4 and sends the final recognition result to the user. On the other hand, it receives the recruitment interview information provided by the user and outputs it to the information extraction unit 12. The information extraction unit 12 extracts the target recruitment enterprise information, target position information, and interview information from the recruitment interview information provided by the user.

[0151] The message sending and receiving unit 11 includes one or more of the following units: an application terminal user interaction unit 110, an email processing unit 111, a mobile short message processing unit 112, and a social media message processing unit 113. The application terminal user interaction unit 110 includes at least an input interface through which recruitment interview information provided by the user via the application terminal can be obtained. In addition, the application terminal user interaction unit 110 may also include a display interface for displaying messages such as the identified risk level, risk content, or corresponding solutions. The email processing unit 111 can identify the recruitment interview information provided by the user via email according to the email address or subject, and can also send messages to the user via email. The mobile short message processing unit 112 identifies the recruitment interview information provided by the user via short message according to the message sending number and subject in the short message received based on the mobile communication network, and can also send messages to the user. The social media message processing unit 113 identifies the recruitment interview information provided by the user via social media from the social media messages or sends messages to the user according to the message sender.

[0152] Figure 12 It is a schematic block diagram of a data processing system according to an embodiment of the present invention. The data processing system in this embodiment includes a user information acquisition module 1, a data collection module 2, an index data generation module 3, and a prediction sample generation module 5, that is Figure 9 or Figure 10 Some of the modules in the false recruitment identification system in constitute a data processing system, which performs information extraction, information crawling, data normalization, etc. processing based on the recruitment interview information provided by the user, so as to obtain prediction samples for evaluating risks by a machine learning model. Specifically, Figure 13 It is a schematic block diagram of a data collection module according to an embodiment of the present invention. In this embodiment, the data collection module includes an index acquisition unit 21, an index analysis unit 22, an information crawling unit 23, and an information matrix construction unit 24.

[0153] The index acquisition unit 21 is configured to read a plurality of indexes respectively applied to the enterprise level, the position level, and the interview level. In this embodiment, the system stores indexes for evaluating risks at each level. The index acquisition unit 21 reads these indexes from the system database and sends them to the index analysis unit 22. The index analysis unit 22 analyzes each index, determines the index reference content required to obtain the index data, and sends the index reference content to the information crawling unit 23. The information crawling unit 23 is connected to the index analysis unit 22 and crawls the corresponding information from the Internet according to the determined index reference content. In one embodiment, the system stores one or more retrieval keywords corresponding to the indexes. For example, when the index acquisition unit 21 reads the index of "number of enterprise employees", the index analysis unit 22 queries its retrieval keyword to obtain the index reference content of "number of enterprise employees / quantity", and the information crawling unit 23 searches and obtains the information of the number of enterprise employees on the official website of the target recruitment enterprise according to the index reference content of "number of enterprise employees / quantity", so as to obtain the information that meets the index of "number of enterprise employees". Another example is that when the index acquisition unit 21 reads the enterprise-level index of "enterprise legal litigation events", the index analysis unit 22 queries its retrieval keyword to obtain the index reference content of "labor / defendant", etc., and the information crawling unit 23 queries the relevant litigation information of the target recruitment enterprise in the recent period (such as five years) according to this content.

[0154] In this embodiment, the information matrix construction unit 24 is connected to the information crawling unit 23 and establishes an information matrix according to the association between the crawled information. The information matrix includes one or more position information released by the target enterprise and the qualification information of the recruitment platforms that release this position information; one or more position information released by one or more similar enterprises of the target enterprise and the qualification information of the recruitment platforms that release this position information. Among them, the positions released by the target enterprise include the target position and a second position different from the target position; the recruitment platforms include one or more target recruitment platforms that release the target position provided by the target enterprise and one or more second recruitment platforms used to release the second position information different from the target position.

[0155] Figure 14Principle block diagram of the indicator data generation module according to an embodiment of the present invention. In this embodiment, the indicator data generation module includes a data cleaning unit 31, a single indicator extraction unit 32, and a composite indicator calculation unit 33. Among them, the data cleaning unit 31 performs data cleaning on the crawled original information, including removing some network symbols, punctuation marks, querying the stop word list to remove stop words, and other processing. The single indicator extraction unit 32 is connected to the data cleaning unit 31 to extract and normalize single indicator data from the cleaned data. The single indicators are, for example, some indicators that do not require complex calculations and processing, such as "number of employees in the enterprise", "number of branch companies", "registered capital", "job title", "educational requirements", and other indicators. For these indicators, the indicator data can be extracted by identifying keywords and processed into string vectors, floating-point values, or codes according to specific indicators. The composite indicator calculation unit 33 is connected to the single indicator extraction unit 32 and is configured to calculate more than one single indicator data according to the composite indicator calculation rule to obtain composite indicator data. The composite indicators are, for example, indicators such as "input degree of the recruitment platform", "qualification level of the recruitment platform", "external consistency coefficient of the target position", etc. The calculation method is as described in the foregoing description of the method part and will not be elaborated here.

[0156] Figure 15It is a schematic block diagram of a model training module in a data processing system according to another embodiment of the present invention. In this embodiment, the model training module 6 includes a training data set unit 61, a model training unit 62, a user feedback monitoring unit 63, a sample annotation unit 64, and a model update unit 65. The training data set unit 61 is used to provide a data set for training a model, which respectively includes an enterprise-level training data subset, a position-level training data subset, and an interview-level training data subset according to the type of the trained model. Further, each training data subset also includes a training subset and a validation subset. The model training unit 62 trains the model with the data in the corresponding type of training data subset according to the type of the trained model, and respectively obtains an enterprise-level machine learning model, a position-level machine learning model, and an interview-level machine learning model. In one embodiment, the model training unit 62 trains the model with mutually independent training sets respectively, so as to obtain three independent machine learning models. In another embodiment, according to the risk level of the enterprise-level machine learning model, the position-level training data subset respectively includes different subsets composed of training data with corresponding enterprise risks, and different position-level machine learning models are obtained according to different training subsets. Similarly, corresponding to the interview-level training data subset, which includes multiple training subsets composed of training data with specific corresponding enterprise risk levels and position-level risk levels, different interview-level machine learning models are obtained according to different training subsets. As an example involved in the description of the method of the present invention as mentioned above, the enterprise-level machine learning model M1, the position-level machine learning models M2.1, M2.2, and the interview-level machine learning models M3.1, M3.2, M3.3, and M3.4 are obtained through different training subsets. In order to expand the training set, the user feedback monitoring unit 63 monitors the feedback information of the user on the false recruitment recognition result, and the feedback information at least includes confirmation information of having risk or no risk. The sample annotation unit 64 is connected to the user feedback monitoring unit 63, and based on the user's feedback information, risks are annotated for the data for obtaining the recognition result and added to the corresponding training data subset, so as to achieve the purpose of enriching the training data. The model update unit 65 is connected to the training data set unit 61, and is used to monitor the update condition, and when the model update condition is met, send an update notice to the model training unit 62. The model training unit 62 trains, optimizes, and updates the model with the training data. Among them, the update condition is, for example, reaching a preset update period, such as updating the model once a week / month, or the number of newly added training samples reaching a threshold. Therefore, the model update unit 65 times after each model update, and sends an update notice to the model training unit 62 when the timing period is reached. Or the model update unit 65 counts the newly added training data in the training data set, and when the number of newly added training data reaches the threshold, notifies the model training unit to optimize and update the original machine learning model with the current training data.

[0157] Figure 16 It is a block diagram of the principle of a false recruitment warning system according to an embodiment of the present invention. In this embodiment, the warning system includes a user information acquisition module 1, a data collection module 2, an index data generation module 3, a risk assessment module 4, and a warning module 7. Among them, the user information acquisition module 1, the data collection module 2, the index data generation module 3, and the risk assessment module 4 are the same as the modules in the false recruitment identification system in the foregoing embodiment, and will not be described herein again.

[0158] The warning module 7 is connected to the risk assessment module 4, and in response to the evaluated risk, traces back and analyzes the index data corresponding to the risk to obtain a warning message and provides it to the user.

[0159] Figure 17 It is a block diagram of the principle of the warning module according to an embodiment of the present invention. In this embodiment, the warning module includes a risk monitoring unit 71, a risk content determination unit 72, a warning message generation unit 73, and a warning message sending unit 74.

[0160] The risk monitoring unit 71 is connected to the risk assessment module 4 to monitor whether the risk assessment module 4 has evaluated a risk. When it detects that the risk assessment module 4 has evaluated a risk, it sends a notification to the risk content determination unit 72. The risk content determination unit 72 traverses the index data used to evaluate the risk based on the evaluated risk type to obtain abnormal index data, and determines the risk content based on the content of the one or more abnormal index data or their association relationship. Among them, any one of the foregoing methods can be used to determine the abnormal index data, and will not be described herein again. The warning message generation unit 73 is connected to the risk content determination unit 72 to generate a warning message based on the risk type and the risk content. The warning message sending unit 74 is connected to the warning message generation unit 73 and is used to provide the warning message to the user. Among them, the warning message sending unit 74 can be one or more of an interactive information sending unit, an email processing unit, a mobile short message processing unit, a mobile short message processing unit, and a social media message processing unit. That is to say, the warning message sending unit 74 can be combined with the user information acquisition module into one module, and the module is implemented to have a two-way function of both message reception and sending. For specific reference, see the foregoing embodiment, and will not be described herein again.

[0161] Figure 18 It is a block diagram of the principle of the warning module according to another embodiment of the present invention. In this embodiment, it is the same as Figure 17Compared with the embodiment of [], except for adding the prediction unit 75, the functions of other modules are the same and will not be elaborated here. The prediction unit 75 is connected to the risk content determination unit 72 and is configured to predict risk events and risk response plans according to the risk content, and add the risk events and their response plans to the warning information. For example, the prediction unit 75 queries various suggestions and corresponding plans preset in the system database according to the risk content, so as to obtain the matching suggestions, and combines multiple suggestions or corresponding plans and adds them to the warning information. Another example is that various possible risk events and corresponding response plans corresponding to the risk content are preset in the system database, and these information can be added to the warning information and sent to the user together.

[0162] Figure 19 is the principle block diagram of the warning module according to another embodiment of the present invention. In this embodiment, compared with Figure 18 the embodiment of [], except for adding the risk report generation unit 76, the functions of other modules are the same and will not be elaborated here. The risk report generation unit 76 is connected to the warning information generation unit 73 and is configured to generate a risk report according to the content in the warning information. Correspondingly, the warning information sending unit stores the risk report in a predetermined location or sends it to the user.

[0163] The above embodiments are only for illustrating the present invention and are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can make various changes and modifications without departing from the scope of the present invention. Therefore, all equivalent technical solutions should also fall within the scope of the disclosure of the present invention.

Claims

1. A false recruitment warning system, characterized in that, Including: A recruitment information acquisition module, which is used to receive the recruitment interview information provided by the user, and extract the target recruitment enterprise information, target position information and interview information from the user's recruitment interview information; An index data collection module, which is used to perform information crawling on the network respectively according to multiple indexes in the enterprise level, position level and interview level based on the extracted target recruitment enterprise information, target position information and interview information, and obtain enterprise-level index data, position-level index data and interview-level index data after processing the crawling results; An enterprise-level machine learning model, which is used to receive the target recruitment enterprise information and the crawled enterprise-level index data, and is used to output an enterprise-level risk assessment result of the target recruitment enterprise; Among them, the enterprise-level risk assessment result includes: the target recruitment enterprise belongs to an enterprise-level risky enterprise, or the target recruitment enterprise belongs to an enterprise-level risk-free enterprise; A first low-risk position-level machine learning model, which is used to receive the information of the target recruitment enterprise that is enterprise-level risk-free and the crawled position-level index data, and is used to output a position-level risk assessment result of the target recruitment enterprise that is enterprise-level risk-free; A first high-risk position-level machine learning model, which is used to receive the information of the target recruitment enterprise that is enterprise-level risky and the crawled position-level index data, and is used to output a position-level risk assessment result of the target recruitment enterprise that is enterprise-level risky; Among them, the position-level risk assessment result includes: the target recruitment enterprise belongs to a position-level risky enterprise, or the target recruitment enterprise belongs to a position-level risk-free enterprise; A first low-risk interview-level machine learning model, which is used to receive the information of the target recruitment enterprise that is enterprise-level risk-free and occupation-level risk-free and the crawled interview-level index data, and is used to output an interview-level risk assessment result of the target recruitment enterprise that is enterprise-level risk-free and occupation-level risk-free; A second high-risk interview-level machine learning model, which is used to receive the information of the target recruitment enterprise that is enterprise-level risk-free and position-level risky and the crawled interview-level index data, and is used to output an interview-level risk assessment result of the target recruitment enterprise that is enterprise-level risk-free and position-level risky; A third high-risk interview-level machine learning model, which is used to receive the information of the target recruitment enterprise that is enterprise-level risky and position-level risk-free and the crawled interview-level index data, and is used to output an interview-level risk assessment result of the target recruitment enterprise that is enterprise-level risky and position-level risk-free; A fourth high-risk interview-level machine learning model, which is used to receive the information of the target recruitment enterprise that is enterprise-level risky and position-level risky and the crawled interview-level index data, and is used to output an interview-level risk assessment result of the target recruitment enterprise that is enterprise-level risky and position-level risky; Among them, the interview-level risk assessment result includes: the target recruitment enterprise belongs to an interview-level risky enterprise, or the target recruitment enterprise belongs to an interview-level risk-free enterprise; An early warning module, which is used to respond to the interview-level risk assessment result, trace back and analyze the indicator data corresponding to the risk to obtain early warning information, and provide it to the user.

2. The system according to claim 1, wherein The early warning module responds to the interview-level risk assessment result, traces back and analyzes the indicator data corresponding to the risk, including: the early warning module traverses multiple indicator data for evaluating risks according to the interview-level risk assessment result to obtain the contribution degree of each indicator data to the evaluated risk, determines the abnormal indicators for those with contribution degrees greater than the threshold, and generates the traced-back risk content based on the content of the abnormal indicators and / or the correlation relationship of the content of multiple abnormal indicators.

3. The system according to claim 1, wherein The early warning module is also used to determine a risk response plan according to the risk content in the early warning information, and add the risk response plan to the early warning information.

4. The system according to claim 1, wherein The early warning module is also used to generate a risk report according to the content in the early warning information, and store the risk report at a predetermined location or send it to the user.

Citation Information

Patent Citations

  • False recruitment information detection method based on cascade forest

    CN113704409A

  • False recruitment position detection method based on deep learning

    CN113506084A

  • Recruitment method and system based on block chain

    CN113673234A