Unemployment population identification method and device based on big data, equipment and medium

CN116916263BActive Publication Date: 2026-08-07CHINA MOBILE GRP GUANGDONG CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MOBILE GRP GUANGDONG CO LTD
Filing Date
2022-12-19
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]本发明提供一种基于大数据的失业人群识别方法、装置、设备及介质,用以解决现有技术中失业状态识别的精准度低的缺陷,实现精确识别出有就业需求和无求职需求的失业人群数据,实现提高对于大数据识别人员失业状态的覆盖率、准确性和精细度

Benefits of technology

[0043] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the big data-based method for identifying unemployed people as described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116916263B_ABST
    Figure CN116916263B_ABST
Patent Text Reader

Abstract

The application provides a big data-based unemployed population identification method, device, equipment and medium, comprising: processing user mobile phone signaling data to obtain a full set of to-be-identified data; according to a plurality of specific forms of population characteristics, screening and removing data corresponding to the population characteristics from the full set of to-be-identified data to obtain unemployed population data; and according to APP types, APP use time, APP use traffic, call main and called data and code scanning data, the unemployed population data is subdivided into high-probability unemployed population data, medium-probability unemployed population data and unemployed population data without job-seeking intention. The application is used to solve the defect of low precision of unemployment state identification in the prior art, accurately identify unemployed population data with employment demand and no job-seeking demand, and improve the coverage, accuracy and fineness of big data identification of the unemployed state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data application technology, and in particular to a method, device, equipment and medium for identifying unemployed people based on big data. Background Technology

[0002] In economic theory, the working-age population can be divided into three categories based on their employment status: employed, unemployed, and non-labor force. If they have a job, they are considered employed; if they are unemployed but able to work and seeking employment, they are unemployed; and if they are unemployed and neither seeking nor able to work, they are considered non-labor force. For example, the unemployed can either become employed through finding a job or become non-labor force by leaving the labor market.

[0003] Existing technical solutions for identifying employment status and unemployment status mainly include proactive registration, surveys, and big data-based unemployment population models. However, existing big data-based models primarily rely on mobile phone signaling and location data, which has a limited data dimension and suffers from low accuracy in identifying unemployment status, making it difficult to obtain detailed information on employment and unemployment. Summary of the Invention

[0004] This invention provides a method, apparatus, device, and medium for identifying unemployed individuals based on big data, which addresses the shortcomings of low accuracy in existing technologies for identifying unemployment status. It enables accurate identification of data on unemployed individuals with employment needs and those without job-seeking needs, thereby improving the coverage, accuracy, and precision of big data-based identification of the unemployment status of individuals.

[0005] This invention provides a method for identifying unemployed individuals, comprising:

[0006] Collect user mobile phone signaling data and process the user mobile phone signaling data to obtain the complete set of data to be identified;

[0007] Based on various specific population characteristics, data corresponding to the population characteristics are filtered and removed from the entire set of data to be identified to obtain unemployment population data;

[0008] Based on the type of APP corresponding to the user ID, APP usage time, APP data usage, call origination and reception data, and QR code scanning data, the unemployment data is further subdivided into high-probability unemployment data, medium-probability unemployment data, and unemployment data with no job-seeking intention.

[0009] According to the present invention, an unemployed population identification method based on big data is provided, wherein the population characteristics include non-employed people, confirmed group users, and college students;

[0010] Based on various specific population characteristics, data corresponding to the population characteristics are filtered and removed from the entire dataset to be identified, resulting in unemployment population data, including:

[0011] By using the age field, data in the non-employment age range of the entire dataset to be identified are removed, so that the age range of the target object is within the preset age range;

[0012] By using the group network affiliation field, user data with group network attributes, data inferred to be college students, and user data with rural household registration are removed to obtain the unemployment population data.

[0013] According to the present invention, an unemployed population identification method based on big data is provided, wherein the population characteristics include people working in fixed locations;

[0014] Based on various specific population characteristics, data corresponding to the population characteristics are filtered and removed from the entire dataset to be identified, resulting in unemployment population data, including:

[0015] Based on the latitude and longitude of the base station corresponding to the user ID, the entire set of data to be identified is encoded and clustered so that the latitude and longitude of several base stations are clustered into a geohash code;

[0016] Each user ID is arranged in chronological order of its geohash encoding, and the duration of continuous residence for each user ID under different geohash encodings is calculated.

[0017] The data set to be identified is defined as the data of fixed-location working population that meets the following criteria: the continuous dwell time within a day is greater than a preset dwell time threshold, the number of days corresponding to the continuous dwell time being greater than the preset dwell time threshold is greater than a preset number of days threshold, and the number of geohashes corresponding to a single user ID with a continuous dwell time greater than the preset dwell time threshold within a day is greater than a preset number threshold.

[0018] By removing the data on people working at fixed locations from the dataset to be identified, we obtain data on unemployed people.

[0019] According to the present invention, a method for identifying unemployed people based on big data is provided, wherein the population characteristics include new business workers, which refers to people who do not have a fixed office location but are presumed to be employed;

[0020] Based on various specific population characteristics, data corresponding to the population characteristics are filtered and removed from the entire dataset to be identified, resulting in unemployment population data, including:

[0021] Obtain the APP type, APP usage time, and APP data usage corresponding to the user ID in the complete set of data to be identified;

[0022] The data corresponding to the following in the complete set of data to be identified are: the APP type belongs to the characteristic APP type set; the average daily number of times the APP is used within the period of the APP usage time is greater than a preset number of times threshold; and the average daily traffic usage within the period of the APP traffic usage is greater than a preset traffic threshold. These data are used as the data of the working population in the new business format.

[0023] After removing the data on the new business types working in the dataset to be identified, the data on the unemployed population is obtained.

[0024] According to the present invention, a method for identifying unemployed individuals based on big data is provided, wherein the characteristics of the population include individuals who continuously log in to office websites;

[0025] Based on various specific population characteristics, data corresponding to the population characteristics are filtered and removed from the entire dataset to be identified, resulting in unemployment population data, including:

[0026] Obtain the URL login data of the user IDs in the complete set of data to be identified;

[0027] The data in the complete set of data to be identified that meets the criteria of logging into a specific website for multiple consecutive weeks and having more than a preset login count threshold per week are considered as data on people who continuously log into the office website.

[0028] After removing the data on people who continuously log into office websites from the entire dataset to be identified, the data on unemployed people is obtained.

[0029] The specific URL mentioned above is an office-type URL.

[0030] According to a big data-based method for identifying unemployed individuals provided by the present invention, the unemployed population data is further subdivided into high-probability unemployed population data, medium-probability unemployed population data, and unemployed population data with no job-seeking intention based on the type of APP corresponding to the user ID, APP usage time, APP data usage, call origination and reception data, and QR code scanning data.

[0031] Obtain the APP type, APP usage time, APP data usage, call origination and reception data, and QR code scanning data corresponding to the user ID in the unemployment population data;

[0032] Data that simultaneously meets the first, second, and third preset conditions from the unemployment population data will be considered as the high-probability unemployment population data.

[0033] The data of the unemployed population that meets any one of the first preset condition, the second preset condition, and the third preset condition shall be used as the medium probability unemployed population data.

[0034] The data of the unemployed population excluding the data of the high-probability unemployed population and the data of the medium-probability unemployed population is taken as the data of the unemployed population with no intention to seek employment.

[0035] The first preset condition includes: the APP type belongs to a specific APP type set, the average daily usage time of the APP is greater than a preset usage time threshold, and the average daily data usage of the APP is greater than a preset data usage threshold.

[0036] The second preset condition includes: during the statistical period, the caller and called party data have made job application calls;

[0037] The third preset condition includes: the scanned data contains a QR code for the job fair within the statistical period.

[0038] The present invention also provides a device for identifying unemployed individuals based on big data, comprising:

[0039] The data processing module is used to collect user mobile phone signaling data and process the user mobile phone signaling data to obtain the complete set of data to be identified;

[0040] The data removal module is used to filter and remove data corresponding to various specific population characteristics from the entire set of data to be identified, thereby obtaining unemployment population data.

[0041] The data segmentation module is used to segment the unemployment data into high-probability unemployment data, medium-probability unemployment data, and unemployment data with no job-seeking intention, based on the type of APP corresponding to the user ID, APP usage time, APP data usage, call origination and reception data, and QR code scanning data.

[0042] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the big data-based unemployment identification method described above.

[0043] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the big data-based method for identifying unemployed people as described above.

[0044] This invention provides a method, apparatus, device, and medium for identifying unemployed individuals based on big data. It obtains unemployed population data by preprocessing and filtering user mobile phone signaling data to extract specific population types. Then, it further subdivides the unemployed population data based on APP type, APP usage time, APP data usage, caller and recipient data, and QR code scanning data. Based on big data analysis methods, this invention introduces more dimensions of mobile phone signaling data. By classifying and subdividing the multi-dimensional mobile phone signaling data according to APP type, APP usage time, APP data usage, caller and recipient data, and QR code scanning data, the unemployed population data is further subdivided into high-probability unemployed population data, medium-probability unemployed population data, and unemployed population data with no job-seeking intention. This further analyzes and accurately identifies unemployed population data with and without job-seeking needs, thereby improving the coverage, accuracy, and precision of big data-based identification of the unemployment status of individuals. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0046] Figure 1 This is one of the flowcharts of the big data-based unemployment identification method provided by the present invention;

[0047] Figure 2 This is the second flowchart of the big data-based method for identifying unemployed people provided by the present invention;

[0048] Figure 3 This is the third flowchart of the big data-based method for identifying unemployed people provided by the present invention;

[0049] Figure 4 This is the fourth flowchart of the big data-based method for identifying unemployed people provided by the present invention;

[0050] Figure 5 This is the fifth flowchart of the big data-based method for identifying unemployed people provided by the present invention;

[0051] Figure 6 This is the sixth flowchart of the big data-based unemployment identification method provided by the present invention;

[0052] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0054] The following is combined with Figures 1-6 This invention describes a big data-based method for identifying unemployed individuals.

[0055] Please refer to Figure 1 The present invention proposes a method for identifying unemployed individuals based on big data, comprising:

[0056] Step 10: Collect user mobile phone signaling data and process the user mobile phone signaling data to obtain the complete set of data to be identified;

[0057] This involves collecting user mobile phone signaling data, which includes user location data, user app usage data, user identity data, and user call detail record (CDR) data. Specifically, user location data refers to the geographical location of the signal base station the user connects to at each moment while using their mobile phone, approximating the user's current location; user app usage data includes the types and names of mobile apps used by the user, as well as corresponding usage time and data usage records; user identity data includes the user's registered identity information such as age and gender; and user CDR data includes the user's call time, duration, and calling / received call details.

[0058] Next, the user's location information, app usage data, identity data, and call detail record (CDR) data from the user's mobile phone signaling data are cleaned to provide accurate information on the user's employment status from different dimensions. After data cleaning and dimensionality reduction, the following complete data set is obtained: The fields of the complete data set include: User ID, Age, Gender, Group Network Affiliation, Base Station ID, Base Station Duration, Base Station Latitude and Longitude, App Type, App Usage Time, App Data Usage, Caller / Called User ID, and Call Duration.

[0059] Further, based on the user's mobile phone signaling data, the complete set of data to be identified is determined, including: obtaining the user's recent mobile phone signaling data, and filling in the missing data in the user's mobile phone signaling data according to the mean of the data distribution of the user's mobile phone signaling data; using the random forest algorithm to rank all of the user's behavioral features by importance, and selecting the user's mobile phone signaling data corresponding to the top several features ranked by identity authentication contribution value as the complete set of data to be identified.

[0060] Data cleaning addresses common data gaps by using the mean method, obtaining user data from other recent time periods and filling in the missing data with the mean of their data distributions. Dimensionality reduction employs a random forest algorithm, ranking all user behavioral features by importance and selecting the top n features with the highest contribution to identity authentication.

[0061] The fields of the complete set of data to be identified include: user ID, age, gender, network affiliation, base station ID, base station duration, base station latitude and longitude, APP type, APP usage time, APP data usage, caller / called user ID, and call duration.

[0062] Step 20: Based on various specific population characteristics, filter and remove data corresponding to the population characteristics from the entire set of data to be identified to obtain unemployed population data;

[0063] Step 30: Based on the type of APP corresponding to the user ID, APP usage time, APP data usage, call origination and reception data, and QR code scanning data, the unemployment data is further subdivided into high-probability unemployment data, medium-probability unemployment data, and unemployment data with no job-seeking intention.

[0064] Among them, the specific characteristics of the population are those who do not currently need to consider employment issues. These characteristics include: people of non-employment age, confirmed group users, college students, people working in fixed locations, people working in new business formats, people who frequently enter the office, and people who frequently log in to the office website.

[0065] In this embodiment, based on various specific forms of population characteristics, data corresponding to specific forms of population characteristics in the entire dataset to be identified are filtered out. This is to filter and remove population characteristics in the entire dataset to be identified that do not require consideration of employment issues, thereby obtaining data on unemployed people who need to consider employment issues. This achieves accurate identification of unemployed people with employment needs based on the characteristics of people who do not require consideration of employment issues.

[0066] Secondly, based on the type of APP corresponding to the user ID, APP usage time, APP data usage, call data, and QR code scanning data, the unemployment data is further subdivided into high-probability unemployment data, medium-probability unemployment data, and unemployment data with no job-seeking intention. Further analysis is then conducted to accurately identify unemployment data with employment needs and those without job-seeking needs.

[0067] This invention provides a big data-based method for identifying unemployed individuals. It preprocesses and filters user mobile phone signaling data to extract specific demographic groups, resulting in unemployed population data. Then, it further subdivides this data based on APP type, APP usage time, APP data usage, caller and recipient data, and QR code scanning data. This invention, based on big data analysis methods, introduces more dimensions of mobile phone signaling data. By classifying and subdividing the multi-dimensional mobile phone signaling data according to APP type, APP usage time, APP data usage, caller and recipient data, and QR code scanning data, the unemployed population data is further subdivided into high-probability unemployed individuals, medium-probability unemployed individuals, and unemployed individuals with no job-seeking intentions. This further analyzes and accurately identifies unemployed individuals with and without job-seeking needs, thereby improving the coverage, accuracy, and precision of big data-based identification of the unemployment status of individuals.

[0068] In one possible embodiment, please refer to Figure 2 The demographic characteristics include non-employed individuals, confirmed group users, and university students.

[0069] Step 20: Based on various specific population characteristics, filter and remove data corresponding to the population characteristics from the entire dataset to be identified, obtaining unemployment population data, including:

[0070] Step 201: Using the age field, remove data whose age field in the entire dataset to be identified falls within the non-employment age range, so that the age range of the target object is within the preset age range;

[0071] Step 202: Using the group network affiliation field, user data with group network attributes, data inferred to be college students, and user data with rural household registration are removed to obtain the unemployment population data.

[0072] In this embodiment, the specific process of removing non-employment age groups, confirmed group users, and college students from the entire dataset to be identified is as follows: some non-target objects are removed from the entire dataset. The entire dataset is input, and user data whose age field falls within the non-employment age range is removed by the age field, so that the age range of the target objects falls within the employment age range. Furthermore, users with group network attributes, those inferred to be college students, and those whose household registration is in rural areas are removed by the group network affiliation field. After removing the above, the results are output to obtain the unemployment population data.

[0073] Among them, the non-employment age range meets the following criteria: male: (A1, B1) and female: (A0, B0). The values ​​of A0, B0, A1, and B1 can be set by relevant international and domestic unemployment definitions, general experience, or tests.

[0074] In this embodiment, the data of users who are not of working age are removed from the entire dataset to be identified by using the age field, and the data of users who are confirmed group users and college students are removed from the entire dataset to be identified by using the group network affiliation field. This accurately removes user data from the dataset to be identified that do not belong to the unemployed population or have no employment needs, thereby improving the accuracy of identifying unemployed people with employment needs.

[0075] In one possible embodiment, please refer to Figure 3 The population characteristics include people who work in fixed locations;

[0076] Step 20: Based on various specific population characteristics, filter and remove data corresponding to the population characteristics from the entire dataset to be identified, obtaining unemployment population data, including:

[0077] Step 211: Based on the latitude and longitude of the base station corresponding to the user ID, encode and cluster the entire set of data to be identified, so as to cluster several base station latitude and longitude into a geohash code;

[0078] Step 212: Arrange the geohash codes of each user ID in chronological order, and calculate the duration of continuous residence of each user ID under different geohash codes;

[0079] Step 213: Collect data that meets the following criteria: the continuous dwell time within a day is greater than a preset dwell time threshold, the number of days corresponding to the continuous dwell time being greater than the preset dwell time threshold is greater than a preset number of days threshold, and the number of geohashes corresponding to a single user ID with a continuous dwell time greater than the preset dwell time threshold within a day is greater than a preset number threshold. Collect data that meets the following criteria: fixed location working population data.

[0080] Step 214: Remove the data on fixed-location working population from the dataset to be identified to obtain data on unemployed population.

[0081] In this embodiment, the specific process for removing the fixed-location working population from the complete dataset to be identified is as follows: the latitude and longitude of the base stations corresponding to the user IDs in the complete dataset to be identified are encoded and clustered using geohash, that is, several base station latitude and longitudes within a suitable range are clustered into one geohash code; the geohash code of each user ID is arranged in chronological order. The following table shows:

[0082] User ID geohash time User 1 ab12cde 202206230955 User 1 ab12cde 202206231000 User 1 ab13cde 202206231005 User 2 ab14cde 202206230955 User 2 ab15cde 202206231000 User 2 ab16cde 202206231005 User 3 ab17cde 202206230955 User 3 ab18cde 202206231000 User 3 ab19cde 202206231005 User 3 ab20cde 202206231010

[0083] Then, calculate the duration of continuous residence for each user ID under different geohash encodings to obtain:

[0084] User ID geohash Continuous stay duration (h) User 1 ab12cde 2 User 1 ab13cde 1 User 2 ab14cde 5 User 2 ab15cde 2.1 User 2 ab16cde 0.4 User 3 ab17cde 1 User 3 ab18cde 0.5

[0085] Then, for the entire dataset to be identified, within a one-week time period, the following conditions are considered: Condition 1 = "continuous dwell time > preset dwell time threshold" on a certain day; Condition 2 = "number of days satisfying condition 1 > preset number of days threshold"; Condition 3 = "number of geohashes of a single user ID satisfying condition 1 > preset number threshold". The preset dwell time threshold, preset number of days threshold, and preset number threshold are determined through general experience and experimentation. Data sets in the entire dataset that simultaneously satisfy conditions 1, 2, and 3 are categorized as the fixed-location working population. Thus, data in the entire dataset that satisfy conditions 1, 2, and 3 are identified as the fixed-location working population. The fixed-location working population in the entire dataset is then removed to obtain the unemployed population data.

[0086] In this embodiment, user data of fixed-location working population in the dataset to be identified is accurately removed, thereby improving the accuracy of identifying data of unemployed people with employment needs.

[0087] In one possible embodiment, please refer to Figure 4 The demographic characteristics include new business workers, which refers to people who do not have a fixed office location but are presumed to be employed.

[0088] Step 20: Based on various specific population characteristics, filter and remove data corresponding to the population characteristics from the entire dataset to be identified, obtaining unemployment population data, including:

[0089] Step 221: Obtain the APP type, APP usage time, and APP data usage corresponding to the user ID in the complete set of data to be identified;

[0090] Step 222: The data in the complete set of data to be identified that the APP type belongs to the characteristic APP type set, the average daily number of times the APP is used within the period of the APP usage time is greater than a preset number of times threshold, and the average daily traffic usage within the period of the APP traffic usage is greater than a preset traffic threshold are taken as the new business format working population data.

[0091] Step 223: Remove the data on the new business workforce from the dataset to be identified to obtain the data on the unemployed population.

[0092] In this embodiment, the new business workforce refers to individuals without a fixed office location but presumed to be employed. The specific process for removing the new business workforce from the entire dataset to be identified is as follows: Obtain the APP type, APP usage time, and APP traffic corresponding to the user IDs in the entire dataset to be identified. The dataset that satisfies "Condition 4: The APP type corresponding to the user ID belongs to a specific APP type set NAME, the average daily usage within the period is greater than the preset usage threshold c0, and the average daily traffic within the period is greater than the preset traffic threshold f0" is recorded as the new business workforce data. The specific APP type set NAME, the preset usage threshold c0, and the preset traffic threshold f0 are set through general experience and testing. For example, the specific APP type set NAME may include "food delivery rider client, live streamer client, e-commerce merchant client, ride-hailing driver client, courier client, etc."

[0093] In this embodiment, data on new business types working in the dataset to be identified is accurately removed, thereby improving the accuracy of identifying data on unemployed people with employment needs.

[0094] In one possible embodiment, please refer to Figure 5 The population characteristics include people who continuously log in to office websites; Step 20: Based on various specific forms of population characteristics, filter and remove data corresponding to the population characteristics from the entire set of data to be identified to obtain unemployment population data, including:

[0095] Step 241: Obtain the URL login data of the user IDs in the complete set of data to be identified;

[0096] Step 242: Select the data in the complete set of data to be identified that meets the criteria of logging into a specific website for multiple consecutive weeks and logging in more than a preset number of times per week as data of people who continuously log into the office website.

[0097] Step 243: Remove the data of people who continuously log in to the office website from the entire set of data to be identified to obtain the data of unemployed people;

[0098] The specific URL mentioned above is an office-type URL.

[0099] In this embodiment, the specific process of removing users who continuously log in to office websites from the entire dataset to be identified is as follows: Obtain the website login data of users in the entire dataset to be identified, and record the dataset that meets the following conditions in the entire dataset to be identified: Condition 8 = "User ID logs in to a specific website NAME4 10 times per week" and Condition 9 = "Consistently meets condition 8 for m0 weeks" as the data of users who continuously log in to office websites. The settings of the specific website NAME4, the preset login number threshold l0, and m0 are confirmed through general experience and testing. The specific website NAME4 can be an office-type website.

[0100] In this embodiment, the system accurately removes individuals who continuously log into office websites from the dataset to be identified, thereby improving the accuracy of identifying data on unemployed individuals with employment needs.

[0101] Furthermore, the specific process for removing data on people continuously entering workplaces from the entire dataset to be identified is as follows: Obtain the QR code data corresponding to the user IDs in the entire dataset to be identified. Record the datasets that satisfy: Condition 5 = "The QR code type corresponding to the user ID belongs to a specific QR code type set NAME2, and the average daily usage frequency > g0" and Condition 6 = "The number of days satisfying condition 5 > h0" as data on people continuously entering workplaces. The settings for NAME2, g0, and h0 are confirmed through general experience and testing; NAME2 is generally the "workplace entry QR code". Afterward, remove the data on people continuously entering workplaces from the entire dataset to be identified to obtain data on unemployed individuals. This achieves accurate removal of data on people continuously entering workplaces from the dataset to be identified, improving the accuracy of identifying unemployed individuals with employment needs.

[0102] In one possible embodiment, please refer to Figure 6 Step 30: Based on the APP type corresponding to the user ID, APP usage time, APP data usage, call origination and reception data, and QR code scanning data, the unemployment data is further subdivided into high-probability unemployed population data, medium-probability unemployed population data, and unemployed population data with no job-seeking intention, including:

[0103] Step 31: Obtain the APP type, APP usage time, APP data usage, call origination and reception data, and QR code scanning data corresponding to the user ID in the unemployment population data.

[0104] Step 32: Select the data in the unemployment population data that simultaneously meets the first preset condition, the second preset condition, and the third preset condition as the high-probability unemployment population data;

[0105] Step 33: Select the data from the unemployment population data that meets any one of the first preset condition, the second preset condition, and the third preset condition as the medium probability unemployment population data;

[0106] Step 34: The data of the unemployed population excluding the data of the high-probability unemployed population and the data of the medium-probability unemployed population is used as the data of the unemployed population with no job-seeking intention;

[0107] The first preset condition includes: the APP type belongs to a specific APP type set, the average daily usage time of the APP is greater than a preset usage time threshold, and the average daily data usage of the APP is greater than a preset data usage threshold.

[0108] The second preset condition includes: during the statistical period, the caller and called party data have made job application calls;

[0109] The third preset condition includes: the scanned data contains a QR code for the job fair within the statistical period.

[0110] In this embodiment, the obtained unemployment data is defined as "potential unemployment data". The unemployment data is further processed to obtain subdivided data on high-probability unemployment, medium-probability unemployment, and unemployment with no job-seeking intention.

[0111] Specifically, the data includes the type of app corresponding to the user ID, app usage time, app data usage, call origination and reception data, and QR code scanning data from the unemployment population data. The dataset that simultaneously meets the following conditions is denoted as "set H": "first preset condition: the app type corresponding to the user ID belongs to a specific app type set NAME5, the average daily usage per week > preset usage threshold n1, and the average daily data usage per week > preset data usage threshold o1", "second preset condition: job search calls were made within the statistical period", and "job fair entry QR code was scanned within the statistical period". Data that meets any one of the first, second, and third preset conditions is used as medium probability unemployment population data.

[0112] Among them, the specific APP type set NAME5, the preset usage threshold n1, the preset traffic threshold o1, and the QR code data settings were confirmed through general experience and testing. NAME5 is selected as a job search APP, the job search phone number is selected as the local human resources and social security unemployment hotline, the caller and callee include job search phone numbers, and the QR code data includes the job fair entry QR code.

[0113] In this embodiment, the mobile signaling data of the unemployed population is divided and subdivided into multi-dimensional categories such as APP type, APP usage time, APP data usage, call origination and reception data, and QR code scanning data. The unemployed population data is further subdivided into high-probability unemployed population data, medium-probability unemployed population data, and unemployed population data with no job-seeking intention. This enables further analysis and accurate identification of unemployed population data with employment needs and those without job-seeking needs, thereby improving the coverage, accuracy, and precision of big data identification of the unemployment status of individuals.

[0114] The following describes the big data-based unemployment identification device provided by the present invention. The big data-based unemployment identification device described below and the big data-based unemployment identification method described above can be referred to and correspond to each other.

[0115] The data processing module is used to collect user mobile phone signaling data and process the user mobile phone signaling data to obtain the complete set of data to be identified;

[0116] The data removal module is used to filter and remove data corresponding to various specific population characteristics from the entire set of data to be identified, thereby obtaining unemployment population data.

[0117] The data segmentation module is used to segment the unemployment data into high-probability unemployment data, medium-probability unemployment data, and unemployment data with no job-seeking intention, based on the type of APP corresponding to the user ID, APP usage time, APP data usage, call origination and reception data, and QR code scanning data.

[0118] Furthermore, the demographic characteristics include non-employed individuals, confirmed group users, and university students;

[0119] The data removal module is also used for:

[0120] By using the age field, data in the non-employment age range of the entire dataset to be identified are removed, so that the age range of the target object is within the preset age range;

[0121] By using the group network affiliation field, user data with group network attributes, data inferred to be college students, and user data with rural household registration are removed to obtain the unemployment population data.

[0122] Furthermore, the population characteristics include people working in fixed locations;

[0123] The data removal module is also used for:

[0124] Based on the latitude and longitude of the base station corresponding to the user ID, the entire set of data to be identified is encoded and clustered so that the latitude and longitude of several base stations are clustered into a geohash code;

[0125] Each user ID is arranged in chronological order of its geohash encoding, and the duration of continuous residence for each user ID under different geohash encodings is calculated.

[0126] The data set to be identified is defined as the data of fixed-location working population that meets the following criteria: the continuous dwell time within a day is greater than a preset dwell time threshold, the number of days corresponding to the continuous dwell time being greater than the preset dwell time threshold is greater than a preset number of days threshold, and the number of geohashes corresponding to a single user ID with a continuous dwell time greater than the preset dwell time threshold within a day is greater than a preset number threshold.

[0127] By removing the data on people working at fixed locations from the dataset to be identified, we obtain data on unemployed people.

[0128] Furthermore, the demographic characteristics include new business employment groups, which refer to people who do not have a fixed office location but are presumed to be employed.

[0129] The data removal module is also used for:

[0130] Obtain the APP type, APP usage time, and APP data usage corresponding to the user ID in the complete set of data to be identified;

[0131] The data corresponding to the following in the complete set of data to be identified are: the APP type belongs to the characteristic APP type set; the average daily number of times the APP is used within the period of the APP usage time is greater than a preset number of times threshold; and the average daily traffic usage within the period of the APP traffic usage is greater than a preset traffic threshold. These data are used as the data of the working population in the new business format.

[0132] After removing the data on the new business types working in the dataset to be identified, the data on the unemployed population is obtained.

[0133] Furthermore, the demographic characteristics include individuals who continuously log into office websites;

[0134] The data removal module is also used for:

[0135] Obtain the URL login data of the user IDs in the complete set of data to be identified;

[0136] The data in the complete set of data to be identified that meets the criteria of logging into a specific website for multiple consecutive weeks and having more than a preset login count threshold per week are considered as data on people who continuously log into the office website.

[0137] After removing the data on people who continuously log into office websites from the entire dataset to be identified, the data on unemployed people is obtained.

[0138] The specific URL mentioned above is an office-type URL.

[0139] Furthermore, the data segmentation module is also used for:

[0140] Obtain the APP type, APP usage time, APP data usage, call origination and reception data, and QR code scanning data corresponding to the user ID in the unemployment population data;

[0141] Data that simultaneously meets the first, second, and third preset conditions from the unemployment population data will be considered as the high-probability unemployment population data.

[0142] The data of the unemployed population that meets any one of the first preset condition, the second preset condition, and the third preset condition shall be used as the medium probability unemployed population data.

[0143] The data of the unemployed population excluding the data of the high-probability unemployed population and the data of the medium-probability unemployed population is taken as the data of the unemployed population with no intention to seek employment.

[0144] The first preset condition includes: the APP type belongs to a specific APP type set, the average daily usage time of the APP is greater than a preset usage time threshold, and the average daily data usage of the APP is greater than a preset data usage threshold.

[0145] The second preset condition includes: during the statistical period, the caller and called party data have made job application calls;

[0146] The third preset condition includes: the scanned data contains a QR code for the job fair within the statistical period.

[0147] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840. The processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a big data-based method for identifying unemployed individuals. This method includes: collecting user mobile phone signaling data and processing the user mobile phone signaling data to obtain a complete set of data to be identified; filtering and removing data corresponding to various specific population characteristics from the complete set of data to be identified to obtain unemployed population data; and further subdividing the unemployed population data into high-probability unemployed population data, medium-probability unemployed population data, and unemployed population data with no job-seeking intention based on the type of APP corresponding to the user ID, APP usage time, APP data usage, call originating and receiving data, and QR code scanning data.

[0148] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0149] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the big data-based unemployment identification method provided by the above methods. The method includes: collecting user mobile phone signaling data and processing the user mobile phone signaling data to obtain a complete set of data to be identified; filtering and removing data corresponding to the population characteristics from the complete set of data to be identified according to various specific population characteristics to obtain unemployment population data; and subdividing the unemployment population data into high-probability unemployment population data, medium-probability unemployment population data, and unemployment population data with no job-seeking intention according to the type of APP corresponding to the user ID, APP usage time, APP data usage, call originating and receiving data, and QR code scanning data.

[0150] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the big data-based unemployment identification method provided by the above methods. The method includes: collecting user mobile phone signaling data and processing the user mobile phone signaling data to obtain a complete set of data to be identified; filtering and removing data corresponding to various specific population characteristics from the complete set of data to be identified to obtain unemployment data; and further subdividing the unemployment data into high-probability unemployment data, medium-probability unemployment data, and unemployment data with no job-seeking intention based on the type of APP corresponding to the user ID, APP usage time, APP data usage, call originating and receiving data, and QR code scanning data.

[0151] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0152] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying unemployed individuals based on big data, characterized in that, include: Collect user mobile phone signaling data and process the user mobile phone signaling data to obtain the complete set of data to be identified; Based on various specific population characteristics, data corresponding to the population characteristics are filtered and removed from the entire set of data to be identified to obtain unemployment population data; Obtain the APP type, APP usage time, APP data usage, call origination and reception data, and QR code scanning data corresponding to the user ID in the unemployment population data; Data that simultaneously meets the first, second, and third preset conditions from the unemployment population data will be considered as high-probability unemployment population data. Data from the unemployment population that meets any one of the first preset condition, the second preset condition, and the third preset condition will be considered as medium-probability unemployment population data. The data of the unemployed population excluding the data of the high-probability unemployed population and the data of the medium-probability unemployed population is regarded as the data of the unemployed population with no intention of seeking employment. The first preset condition includes: the APP type belongs to a specific APP type set, the average daily usage time of the APP is greater than a preset usage time threshold, and the average daily data usage of the APP is greater than a preset data usage threshold. The second preset condition includes: during the statistical period, the caller and called party data have made job application calls; The third preset condition includes: the scanned data contains a QR code for the job fair within the statistical period.

2. The method for identifying unemployed individuals based on big data according to claim 1, characterized in that, The demographic characteristics include non-employed individuals, confirmed group users, and university students. Based on various specific population characteristics, data corresponding to the population characteristics are filtered and removed from the entire dataset to be identified, resulting in unemployment population data, including: By using the age field, data in the non-employment age range of the entire dataset to be identified are removed, so that the age range of the target object is within the preset age range; By using the group network affiliation field, user data with group network attributes, data inferred to be college students, and user data with rural household registration are removed to obtain the unemployment population data.

3. The method for identifying unemployed individuals based on big data according to claim 1, characterized in that, The population characteristics include people who work in fixed locations; Based on various specific population characteristics, data corresponding to the population characteristics are filtered and removed from the entire dataset to be identified, resulting in unemployment population data, including: Based on the latitude and longitude of the base station corresponding to the user ID, the entire set of data to be identified is encoded and clustered so that the latitude and longitude of several base stations are clustered into a geohash code; Each user ID is arranged in chronological order of its geohash encoding, and the duration of continuous residence for each user ID under different geohash encodings is calculated. The data set to be identified is defined as the data of fixed-location working population that meets the following criteria: the continuous dwell time within a day is greater than a preset dwell time threshold, the number of days corresponding to the continuous dwell time being greater than the preset dwell time threshold is greater than a preset number of days threshold, and the number of geohashes corresponding to a single user ID with a continuous dwell time greater than the preset dwell time threshold within a day is greater than a preset number threshold. By removing the data on people working at fixed locations from the dataset to be identified, we obtain data on unemployed people.

4. The method for identifying unemployed individuals based on big data according to claim 1, characterized in that, The demographic characteristics include new business workers, who are people who do not have a fixed office location but are presumed to be employed. Based on various specific population characteristics, data corresponding to the population characteristics are filtered and removed from the entire dataset to be identified, resulting in unemployment population data, including: Obtain the APP type, APP usage time, and APP data usage corresponding to the user ID in the complete set of data to be identified; The data corresponding to the following in the complete set of data to be identified are: the APP type belongs to the characteristic APP type set; the average daily number of times the APP is used within the period of the APP usage time is greater than a preset number of times threshold; and the average daily traffic usage within the period of the APP traffic usage is greater than a preset traffic threshold. These data are used as the data of the working population in the new business format. After removing the data on the new business types working in the dataset to be identified, the data on the unemployed population is obtained.

5. The method for identifying unemployed individuals based on big data according to claim 1, characterized in that, The demographic characteristics include individuals who consistently log into the office website. Based on various specific population characteristics, data corresponding to the population characteristics are filtered and removed from the entire dataset to be identified, resulting in unemployment population data, including: Obtain the URL login data of the user IDs in the complete set of data to be identified; The data in the complete set of data to be identified that meets the criteria of logging into a specific website for multiple consecutive weeks and having more than a preset login count threshold per week are considered as data on people who continuously log into the office website. After removing the data on people who continuously log into office websites from the entire dataset to be identified, the data on unemployed people is obtained. The specific URL mentioned above is an office-type URL.

6. A device for identifying unemployed individuals based on big data, characterized in that, include: The data processing module is used to collect user mobile phone signaling data and process the user mobile phone signaling data to obtain the complete set of data to be identified; The data removal module is used to filter and remove data corresponding to various specific population characteristics from the entire set of data to be identified, thereby obtaining unemployment population data. The data segmentation module is used to obtain the APP type, APP usage time, APP data usage, call origination and reception data, and QR code scanning data corresponding to the user ID in the unemployment population data; and to identify the data in the unemployment population data that simultaneously meets the first preset condition, the second preset condition, and the third preset condition as high-probability unemployment population data. Data from the unemployment population that meets any one of the first preset condition, the second preset condition, and the third preset condition will be considered as medium-probability unemployment population data. The data of the unemployed population excluding the data of the high-probability unemployed population and the data of the medium-probability unemployed population is regarded as the data of the unemployed population with no intention of seeking employment. The first preset condition includes: the APP type belongs to a specific APP type set, the average daily usage time of the APP is greater than a preset usage time threshold, and the average daily data usage of the APP is greater than a preset data usage threshold; the second preset condition includes: the caller and called data have made job search calls within the statistical period; the third preset condition includes: the scanned data contains a job fair entry QR code within the statistical period.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the big data-based method for identifying unemployed people as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the big data-based method for identifying unemployed people as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Big data-based unemployed population dynamic monitoring method

    CN108733774A

  • Crowd occupational type acquisition method and system based on mobile phone signaling, and storage medium

    CN115002680A