A method for extracting key domain data information from big data

By analyzing the training data of new employees, an abnormal impact identification and extraction set is generated, which solves the problem of poor extraction and processing of information of new employees training big data, and achieves more reliable analysis and evaluation.

CN119337272BActive Publication Date: 2025-08-15BEIXI INTELLIGENT FUTURE (SHANGHAI) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411361974.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-08-15
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

In the existing plan, the extraction and processing of new employee training big data information is poor and the analysis and evaluation is poor.

Method used

By obtaining training data for new employees in different positions, analyzing the reasons for unemployment and failed data items, generating an abnormal impact identification and extraction set, and pushing it to supervisors for targeted training exception extraction analysis.

Benefits of technology

Improve the diversity and reliability of data extraction processing for unemployed reasons and failed data items, and provide reliable data support for subsequent training exception extraction analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119337272B_ABST
    Figure CN119337272B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for extracting data information from key domains of big data information, which belongs to the technical field of data information processing. The method is used to solve the technical problems of poor extraction and processing effects and poor analysis and evaluation effects of new employee training big data information in existing solutions. The method processes and analyzes the overall training status of new employees corresponding to different positions based on all training data to obtain several target positions and several target training data. The method processes and extracts data in the dimension of reasons for not joining the job for all target positions obtained through processing, and performs data calculation and analysis mark combination on the extracted data corresponding to different reasons for not joining the job according to the screening and classification to obtain a first abnormal impact identification extraction set. The method processes and extracts data in the dimension of failed data items for all target training data obtained through processing to obtain a second abnormal impact identification extraction set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data information processing, and in particular to a method for extracting key domain data information from big data. Background Art

[0002] The key domains of big data information refer to areas or data sets that are crucial in big data analysis and applications. These key domains usually contain data that have a significant impact on business decisions, market analysis, user behavior understanding, etc.

[0003] The implementation steps of data information extraction from key domains of big data information can be divided into several main stages, from demand analysis to data application. These steps ensure the effective collection, processing and utilization of data. When the existing data information extraction scheme for key domains of big data information is implemented, there is a single supervision and processing method and a single analysis and evaluation method for the information extraction of new employee training and work, which leads to poor extraction and processing effects of big data information for new employee training and poor analysis and evaluation effects. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for extracting key domain data information from big data information, which is used to solve the technical problems of poor extraction and processing effect and poor analysis and evaluation effect of new employee training big data information in existing solutions.

[0005] The purpose of the present invention can be achieved through the following technical solutions:

[0006] A method for extracting key domain data information from big data, comprising:

[0007] Obtain the training data corresponding to all new employees in different positions, and process and analyze the overall training status of new employees corresponding to different positions based on all training data to obtain several target positions and several target training data;

[0008] Data processing and extraction of the non-employment reason dimension are performed on all target positions obtained, and the first abnormal impact identification extraction set of the corresponding dimension is obtained;

[0009] Performing data processing and extraction on the failed data item dimension for all the target training data obtained through processing to obtain a second abnormal impact identification extraction set of the corresponding dimension;

[0010] The first abnormal impact identification extraction set and the second abnormal impact identification extraction set obtained by processing are pushed to the supervisors respectively, and targeted training abnormal extraction analysis prompts are implemented.

[0011] Preferably, when implementing different training supervision processes for all training data corresponding to different positions, the name, gender, education background, training start time, training end time, and training results of the corresponding trainees in all training data for different positions are obtained in sequence;

[0012] When processing and extracting the training result dimension for the training data corresponding to all new employees in different positions, the training results in different training data of different positions are obtained in sequence and analyzed traversally;

[0013] If the training result is normal employment, a local training success instruction is generated and the total number of successful trainees CPi of the corresponding position is increased by one; i is a different position, i = 1, 2, 3, ..., n; n is a positive integer;

[0014] If the training result is failure to join the job, a local training failure instruction will be generated and the total number of training failures SPi for the corresponding position will be increased by one.

[0015] Preferably, when processing and analyzing the overall training status of new employees corresponding to different positions, the total number of successful training CPi and the total number of failed training SPi corresponding to different positions are calculated using a training status identification function, and the overall training identification ZPi corresponding to different positions is output;

[0016] Among them, the expression of the training state recognition function is Where PB is the standard value of overall training requirements corresponding to different positions;

[0017] The overall training flag contains a value of 0 or 1, indicating whether the overall training status of new employees in the corresponding position is normal or abnormal;

[0018] According to the overall training flag with a value of 1, the corresponding position is marked as the target position, and the training data belonging to the training results of all target positions that have not yet been employed are marked as target training data.

[0019] Preferably, the reasons for non-employment corresponding to the training results of non-employment in all target positions are obtained, and all reasons for non-employment are classified and marked;

[0020] If the reason for non-joining is due to personal reasons of the employee, the non-joining employee will be marked as the first non-joining employee;

[0021] If the reason for non-employment is not due to personal reasons, the employee will be marked as the second non-employed employee;

[0022] Obtain all non-joining reasons corresponding to all second non-joining employees, sort and combine all non-joining reasons, deduplicate all identical non-joining reasons in the sorted combinations, and update the total number of occurrences of different non-joining reasons to obtain a non-joining reasons statistics table.

[0023] Preferably, when processing the data on the influence of abnormal reasons corresponding to different reasons for not joining the job according to the statistical table of reasons for not joining the job, the total number of occurrences of different reasons for not joining the job is sequentially calculated using the formula Calculate and obtain the corresponding abnormal cause impact DYj; where NYj is the total number of occurrences of different reasons for non-employment; j is a different reason for non-employment, j = 1, 2, 3, ..., m; m is a positive integer; B is the standard value of the abnormal cause impact.

[0024] Preferably, when performing data analysis on the impact of abnormal reasons to determine the impact of abnormal reasons corresponding to different reasons for non-employment;

[0025] If the impact of the abnormal reason is 0, a normal impact instruction of the abnormal reason is generated, and the corresponding non-employment reason is marked as a normal non-employment reason;

[0026] If the impact of the abnormal reason is not 0, a special impact instruction of the abnormal reason is generated, and the corresponding non-employment reason is marked as a special non-employment reason;

[0027] All the special reasons for non-employment marked by the analysis are sorted and combined to obtain the first abnormal impact identification and extraction set.

[0028] Preferably, data statistics are performed on different data items in all target training data to obtain corresponding statistical sequences of failed genders, failed educational backgrounds, and failed training durations;

[0029] When performing abnormal identification and extraction of different elements in different statistical sequences, the different elements in different statistical sequences are sequentially analyzed by the formula Calculate and obtain the element influence ratio YLk of the corresponding element; where YZk is the value of different elements in different statistical sequences; k is the different elements in different statistical sequences, k = 1, 2, 3, ..., p; p is a positive integer, which is the total number of elements in the statistical sequence.

[0030] Preferably, the difference in the element influence ratio between different elements in the statistical sequence is calculated and marked as the element influence difference YC. When determining the influence status of the abnormal data items between different elements based on the element influence difference, the influence differences of different elements in the statistical sequence are calculated by the element influence identification function and the corresponding abnormal data item influence degree YX is output;

[0031] Among them, the expression of the element influence identification function is Where, YC0 is the standard element influence difference;

[0032] The abnormal data item impact degree includes a value of 0 or 1, indicating whether the abnormal data item impact status between corresponding elements is normal or abnormal.

[0033] Preferably, according to the abnormal data item influence degree of 1, the maximum value element in the two calculation elements is marked as a high abnormal impact data item;

[0034] All high abnormal impact data items are sorted and combined to obtain the second abnormal impact identification and extraction set.

[0035] Preferably, the failure gender statistical sequence includes the total number of female students' training failures and the total number of male students' training failures;

[0036] The statistical series of failed academic qualifications includes the total number of failed college training, the total number of failed undergraduate training, and the total number of failed master's training;

[0037] The failed training duration statistical sequence includes the total number of failures during the first duration, the total number of failures during the second duration, and the total number of failures during the third duration.

[0038] Compared with the existing solutions, the present invention achieves the following beneficial effects:

[0039] The present invention processes and analyzes the overall training status of new employees corresponding to different positions based on all training data, and obtains several target positions and several target training data. It can not only obtain the overall training status of new employees corresponding to different positions, but also provide reliable supervision and analysis data support for subsequent data processing and extraction in different dimensions.

[0040] The present invention processes and extracts data from the dimension of reasons for non-employment for all target positions obtained through processing, and performs data calculation and analysis marking on the extracted data corresponding to different reasons for non-employment that have been screened and classified, thereby obtaining different conventional reasons for non-employment and special reasons for non-employment, extracts and combines all special reasons for non-employment, and obtains the first abnormal impact identification extraction set, thereby improving the diversity and reliability of the data extraction processing of the dimension of reasons for non-employment, and can also provide reliable data support for the dimension of extraction of reasons for non-employment for subsequent training abnormal extraction and analysis prompts.

[0041] The present invention implements data processing and extraction of failed data item dimensions on all target training data obtained through processing, and performs data calculation on different element influence differences in different statistical sequences in sequence through element influence identification functions and outputs corresponding abnormal data item influence degrees. The abnormal data item influence states between different abnormal data items are analyzed and screened through the abnormal data item influence degrees to obtain high abnormal influence data items, and all high abnormal influence data items are sorted and combined to obtain a second abnormal influence identification and extraction set, thereby improving the diversity and reliability of data extraction processing of failed data item dimensions, and also providing reliable failed data item dimension extraction data support for subsequent training abnormal extraction analysis prompts. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The present invention will be further described below with reference to the accompanying drawings.

[0043] Figure 1 This is a flowchart of a method for extracting key domain data information from big data information according to the present invention.

[0044] Figure 2 This is a flowchart for processing and extracting the training result dimension of the training data corresponding to all new employees in different positions in the present invention. DETAILED DESCRIPTION

[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0046] like Figure 1 As shown, the present invention is a method for extracting key domain data information from big data, comprising:

[0047] Obtain the training data corresponding to all new employees in different positions, and process and analyze the overall training status of new employees corresponding to different positions based on all training data to obtain several target positions and several target training data; including:

[0048] When implementing different training supervision methods for all training data corresponding to different positions, obtain the name, gender, education background, training start time, training end time, and training results of the corresponding trainees in all training data for different positions in sequence; the units corresponding to the start time and training end time are both days;

[0049] Among them, different positions include but are not limited to technical positions, sales positions and customer service positions;

[0050] like Figure 2 As shown, when processing and extracting the training result dimension for the training data corresponding to all new employees in different positions, the training results in different training data of different positions are obtained in sequence and traversed and analyzed;

[0051] Among them, the training results are normal employment or non-employment;

[0052] If the training result is normal employment, a local training success instruction is generated and the total number of successful trainees CPi of the corresponding position is increased by one; i is a different position, i = 1, 2, 3, ..., n; n is a positive integer;

[0053] If the training result is failure to join the company, a local training failure instruction will be generated and the total number of training failures SPi for the corresponding position will be increased by one;

[0054] When processing and analyzing the overall training status of new employees corresponding to different positions, the total number of successful training CPi and the total number of failed training SPi corresponding to different positions are calculated through the training status identification function, and the overall training identification ZPi corresponding to different positions is output;

[0055] Among them, the expression of the training state recognition function is Where PB is the standard value of the overall training requirements corresponding to different positions, which can be determined by the actual application requirements of the actual application scenario, or by the median value of all the overall training identifiers corresponding to different positions in history;

[0056] It should be noted that the overall training identifier is used to perform data calculations on all training result data corresponding to different positions, so as to digitally represent the overall training status of new employees corresponding to different positions;

[0057] The overall training flag contains a value of 0 or 1, indicating whether the overall training status of new employees in the corresponding position is normal or abnormal;

[0058] According to the overall training flag with a value of 1, the corresponding position is marked as the target position, and the training data of the training results of all target positions that have not yet been employed are marked as target training data;

[0059] In an embodiment of the present invention, the overall training status of new employees corresponding to different positions is processed and analyzed based on all training data to obtain several target positions and several target training data. This can not only obtain the overall training status of new employees corresponding to different positions, but also provide reliable supervision and analysis data support for subsequent data processing and extraction in different dimensions.

[0060] Data processing and extraction of the non-employment reason dimension is performed on all acquired target positions to obtain the first abnormal impact identification extraction set of the corresponding dimension; including:

[0061] Obtain the reasons for non-employment corresponding to the training results for all target positions, and classify and mark all reasons for non-employment;

[0062] The reasons for non-employment shall be determined by the trainers and training managers;

[0063] If the reason for non-joining is due to personal reasons of the employee, the non-joining employee will be marked as the first non-joining employee;

[0064] If the reason for non-employment is not due to personal reasons, the employee will be marked as the second non-employed employee;

[0065] In the embodiment of the present invention, by performing data analysis and classification marking on the reasons for non-employment, reliable screening data support can be provided for subsequent analysis of the impact of abnormal reasons;

[0066] Obtain all non-joining reasons corresponding to all second non-joining employees, sort and combine all non-joining reasons, remove duplicates from all identical non-joining reasons in the sorted combinations, and update the total number of occurrences of different non-joining reasons to obtain a non-joining reason statistics table;

[0067] In the embodiment of the present invention, by sorting and combining all the reasons for non-employment marked by screening and removing duplicates, a statistical table of reasons for non-employment is obtained, thereby improving the reliability and diversity of the non-employment reason extraction process;

[0068] When processing the data on the impact of abnormal reasons corresponding to different reasons for non-employment according to the statistical table of reasons for non-employment, the total number of occurrences of different reasons for non-employment is calculated in turn by the formula Calculate the corresponding abnormal cause impact DYj; where NYj is the total number of occurrences of different reasons for non-employment; j is a different reason for non-employment, j = 1, 2, 3, ..., m; m is a positive integer; B is the standard value of the abnormal cause impact; [*] is the rounding function, indicating that the maximum integer that does not exceed the real number * is obtained;

[0069] It should be noted that the abnormal cause impact is used to calculate the extracted data corresponding to different reasons for non-employment, so as to digitally represent the impact of abnormal causes corresponding to different reasons for non-employment;

[0070] When conducting data analysis on the impact of abnormal reasons to determine the impact of abnormal reasons corresponding to different reasons for non-employment;

[0071] If the impact of the abnormal reason is 0, a normal impact instruction of the abnormal reason is generated, and the corresponding non-employment reason is marked as a normal non-employment reason;

[0072] If the impact of the abnormal reason is not 0, a special impact instruction of the abnormal reason is generated, and the corresponding non-employment reason is marked as a special non-employment reason;

[0073] Sort and combine all the special reasons for non-employment marked by the analysis to obtain the first abnormal impact identification and extraction set;

[0074] In an embodiment of the present invention, data processing and extraction of the dimension of reasons for non-employment are implemented for all target positions obtained through processing, and data calculation and analysis marking are performed on the extracted data corresponding to different reasons for non-employment that are screened and classified, so as to obtain different conventional reasons for non-employment and special reasons for non-employment, and extract and combine all special reasons for non-employment to obtain the first abnormal impact identification extraction set, thereby improving the diversity and reliability of the data extraction processing of the dimension of reasons for non-employment, and at the same time providing reliable data support for the dimension of extraction of reasons for non-employment for subsequent training abnormal extraction and analysis prompts.

[0075] Perform data processing and extraction on the failed data item dimension for all target training data obtained through processing to obtain a second abnormal impact identification extraction set of the corresponding dimension; including:

[0076] Perform data statistics on different data items in all target training data to obtain the corresponding statistical series of failed gender, failed academic qualifications, and failed training duration;

[0077] Among them, the failure gender statistics series includes the total number of female training failures and the total number of male training failures;

[0078] The statistical series of failed academic qualifications includes the total number of failed college training, the total number of failed undergraduate training, and the total number of failed master's training;

[0079] The failure training duration statistical sequence includes the total number of failures during the first duration, the total number of failures during the second duration, and the total number of failures during the third duration;

[0080] Specifically, the first duration may be 1-7 days, the second duration may be 8-14 days, and the third duration may be more than 14 days;

[0081] In the embodiment of the present invention, by processing and statistically analyzing the failed gender statistical series, failed academic record statistical series, and failed training duration statistical series for all target training data, the reliability and comprehensiveness of the extraction and processing of different data items can be effectively improved;

[0082] When performing abnormal identification and extraction of different elements in different statistical sequences, the different elements in different statistical sequences are sequentially analyzed by the formula Calculate and obtain the element influence ratio YLk of the corresponding element; where YZk is the value of different elements in different statistical sequences; k is different elements in different statistical sequences, k = 1, 2, 3, ..., p; p is a positive integer, which is the total number of elements in the corresponding statistical sequence;

[0083] Furthermore, the difference in the element influence ratio between different elements in the statistical sequence is calculated and marked as the element influence difference YC. When determining the influence status of abnormal data items between different elements based on the element influence difference, the influence difference of different elements in the statistical sequence is calculated by the element influence identification function and the corresponding abnormal data item influence degree YX is output;

[0084] Among them, the expression of the element influence identification function is Where YC0 is the standard element influence difference, which can be determined by the actual application requirements of the actual application scenario, or by the median value of the influence of all abnormal data items in the history corresponding to different abnormal data items;

[0085] It should be noted that the abnormal data item impact degree is used to perform data calculation and digital representation on the abnormal data item impact status between different abnormal data items;

[0086] The impact degree of abnormal data items contains a value of 0 or 1, indicating whether the abnormal data items between corresponding elements have an impact on the state of normal or abnormal;

[0087] According to the abnormal data item influence with a value of 1, the maximum value element in the two calculation elements is marked as a high abnormal impact data item;

[0088] For example, the total number of female and male training failures in the failed gender statistical sequence has an element influence of 47% and 53% respectively, and the element influence difference between the two elements is 6%. In this case, the element influence difference of 6% is not greater than the standard element influence difference of 20%. Therefore, the influence status of the abnormal data item between the total number of female and male training failures in the failed gender statistical sequence is determined to be normal.

[0089] If the corresponding element influence ratios are 37% and 63% respectively, the element influence difference between the two elements is 26%. In this case, the element influence difference of 26% is greater than the standard element influence difference of 20%. Therefore, the abnormal data item impact status between the total number of female training failures and the total number of male training failures contained in the failed gender statistical sequence is determined to be abnormal. In this case, the total number of male training failures with an element influence ratio of 63% is a high abnormal influence data item.

[0090] Sort and combine all high abnormal impact data items to obtain a second abnormal impact identification and extraction set;

[0091] In an embodiment of the present invention, data processing and extraction of failed data item dimensions are performed on all target training data obtained through processing, and data calculation is performed in sequence on different element influence differences in different statistical sequences through element influence identification functions, and the corresponding abnormal data item influence degrees are output. The abnormal data item influence states between different abnormal data items are analyzed and screened through the abnormal data item influence degrees to obtain high abnormal influence data items, and all high abnormal influence data items are sorted and combined to obtain a second abnormal influence identification and extraction set, thereby improving the diversity and reliability of the data extraction processing of the failed data item dimensions, and at the same time, providing reliable failed data item dimension extraction data support for subsequent training abnormal extraction analysis prompts.

[0092] The first abnormal impact identification extraction set and the second abnormal impact identification extraction set obtained by processing are pushed to the supervisors respectively, and targeted training abnormal extraction analysis prompts are implemented.

[0093] Among them, targeted training anomaly extraction analysis prompts are implemented based on all special reasons for non-employment in the first abnormal impact identification and extraction set and all high abnormal impact data items in the second abnormal impact identification and extraction set, which can provide reliable multi-dimensional training extraction data support for the implementation and management of subsequent new employee training.

[0094] In addition, the formulas involved in the above are all calculated by removing dimensions and taking their numerical values. They are a formula that is closest to the actual situation obtained by collecting a large amount of data and simulating it through simulation software.

[0095] In the several embodiments provided by the present invention, it should be understood that the disclosed methods can be implemented in other ways. For example, the above-described embodiments of the invention are merely illustrative. For example, the division of modules is only a logical function division, and other division methods may be used in actual implementation.

[0096] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of these modules may be selected to achieve the objectives of this embodiment based on actual needs.

[0097] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing module, each module may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or hardware plus software functional modules.

[0098] It is obvious to a person skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, but that the present invention can be implemented in other specific forms without departing from the essential characteristics of the present invention.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for extracting key domain data information from big data, characterized in that: include: Obtain the training data corresponding to all new employees in different positions, and process and analyze the overall training status of new employees corresponding to different positions based on all training data to obtain several target positions and several target training data; Data processing and extraction of the non-employment reason dimension are performed on all target positions obtained, and the first abnormal impact identification extraction set of the corresponding dimension is obtained; Performing data processing and extraction on the failed data item dimension for all the target training data obtained through processing to obtain a second abnormal impact identification extraction set of the corresponding dimension; Push the first abnormal impact identification and extraction set and the second abnormal impact identification and extraction set obtained by processing to the supervisor respectively, and implement targeted training abnormal extraction and analysis prompts; Among them, data statistics are performed on different data items in all target training data to obtain the corresponding statistical series of failed gender, failed academic qualifications, and failed training duration; When performing anomaly identification and extraction of different elements in different statistical sequences, the element influence ratio of the corresponding elements is obtained by calculating the different elements in different statistical sequences in turn; Calculate the difference in the element influence ratio between different elements in the statistical sequence and mark it as the element influence difference. When determining the influence status of abnormal data items between different elements based on the element influence difference, calculate the data of the influence difference of different elements in the statistical sequence through the element influence identification function and output the corresponding abnormal data item influence degree YX; Among them, the expression of the element influence identification function is ; In the formula, YC is the element influence difference; YC0 is the standard element influence difference; The abnormal data item impact degree includes a value of 0 or 1, indicating whether the abnormal data item impact status between corresponding elements is normal or abnormal.

2. The method for extracting key domain data information from big data according to claim 1, characterized in that: When implementing different training supervision methods for all training data corresponding to different positions, obtain the name, gender, education background, training start time, training end time, and training results of the corresponding trainees in all training data for different positions in sequence; When processing and extracting the training result dimension for the training data corresponding to all new employees in different positions, the training results from different training data of different positions are obtained in sequence and analyzed traversally; If the training result is normal employment, a local training success instruction is generated and the total number of successful trainees CPi of the corresponding position is increased by one; i is a different position, i=1, 2, 3, ..., n; n is a positive integer; If the training result is failure to join the job, a local training failure instruction will be generated and the total number of training failures SPi for the corresponding position will be increased by one.

3. The method for extracting key domain data information from big data according to claim 2, characterized in that: When processing and analyzing the overall training status of new employees corresponding to different positions, the total number of successful training CPi and the total number of failed training SPi corresponding to different positions are calculated through the training status identification function, and the overall training identification ZPi corresponding to different positions is output; Among them, the expression of the training state recognition function is ; Where PB is the standard value of overall training requirements corresponding to different positions; The overall training flag contains a value of 0 or 1, indicating whether the overall training status of new employees in the corresponding position is normal or abnormal; According to the overall training flag with a value of 1, the corresponding position is marked as the target position, and the training data belonging to the training results of all target positions that have not yet been employed are marked as target training data.

4. The method for extracting key domain data information from big data according to claim 3, characterized in that: Obtain the reasons for non-employment corresponding to the training results for all target positions, and classify and mark all reasons for non-employment; If the reason for non-joining is due to personal reasons of the employee, the non-joining employee will be marked as the first non-joining employee; If the reason for non-employment is not due to personal reasons, the employee will be marked as the second non-employed employee; Obtain all non-joining reasons corresponding to all second non-joining employees, sort and combine all non-joining reasons, deduplicate all identical non-joining reasons in the sorted combinations, and update the total number of occurrences of different non-joining reasons to obtain a non-joining reasons statistics table.

5. The method for extracting key domain data information from big data according to claim 4, characterized in that: When processing the data on the impact of abnormal reasons corresponding to different reasons for non-employment according to the statistical table of reasons for non-employment, the total number of occurrences of different reasons for non-employment is calculated in turn by the formula Calculate and obtain the corresponding abnormal cause impact DYj; where NYj is the total number of occurrences of different reasons for non-employment; j is a different reason for non-employment, j = 1, 2, 3, ..., m; m is a positive integer; B is the standard value of the abnormal cause impact, and [*] is the rounding function, which means obtaining the maximum integer that does not exceed the real number *.

6. The method for extracting key domain data information from big data according to claim 5, characterized in that: When conducting data analysis on the impact of abnormal reasons to determine the impact of abnormal reasons corresponding to different reasons for non-employment; If the impact of the abnormal reason is 0, a normal impact instruction of the abnormal reason is generated, and the corresponding non-employment reason is marked as a normal non-employment reason; If the impact of the abnormal reason is not 0, a special impact instruction of the abnormal reason is generated, and the corresponding non-employment reason is marked as a special non-employment reason; All the special reasons for non-employment marked by the analysis are sorted and combined to obtain the first abnormal impact identification and extraction set.

7. The method for extracting key domain data information from big data according to claim 6, characterized in that: Different elements in different statistical sequences are sequentially converted into Calculate and obtain the element influence ratio YLk of the corresponding element; where YZk is the value of different elements in different statistical sequences; k is the different elements in different statistical sequences, k = 1, 2, 3, ..., p; p is a positive integer, which is the total number of elements in the statistical sequence.

8. The method for extracting key domain data information from big data according to claim 1, wherein the maximum value element among the two calculation elements is marked as a high abnormal impact data item according to the abnormal data item impact degree of 1; All high abnormal impact data items are sorted and combined to obtain the second abnormal impact identification and extraction set.

9. The method for extracting key domain data information from big data according to claim 7, characterized in that: The failure gender statistics series includes the total number of female training failures and the total number of male training failures; The statistical series of failed academic qualifications includes the total number of failed college training, the total number of failed undergraduate training, and the total number of failed master's training; The failed training duration statistical sequence includes the total number of failures during the first duration, the total number of failures during the second duration, and the total number of failures during the third duration.