Human resource text information management method and device based on big data and medium
Through the human resource text information management method based on big data technology, the representative characteristics of text data collected through multi-channel are extracted and processed, and the problem of low text data processing efficiency under the traditional human resource management model is solved, and more efficient and accurate human resource management decisions are achieved.
Patent Information
- Application Number
- CN202510058330.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The traditional human resource management model is difficult to efficiently process and analyze large amounts of unstructured text data, resulting in inefficient information processing and analysis, and the inability to make full use of the human resource data within the enterprise, and the accuracy and effectiveness of decision-making are affected.
The human resource text information management method based on big data technology is adopted, and by obtaining text data collected through multiple channels, representative features are extracted, processing, feature classification and grouping, the feature score of text data is calculated, and the analysis user is pushed.
It significantly improves the efficiency and accuracy of human resource management, reduces manual operations and human errors, saves time and costs, promotes information sharing and collaboration, and helps enterprises build more scientific decision-making models.
Smart Images

Figure CN119988580A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data technology, and specifically to a human resources text information management method, device and medium based on big data. Background Art
[0002] The traditional human resource management model mainly relies on paper documents and simple database systems. This reliance makes the efficiency of information processing and analysis seriously low, and it is impossible to fully explore and utilize the valuable human resource data within the enterprise. Many companies often rely solely on personal experience and intuition to make decisions in key recruitment, performance evaluation, employee training and resignation management processes. This method not only lacks scientific basis, but also greatly affects the accuracy and effectiveness of decision-making. At the same time, this traditional management method has also led to serious information islands, insufficient information sharing between departments, and the inability to achieve comprehensive analysis and collaborative work, further increasing the complexity and cost of management. With the changes in the business environment, companies have continuously accumulated a large amount of unstructured text data, including resumes, performance evaluation reports, training records, employee feedback, and resignation interview records, which contain rich information and insights. However, since these data usually exist in an unstructured form, traditional data processing methods are difficult to efficiently extract effective information from them, resulting in companies being unable to truly use these data to optimize management processes and decisions.
[0003] Therefore, there is an urgent need for a solution based on big data technology to effectively manage and utilize these rich text data, thereby improving the quality and efficiency of human resource management. Summary of the invention
[0004] In order to solve the above technical problems, the purpose of the embodiments of the present application is to provide a human resources text information management method, device and machine-readable storage medium based on big data.
[0005] In a first aspect, the present application provides a human resources text information management method based on big data, comprising:
[0006] Obtain text data of the object to be analyzed collected from multiple channels;
[0007] Extracting representative features of the text data, wherein the representative features are obtained based on basic information of each feature in the text data, wherein the basic information includes a code, a frequency, and a length of each feature;
[0008] Processing the representative features to transform the representative features into features of another dimension;
[0009] Performing feature classification and feature grouping on the processed representative features, wherein the feature classification is used to evaluate the rank of the representative features, and the feature grouping is used to aggregate representative features belonging to similar features;
[0010] The feature scores of the text data of the object to be analyzed are calculated according to the results of the feature grouping and feature classification of the representative features.
[0011] Preferably, representative features of the text data are extracted, and the representative features are obtained based on basic information of each feature in the text data, wherein the basic information includes the encoding, frequency and length of each feature, including:
[0012] Formula (1) is used to extract the representative features of the text data:
[0013]
[0014] Among them, D represents the representative feature, T represents the text data, and w k represents the kth word in the text data T, f 1 Represents the word segmentation encoding function, f 2 represents the word segmentation probability density function, L(w k ,T) represents the kth word w in the text data T k The length of K(w k ,T) represents the kth word w k The number of times it appears in the text data T, γ represents the word segmentation probability influence factor, and δ represents the kth word segmentation w k where N represents the total number of tokens in the text data T and |||||| represents the norm symbol.
[0015] Preferably, processing the representative feature to convert the representative feature into a feature of another dimension comprises:
[0016] The representative feature D is processed using formula (2):
[0017]
[0018] Among them, D represents the representative feature, D 1 represents the representative features after preprocessing, that is, the features of another dimension, μ represents the mean of the representative feature D, σ represents the variance of the representative feature D, and a 1 represents the smoothing factor of the representative feature D, α represents the regularization factor of the representative feature D, β represents the penalty factor of the representative feature D, b represents the scaling factor of the representative feature D, and c represents the mean smoothing factor of the representative feature D.
[0019] Preferably, the processed representative features are subjected to feature classification and feature grouping in feature grouping, including:
[0020]
[0021] Among them, R 1 represents the result of feature grouping, D 1 represents the representative features after preprocessing, that is, the features of another dimension, μ represents the mean of the representative features D, V j Denotes the representative feature D after preprocessing 1 The j-th central feature in , σ represents the variance of the representative feature D, D represents the representative feature, |||| represents the norm sign, and c represents the mean smoothing factor of the representative feature D.
[0022] Preferably, the processed representative features are subjected to feature classification and feature classification in feature grouping, including:
[0023]
[0024] Among them, R 2 represents the result of feature classification, D 1 represents the representative feature after preprocessing, that is, the feature of another dimension, b represents the scaling factor of the representative feature D, |||| represents the norm symbol, and w k represents the kth word in the text data T, p i Denotes the representative feature D after preprocessing 1 The weight factor of the i-th feature in , ε represents the representative feature D after preprocessing 1 The penalty factor, θ i Denotes the representative feature D after preprocessing 1 The scaling factor for the i-th feature in .
[0025] Preferably, the feature score of the text data of the object to be analyzed is calculated according to the result of the feature grouping and feature classification of the representative features, including:
[0026] Formula (5) is used to calculate the feature score of the object to be analyzed:
[0027]
[0028] Among them, R represents the feature score, R 1 Represents the result of feature grouping, R 2 represents the result of feature classification, X 1 The result R representing the feature score 1 The weighted impact factor, X 2 The result R represents the feature grouping 2 The weight influence factor, τ1 The result R representing the feature score 1 The characteristic scaling factor, τ 2 The result R represents the feature grouping 2 The characteristic scaling factor, τ 3 , τ 4 Represents the characteristic adjustment factor.
[0029] Preferably, after calculating the feature score of the text data of the object to be analyzed according to the results of the feature grouping and feature classification of the representative features, the human resources text information management method based on big data further includes:
[0030] The feature score of the object to be analyzed is pushed to the analysis user.
[0031] Preferably, before extracting the representative features of the text data, the human resources text information management method based on big data further includes:
[0032] The text data is cleaned.
[0033] In a second aspect, the present application provides a human resources text information management device based on big data, comprising:
[0034] a memory configured to store instructions; and
[0035] The processor is configured to call instructions from the memory to execute the above-mentioned human resources text information management method based on big data.
[0036] In a third aspect, the present application provides a machine-readable storage medium having instructions stored thereon, the instructions being used to enable a machine to execute the above-mentioned human resources text information management method based on big data.
[0037] The method of the present invention comprises: obtaining text data of an object to be analyzed collected from multiple channels; extracting representative features of the text data, wherein the representative features are obtained based on basic information of each feature in the text data, wherein the basic information includes the coding, frequency and length of each feature; processing the representative features to convert the representative features into features of another dimension; performing feature classification and feature grouping on the processed representative features, wherein the feature classification is used to evaluate the rank of the representative features, and the feature grouping is used to aggregate representative features belonging to similar features; and calculating the feature score of the text data of the object to be analyzed according to the results of the feature grouping and feature classification of the representative features. In the present invention, by adopting big data technology and intelligent data analysis methods, enterprises can efficiently process and analyze a large amount of unstructured text data, significantly improving the efficiency and accuracy of human resource management. The automation of data collection, storage and analysis processes reduces manual operations and human errors, saving time and costs. The present invention reduces the information island phenomenon in the traditional human resource management model and promotes information sharing and collaboration between departments. Each department can access relevant data in real time and conduct comprehensive analysis, so as to better work together and form a joint force, which is conducive to improving the overall management level. With the help of data feature analysis, the present invention can help enterprises build more scientific decision-making models and provide data-driven support for recruitment, performance evaluation, training and development, etc. This data-based decision-making method can improve the accuracy of decision-making and promote the development of human resource management towards intelligence and digitization.
[0038] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present application, but do not constitute a limitation on the embodiments of the present application. In the accompanying drawings:
[0040] Figure 1 A flowchart of a human resources text information management method based on big data provided in an embodiment of the present application;
[0041] Figure 2 A structural block diagram of a human resources text information management device based on big data provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific implementation methods described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0043] It should be noted that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back...), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0044] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist, and is not in the flow chart of the human resources text information management method based on big data required by this application.
[0045] In this application, text information can be equivalent to text data. Data can be equivalent to data features in some cases. Some places in this article only explain one situation (for example, analysis and evaluation of historical data of employees in recruitment or internal enterprises). This is not a limitation on other situations, but can include any equivalent situation (for example, text analysis, employee analysis based on employee data, etc.), which will not be elaborated here. This application is not limited to the field of human resources, but can be used for any text analysis method that adopts the technical solution in this application.
[0046] like Figure 1 As shown, the human resources text information management method based on big data of an embodiment of the present invention may include the following steps:
[0047] Step S10: Acquire text data of the object to be analyzed collected from multiple channels;
[0048] In an embodiment of the present invention, the data collection module: data collection automatically obtains relevant structured text data and unstructured text data from multiple sources such as recruitment websites, internal management systems, employee feedback platforms, etc., including resumes, performance evaluation reports, training records, employee feedback and exit interview records.
[0049] First, recruitment websites are an important channel for obtaining information about potential talents. In compliance with laws and regulations, data interfaces are established with mainstream recruitment platforms (such as Zhaopin.com, 51Job, and BOSS Zhipin) to achieve automatic collection of resume data. These resumes contain the basic information of job seekers (name, gender, age, contact information, etc.), educational background (education, major, graduation school, graduation time, etc.), work experience (work unit, position, working hours, work responsibilities and performance, etc.), skill certificates, and career planning. During the collection process, web crawler technology and data interface technology are used to regularly capture the latest resume data from recruitment websites according to preset rules and frequencies, and store them in the data warehouse.
[0050] Secondly, the internal management system of an enterprise covers a wealth of information related to human resources. The internal management system of an enterprise may include a human resources information system (HRIS), a performance management system, a training management system, etc. Human Resources Information System: This system records the basic file information, attendance records, salary and benefits data, labor contract information, etc. of employees. By connecting with the HRIS system, the automatic extraction and integration of these data can be achieved. For example, the employee's entry time, department, position and other information can be obtained from the employee's basic file; the employee's attendance days, number of late arrivals and early departures and other data can be extracted from the attendance records to provide basic support for subsequent performance evaluation and human resources analysis. In the performance management system, the performance evaluation report is the key basis for understanding the employee's work performance. The performance management system automatically collects information such as the employee's performance evaluation results, superior evaluation, self-evaluation, and performance improvement plan. These data can reflect the employee's work goal completion, work ability and work attitude in a certain period of time, which helps the company identify high-performance and low-performance employees and provide a reference for the promotion, salary adjustment and training development of employees.
[0051] In the data collection process, we first need to analyze the page structure and data format of the recruitment website and determine the crawler's crawling rules and paths. Then, we use professional crawler tools (such as Scrapy, BeautifulSoup, etc.) to write crawler programs, simulate browsers to access the recruitment website, and extract the required resume information according to the preset rules. We also check the legitimacy and integrity of the collected data.
[0052] Data connection with the internal management system and employee feedback platform is mainly achieved through data interface technology. The internal management system of an enterprise usually provides standard data interfaces (such as API interfaces), which can be called to automatically transfer data in the system to the data warehouse in accordance with the agreed data format and protocol. At the same time, in order to ensure the security and stability of the data, the data interface needs to be authenticated and encrypted to prevent data leakage and illegal access.
[0053] In this embodiment, in order to ensure the quality of the collected data, a strict data quality control mechanism is adopted. For example, data verification and data auditing are performed to ensure the quality of the collected data and provide high-quality data for subsequent decision-making support. Specifically, during the data collection process, the collected data is verified in real time to check the integrity, accuracy and legality of the data. For example, check whether the required fields in the resume are filled in completely, whether the scores in the performance evaluation report are within a reasonable range, etc. For data that does not meet the requirements, it is marked and processed in a timely manner to ensure the quality of the data. In addition, the collected data is audited regularly, especially for some key information and important data. Data problems found during the audit process are fed back to the system in a timely manner for data marking. Furthermore, it is necessary to check the consistency of data from different data sources to ensure that the information of the same employee is consistent in different systems. For example, the records of the basic information of employees in the HRIS system and the recruitment system should be consistent. If inconsistencies are found, they need to be checked and corrected in a timely manner to avoid data conflicts and errors.
[0054] Step S20: extracting representative features of the text data, wherein the representative features are obtained based on basic information of each feature in the text data, wherein the basic information includes the code, frequency and length of each feature;
[0055] Before feature extraction of text data or text information, it is necessary to further perform necessary preprocessing operations on the text data or text information. Specifically, it includes data cleaning, word segmentation, part-of-speech tagging and named entity recognition of text data or text information. In data cleaning, the collected text data or text information is cleaned to remove noise data, such as HTML tags, JavaScript codes, redundant blank characters, special symbols, etc. At the same time, typos and abbreviations in the text are corrected and expanded to improve the readability and accuracy of the text. According to the language type of the text, select a suitable word segmentation tool (such as Chinese Jieba word segmentation, LTP, English NLTK, Stanford CoreNLP, etc.) to segment the text into independent words or phrases. Word segmentation is the basis for subsequent text analysis and can convert text into basic units that can be processed by computers. The text after word segmentation is tagged with parts of speech to determine the part of speech of each word (such as noun, verb, adjective, etc.), and the named entity recognition technology is used to identify entity information such as names of people, places, organizations, and products in the text. This information helps to further understand the semantics and structure of the text.
[0056] Determine the type of features to be extracted based on the characteristics of the object to be analyzed and the purpose of the analysis. Common features include lexical features (such as keywords, high-frequency words, subject words, etc.), semantic features (such as sentiment orientation, semantic similarity, etc.), and structural features (such as sentence length, paragraph structure, etc.). For example, when analyzing product user feedback, the product function name, user evaluation sentiment words (such as "satisfied", "disappointed"), etc. can be used as important features; when analyzing industry reports, industry keywords, the number of times policies and regulations are mentioned, etc. may be key features.
[0057] After word segmentation, the text is traversed to identify different features. These features can be specific concepts related to human resource management, such as "recruitment", "performance evaluation", "employee training", "resignation management", etc.; they can also be words describing management methods, data status, etc., such as "traditional management model", "information island", "unstructured text data", etc.
[0058] Count the number of times each feature appears in the text. You can create a dictionary data structure to record the frequency of each feature, with the number of times the feature appears as the value. For example, if the feature "recruitment" appears 3 times in the text, then the corresponding frequency value in the dictionary is {X:3}, where X can be a coded value set as needed or other value used for representation. In order to make the feature frequencies between different texts comparable, the feature frequencies can be normalized.
[0059] In addition, for each feature, the number of characters or words it contains can be calculated as the length of the feature. For example, the feature "traditional management mode" contains 5 characters, so its length is 5; the feature "unstructured text data" contains 7 characters, so its length is 7. In order to more comprehensively describe the length information of text features, the average length of all features can be calculated. The sum of the lengths of all features is divided by the total number of features to obtain the average feature length. This indicator can reflect the overall length level of features in the text.
[0060] Different weights are assigned to the feature attributes, frequency, and length of the feature according to actual needs and business background. For example, if you think that the feature frequency is more important in reflecting text information, you can assign a higher weight to the frequency; if you are more concerned about the specific identification of the feature, you can appropriately increase the weight of the encoding; if you want to highlight the detailed description of the feature, you can increase the weight of the length.
[0061] Finally, calculate the comprehensive score. For each feature, multiply its feature attribute value, frequency, and length by the corresponding weight, and then add them together to get the comprehensive score of each feature. The calculation formula can be expressed as: comprehensive score = feature attribute weight × feature attribute value + frequency weight × frequency value + length weight × length value.
[0062] Of course, in other embodiments, other basic information of each feature of the text information, such as the criticality of the text, the relevance of the text to the analysis target, and other basic information may be further considered to extract representative features.
[0063] Finally, we screen representative features, sort all features according to their comprehensive scores, and select several features with higher scores as representative features. These representative features can summarize the core information of the text to a certain extent and reflect the main content and key points discussed in the text.
[0064] Through the above steps, the processor can extract representative features based on the characteristic attributes, frequency and length of each feature of the text information, so as to have a deeper understanding of the text content and provide strong support for subsequent data analysis and decision-making.
[0065] Step S30: Processing the representative feature to convert the representative feature into a feature of another dimension;
[0066] The collected unstructured text data or data features are sorted and cleaned, and converted into data or data features (or features of another dimension) in an analyzable format. By removing interference factors in the data, the data quality and analysis accuracy are improved, making the results of the subsequent analysis stage more reliable.
[0067] The sources of unstructured text data are extremely rich and diverse, including but not limited to recruitment media platforms, internal corporate documents, customer feedback records, etc. Data from different sources have different characteristics and formats. These data have no fixed format and structure, and may contain various special characters, HTML tags, typesetting formats, etc. For example, the text on a web page may be interspersed with a large number of HTML tags to control the page layout and style. These tags have no practical significance for the analysis of text content, but will interfere with data processing. There is often a lot of noise and redundant information in the data, such as advertising content, repeated text fragments, irrelevant links, etc. Use regular expressions or text processing tools to remove special characters (such as punctuation, emoticons, mathematical symbols, etc.) and HTML tags in the text. This step can simplify the text into plain text form for subsequent processing. For example, for text containing HTML tags " This is a paragraph Bold Text ”, after processing, we can get “This is a bold text”. The text format conversion method is used to convert the text into a unified uppercase and lowercase format to eliminate the impact of uppercase and lowercase differences on the analysis. At the same time, the abbreviations and abbreviations in the text are expanded to make them have complete semantics. For example, "don't" is expanded to "do not", "it's" is expanded to "it is", etc. In addition, by comparing the similarity of the text, duplicate text fragments in the data are identified and removed. At the same time, according to business needs and data characteristics, rules are formulated to remove noise information, such as advertising content, irrelevant links, etc. For example, advertising text can be identified and deleted from the data by matching specific keywords or regular expressions.
[0068] After completing data cleaning, the processed data is verified and quality checked to ensure the accuracy and completeness of the data. This can be achieved through basic features of statistical data (such as text length distribution, vocabulary frequency, etc.), manual sampling checks, etc. If problems are still found in the data, it is necessary to return to the previous steps for further processing.
[0069] After the above-mentioned sorting and cleaning steps, the unstructured text data or data features are converted into data or data features (or features of another dimension) in an analyzable format.
[0070] By organizing and cleaning unstructured text data or data features and converting them into analyzable data or data features (or features of another dimension), data quality can be effectively improved, laying a solid foundation for subsequent data analysis and mining.
[0071] Step S40: performing feature classification and feature grouping on the processed representative features, wherein the feature classification is used to evaluate the rank of the representative features (i.e., the level of feature importance), and the feature grouping is used to cluster representative features belonging to similar features;
[0072] The goal of feature grouping is to aggregate features with similar attributes, such as skills, work experience, and feedback type, by analyzing and distinguishing the features of resumes, performance evaluations, and employee feedback. Specifically, this process is achieved by minimizing the squared distance between each feature and the central feature. In this way, the system can efficiently identify and summarize text features that are similar in attributes or semantics, so that they can be clustered together.
[0073] The advantage of feature grouping is that it reduces data redundancy, optimizes data structure, and makes feature representation more compact. This compact feature representation can not only improve the efficiency of data storage, but also significantly speed up subsequent data processing. By aggregating similar features, the system can more easily gain insight into the overall trend of the data and provide more consistent and accurate analysis results.
[0074] The purpose of feature classification is to systematically distinguish and identify features extracted from documents such as resumes and performance evaluations. For example, this process can include classifying job applicants in resumes according to skill types (such as technical, management, etc.), and grading performance evaluation results to clarify the standards for high-performing and low-performing employees. Through feature classification, the system can not only effectively identify and mark the category corresponding to each feature, but also help human resource managers quickly and accurately obtain the necessary information when evaluating employee performance. This classification result provides a basis for subsequent decision-making, enabling more scientific and reasonable judgments based on data when conducting talent recruitment, performance analysis, and employee development.
[0075] Step S50: Calculate the feature score of the text data of the object to be analyzed according to the results of the feature grouping and feature classification of the representative features.
[0076] The processor calculates the feature score by combining the results of feature grouping and feature classification, thereby providing flexible, dynamic and data-driven support for human resource management. It not only improves the efficiency and accuracy of decision-making, but also promotes the overall optimization of talent management, providing a solid technical foundation for enterprises in human resources. The processor uses the results of feature grouping and feature classification to calculate the feature score, thereby providing support for human resource analysis. By combining the results of feature grouping and feature classification, the processor can calculate the feature score, which the human resources department can use to analyze the employee status, help identify high-potential employees, evaluate training needs or formulate improvement measures. For example, by analyzing performance scores, employees with outstanding work performance can be found and provided with more career development opportunities. Or in the recruitment process, by calculating the feature scores of job applicants' resumes, the best candidates for the job requirements can be quickly screened out, thereby improving recruitment efficiency. Analysis based on feature scores can help management formulate more targeted staffing, training and development strategies. By identifying the relationship between employee characteristics and business performance, the overall performance of the organization can be promoted. The processor can integrate data features from different sources to form a unified feature score, avoid information islands, and thus improve data usage efficiency.
[0077] The method of the present invention comprises: obtaining text data or text information of an object to be analyzed collected from multiple channels; extracting representative features of the text data or text information, wherein the representative features are obtained based on basic information of each feature in the text data or text information, wherein the basic information includes the coding, frequency and length of each feature; processing the representative features to convert the representative features into features of another dimension; performing feature classification and feature grouping on the processed representative features, wherein the feature classification is used to evaluate the rank of the representative features, and the feature grouping is used to aggregate representative features belonging to similar features; and calculating the feature score of the text data or text information of the object to be analyzed according to the results of the feature grouping and feature classification of the representative features. In the present invention, by adopting big data technology and intelligent data analysis methods, enterprises can efficiently process and analyze a large amount of unstructured text data, significantly improving the efficiency and accuracy of human resource management. The automation of data collection, storage and analysis processes reduces manual operations and human errors, saving time and costs. The present invention reduces the information island phenomenon in the traditional human resource management model and promotes information sharing and collaboration across departments. Each department can access relevant data in real time and conduct comprehensive analysis, so as to work together better and form a joint force, which helps to improve the overall management level. With the help of data feature analysis, the present invention can build a more scientific decision-making model and provide data-driven support for recruitment, performance evaluation, training and development. This data-based decision-making method can improve the accuracy of decision-making and promote the development of human resource management towards intelligence and digitization.
[0078] In another embodiment of the present invention, representative features of the text data are extracted, and the representative features are obtained based on basic information of each feature in the text data, and the basic information includes the encoding, frequency and length of each feature, including:
[0079] Formula (1) is used to extract the representative features of the text information:
[0080]
[0081] Among them, D represents the representative feature, T represents the text data, and w k represents the kth word in the text data T, f 1 Represents the word segmentation encoding function, f 2 represents the word segmentation probability density function, L(w k ,T) represents the kth word w in the text data T k The length of K(w k , T) represents the kth participle w k The number of times it appears in the text data T, γ represents the word segmentation probability influence factor, and δ represents the kth word segmentation wk , N represents the total number of word segments in the text data T, and |||| represents the norm symbol.
[0082] Specifically, in this embodiment, the method of formula (1) is adopted to extract the representative features of the text data. In addition to the method of using the weight value calculation method in the above embodiment to calculate the score of each feature of the text data and determine the representative features according to the score, this embodiment adopts a complex linear calculation method to extract the representative features of the text, and comprehensively considers the length, frequency and word segmentation probability of the text to extract the representative features.
[0083] Specifically, after receiving the text data, the feature extraction module will extract features and apply the above text features to the formula (1) for calculation. The calculation method of formula (1) can quantify the value of each feature. By combining the encoding, frequency, length and other factors of each feature, the representative features finally obtained are representative and reliable.
[0084] The calculation method of this embodiment is used to conduct a comprehensive evaluation in combination with the frequency and importance of the word to ensure that features with high discrimination and representativeness are extracted. By introducing factors such as word length, frequency of occurrence, and word segmentation probability, feature extraction is made more accurate and comprehensive. This complex linear relationship formula is selected to ensure that different text features can be fully quantified to provide accurate numerical support for subsequent analysis.
[0085] In another embodiment of the present invention, processing the representative feature to convert the representative feature into a feature of another dimension includes:
[0086] The representative feature D is processed using formula (2):
[0087]
[0088] Among them, D represents the representative feature, D 1 represents the representative features after preprocessing, that is, the features of another dimension, μ represents the mean of the representative feature D, σ represents the variance of the representative feature D, and a 1 represents the smoothing factor of the representative feature D, α represents the regularization factor of the representative feature D, β represents the penalty factor of the representative feature D, b represents the scaling factor of the representative feature D, and c represents the mean smoothing factor of the representative feature D.
[0089] The method of this embodiment is used to process the text data features, and by removing interference factors and applying regularization, the clarity and accuracy of the data can be significantly improved. In addition, comprehensive consideration of the mean, variance and penalty factor can help avoid overfitting and improve the generalization ability of the model. Through this processing formula, it is ensured that the representative features obtained after preprocessing are more representative and suitable for various algorithms in subsequent analysis. The method of this embodiment can help correct potential deviations and noise effects.
[0090] In another embodiment of the present invention, the processed representative features are subjected to feature classification and feature grouping in feature grouping, including:
[0091]
[0092] Among them, R 1 represents the result of feature grouping, D 1 represents the representative features after preprocessing, that is, the features of another dimension, μ represents the mean of the representative features D, V j Denotes the representative feature D after preprocessing 1 The j-th central feature in , σ represents the variance of the representative feature D, D represents the representative feature, |||| represents the norm sign, and c represents the mean smoothing factor of the representative feature D.
[0093] Reduce data redundancy: By clustering similar features together, valuable information can be highlighted and the complexity of processing can be reduced.
[0094] Efficient feature representation: Compact feature representation accelerates subsequent analysis and computation.
[0095] Reasons for selection:
[0096] Choosing such a grouping method can better provide a clear data structure for subsequent classification and analysis, making data processing more efficient and consistent.
[0097] In another embodiment of the present invention, the processed representative features are subjected to feature classification and feature classification in feature grouping, including:
[0098]
[0099] Among them, R 2 represents the result of feature classification, D 1 represents the representative feature after preprocessing, that is, the feature of another dimension, b represents the scaling factor of the representative feature D, |||| represents the norm symbol, and w k represents the kth word in the text data T, p i Denotes the representative feature D after preprocessing 1The weight factor of the i-th feature in , ε represents the representative feature D after preprocessing 1 The penalty factor, θ i Denotes the representative feature D after preprocessing 1 The scaling factor for the i-th feature in .
[0100] The method of this embodiment improves the interpretability of features by clarifying the classification, which facilitates HR decision-making. It ensures that the standards for efficient and inefficient employees are systematized, helping to take appropriate management measures in a timely manner. Such a classification method aims to clearly distinguish feature attributes, making evaluation and decision-making more efficient and in line with business needs.
[0101] In another embodiment of the present invention, calculating the feature score of the text data of the object to be analyzed according to the result of the feature grouping and feature classification of the representative features includes:
[0102] Formula (5) is used to calculate the feature score of the object to be analyzed:
[0103]
[0104] Among them, R represents the feature score, R 1 Represents the result of feature grouping, R 2 represents the result of feature classification, X 1 The result R representing the feature score 1 The weighted impact factor, X 2 The result R represents the feature grouping 2 The weight influence factor, τ 1 The result R representing the feature score 1 The characteristic scaling factor, τ 2 The result R represents the feature grouping 2 The characteristic scaling factor, τ 3 , τ 4 Represents the characteristic adjustment factor.
[0105] In a complex and ever-changing business environment, human resource management is crucial to the success of an enterprise. In order to evaluate employees more comprehensively and accurately, this paper introduces an innovative comprehensive evaluation method that integrates key features such as multidimensional characteristics, exponential growth smoothness, dynamic adjustment capabilities, avoiding information islands, and supporting decision-making.
[0106] In traditional employee evaluation, only a single or a few characteristics are often focused on, which makes the understanding of employees not comprehensive and in-depth. The comprehensive evaluation method of the present invention fully considers multi-dimensional characteristics and covers all aspects of employees at work, including but not limited to work performance, professional skills, teamwork ability, innovative thinking, communication ability and learning ability.
[0107] In order to process these rich feature information more effectively, the present invention uses advanced data analysis algorithms to group and classify features. For example, the features directly related to work results are grouped into one group, such as project completion rate, work quality, etc.; the features reflecting the personal ability and quality of employees are grouped into another group, such as professional knowledge level, problem-solving ability, etc.; the features reflecting the social and teamwork aspects of employees are grouped into one group, such as teamwork tacit understanding, communication effect, etc.; in addition, these features are further classified into different levels, that is, graded. Through such grouping and classification, the position and role of each feature in the overall evaluation can be more clearly understood.
[0108] Then, the present invention combines the results of these groupings and classifications using an algorithmic formula. This formula is not a simple linear addition, but a comprehensive consideration of the interrelationships and weights between the various features. For example, for different types of jobs, certain features may have higher weights. In this way, the final score can fully reflect the employee's performance in all dimensions, presenting us with a complete and three-dimensional portrait of the employee.
[0109] In the comprehensive evaluation, it is considered that different characteristics have different importance to employee performance and potential. In order to more accurately reflect this difference, the present invention introduces the property of exponential growth smoothness.
[0110] Specifically, the present invention uses an exponential function to process feature scores. The characteristics of the exponential function make the feature score more sensitive to the importance of grouping and classification. For those low-scoring features, their contribution to the final score will not be too large. This is because the curve of the exponential function is relatively flat in the low-scoring area, and as the score decreases, its impact on the final score will decrease rapidly. This is also one of the purposes of the present invention to use this calculation method to calculate the final feature score.
[0111] On the contrary, for high-scoring features, the curve of the exponential function becomes steeper in the high-scoring area, which means that the impact of high-scoring features will be significantly amplified. For example, an employee who excels in innovative thinking will have his or her high score on this feature fully reflected in the final score through the effect of the exponential function, thereby widening the gap with other employees who perform averagely in this aspect. This approach greatly improves the differentiation of scores, allowing analysts to more accurately identify employees who have outstanding advantages in certain key aspects.
[0112] By adjusting the weight factor and the scaling factor, the calculation of the feature score can be dynamically adjusted according to the actual situation. For example, when the company is in the business expansion stage, it may pay more attention to the market development ability and teamwork ability of employees. At this time, the weights of these two aspects of the characteristics can be appropriately increased; when the company enters the technological innovation stage, the requirements for employees' professional skills and innovative thinking are higher, and we can adjust the weights of these features accordingly.
[0113] In order to avoid the adverse effects of information silos on employee evaluation, our comprehensive evaluation module integrates feature scores from different sources to form a unified view. This process is not just a simple aggregation of data, but also a deep integration and analysis of information from different sources.
[0114] For example, we can combine the basic information and training records of employees recorded in the human resources management system with the actual performance of employees in the project reflected in the project management system, and then integrate them into the quantitative data in the performance appraisal system, so as to comprehensively and accurately evaluate the work ability and potential of employees. By eliminating data silos, we can have a more comprehensive information basis when making decisions and avoid decision-making errors caused by incomplete information.
[0115] The final feature score is not just a number, it has important practical application value and can directly support a series of decisions in human resource management.
[0116] In terms of identifying high-potential employees, the characteristic scores obtained through comprehensive evaluation can help us accurately screen out those employees who perform well in multiple dimensions and have great development potential. These employees may be the core strength of the company in the future, and providing them with more development opportunities and resources will help improve the overall competitiveness of the company.
[0117] In the process of precision recruitment, we can match the key characteristics required for the recruitment position with the characteristic scores of job seekers, so as to more accurately screen out talents that meet the requirements of the position. This data-based recruitment method can improve recruitment efficiency and quality and reduce recruitment risks.
[0118] For optimizing training needs, feature scores can help us find out where employees are lacking, so that we can develop targeted training plans. For example, if an employee scores low in communication skills, we can arrange relevant communication skills training courses for him to help him improve his abilities and achieve common development of individuals and enterprises.
[0119] In summary, the formula of the embodiment of the present invention combines the results of feature grouping and classification together so that the final score can reflect more comprehensive information. This comprehensive evaluation helps to better understand the performance and potential of employees. In addition, the use of exponential functions makes the feature scores more sensitive to the importance of grouping and classification. The low-scoring features will not contribute too much to the final score, while the impact of high-scoring features will be amplified, thereby improving the discrimination of the scores. By adjusting the weight factors and scaling factors, the calculation of feature scores can be dynamically adjusted as needed to adapt to different business scenarios and management goals. By integrating feature scores from different sources, the module can form a unified view, eliminate data silos, and help form a comprehensive information basis when making decisions. The final feature score can directly support a series of decisions in human resource management, help identify high-potential employees, accurately recruit, optimize training needs, etc.
[0120] In another embodiment of the present invention, after calculating the feature score of the object to be analyzed according to the results of the feature grouping and feature classification of the representative features, the human resources text information management method based on big data further includes:
[0121] The feature score of the object to be analyzed is pushed to the analysis user.
[0122] In the embodiment of the present invention, the obtained feature scores are finally pushed to users who need the data, such as the human resources analysis department. Based on the feature scores, the human resources analysis department can make decisions according to its own needs.
[0123] The content of this scheme will be fully described below with specific embodiments:
[0124] Suppose you need to analyze a batch of software engineer resumes to identify high-potential candidates and perform effective staffing. The processor will process the data in the following steps:
[0125] 1. Data Collection
[0126] The processor first obtains a set of resume data, each resume contains the following information: name
[0127] Contact Details
[0128] Education
[0129] Professional skills (such as programming languages, tools, etc.)
[0130] Project Experience
[0131] Work Experience
[0132] Personal Introduction
[0133] Among them, the source data example of the above resume data is:
[0134] {
[0135] "Name":"Zhang San",
[0136] "Contact information":"123456789",
[0137] "Educational Background":"Computer Science and Technology",
[0138] "Professional skills": ["Python", "Java", "Machine Learning"],
[0139] "Project experience": ["Intelligent recommendation system development", "Big data processing platform"],
[0140] "Work Experience": ["Company A (2019-2021)", "Company B (2022 to present)"],
[0141] "Personal introduction": "Passionate about technology and good at learning new technologies."
[0142] }
[0143] 2. Feature extraction
[0144] Assume that the following features are of interest:
[0145] Skills (Python, Java, Machine Learning)
[0146] Number of projects
[0147] Years of work experience
[0148] The result of feature extraction will look like this:
[0149] feature Python Hava Machine Learning Number of projects Years of work experience Zhang San 1 1 1 2 5
[0150] explain:
[0151] Skills such as Python, Java, and machine learning are coded as 1 if they are present and 0 if they are not present.
[0152] The number of projects and years of work experience are expressed as integer fields.
[0153] 3. Feature preprocessing
[0154] During the feature preprocessing phase, the processor will check and clean the data, assuming there are no missing values, and keep the data in binary representation.
[0155] 4. Feature Grouping
[0156] The characteristics are grouped so that they can be aggregated into a few main categories. The processor aggregates skills, project experience, and work experience into the following categories:
[0157] Programming skills: Using binary code to represent quantities
[0158] Project experience: using project quantity to express
[0159] Work experience: Use years of experience to indicate
[0160] The structured results of feature grouping (assuming the number of programming skills is in the interval of 1-3):
[0161] Grouping Number of programming skills Project Experience Work Experience Zhang San 3 2 5
[0162] 5. Feature Classification
[0163] In the feature classification stage, we divide coding into basic skills and advanced skills. Based on skill assessment:
[0164] Basic skills
[0165] Java
[0166] Advanced Skills
[0167] Python
[0168] Machine Learning
[0169] The calculated feature classification results are:
[0170] Classification Number of basic skills Number of advanced skills Zhang San 1 2
[0171] 6. Feature score calculation
[0172] Results R based on feature grouping 1 And the result of feature classification R 2 To calculate the feature score R. Assume the following weights and adjustment factors are used:
[0173] Results R based on feature grouping 1 And the result of feature classification R 2 To calculate the feature score R. Assume the following weights and adjustment factors are used:
[0174] ·X 1 =0.5 (weight of feature grouping)
[0175] ·X 2 =1.5 (weight of feature classification)
[0176] ·τ 1 =1
[0177] ·τ 2 =1
[0178] ·τ 3 =1
[0179] ·τ 4 =1
[0180] Apply the formula to calculate:
[0181]
[0182] Substitute the numerical calculation:
[0183] Assume R 1 Total value = 3 + 2 + 5 = 10
[0184] Assume R 2 Total value = 1 + 2 = 3
[0185] Finally, the value of R can be obtained:
[0186] Assume the calculation result is R = 85
[0187] 7. Results Analysis
[0188] By feature score R:
[0189] Identifying high-potential employees: Based on the characteristic scores, Zhang San has a higher score (85) and can be considered a high-potential employee.
[0190] Develop recruitment strategies: The HR team can quickly decide whether to interview or further screen the candidate based on the score.
[0191] Through feature scoring and corresponding characteristic analysis, personnel selection and decision support are completed, making the recruitment process more efficient and targeted.
[0192] Figure 2 The structural block diagram of a human resources text information management device based on big data provided by the embodiment of the present application. Figure 2 As shown, the human resources text information management device based on big data may include:
[0193] Memory 210, configured to store instructions; and
[0194] The processor 220 is configured to call instructions from the memory 210 and implement the above-mentioned human resources text information management method based on big data when executing the instructions.
[0195] An embodiment of the present invention further provides a human resources text information management system based on big data, and the system may include:
[0196] A data collection module is used to obtain text data of the object to be analyzed collected from multiple channels;
[0197] A feature extraction module, used for extracting representative features of the text data, wherein the representative features are obtained based on basic information of each feature in the text data, wherein the basic information includes a code, a frequency and a length of each feature;
[0198] A feature preprocessing module, used for processing the representative feature to convert the representative feature into a feature of another dimension;
[0199] A text data management module, used for performing feature classification and feature grouping on the processed representative features, wherein the feature classification is used for evaluating the rank of the representative features, and the feature grouping is used for clustering representative features belonging to similar features;
[0200] The text data analysis module is used to calculate the feature score of the text data of the object to be analyzed according to the results of the feature grouping and feature classification of the representative features.
[0201] The embodiments of the present invention provide a human resources text information management system based on big data (i.e., a decision support system based on data analysis), a specific machine learning model algorithm designed for human resources management needs, and specific methods and technical details for acquiring, cleaning, and analyzing unstructured text data, so as to help human resources management personnel make more scientific decisions.
[0202] An embodiment of the present application also provides a machine-readable storage medium having instructions stored thereon, the instructions being used to enable a machine to execute the above-mentioned human resources text information management method based on big data.
[0203] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0204] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0205] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0206] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0207] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0208] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0209] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0210] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0211] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. Human resources text information management method based on big data, including: Obtain text data of the object to be analyzed collected from multiple channels; Extracting representative features of the text data, wherein the representative features are obtained based on basic information of each feature in the text data, wherein the basic information includes a code, a frequency, and a length of each feature; Processing the representative features to transform the representative features into features of another dimension; Performing feature classification and feature grouping on the processed representative features, wherein the feature classification is used to evaluate the rank of the representative features, and the feature grouping is used to cluster representative features belonging to similar features; The feature scores of the text data of the object to be analyzed are calculated according to the results of the feature grouping and feature classification of the representative features.
2. The human resources text information management method based on big data according to claim 1, wherein: Extracting representative features of the text data, wherein the representative features are obtained based on basic information of each feature in the text data, wherein the basic information includes the encoding, frequency and length of each feature, including: Formula (1) is used to extract the representative features of the text data: Among them, D represents the representative feature, T represents the text data, and w k represents the kth word in the text data T, f1 represents the word encoding function, f2 represents the word probability density function, L(w k ,T) represents the kth word w in the text data T k The length of K(w k ,T) represents the kth word w k The number of times it appears in the text data T, γ represents the word segmentation probability influence factor, and δ represents the kth word segmentation w k where N represents the total number of tokens in the text data T and |||||| represents the norm symbol.
3. The human resources text information management method based on big data according to claim 2, wherein: Processing the representative features to convert the representative features into features of another dimension includes: The representative feature D is processed using formula (2): Among them, D represents the representative feature, D1 represents the representative feature after preprocessing, that is, the feature of another dimension, μ represents the mean of the representative feature D, σ represents the variance of the representative feature D, a1 represents the smoothing factor of the representative feature D, α represents the regularization factor of the representative feature D, β represents the penalty factor of the representative feature D, b represents the scaling factor of the representative feature D, and c represents the mean smoothing factor of the representative feature D.
4. The human resources text information management method based on big data according to claim 1, wherein: The processed representative features are classified and grouped into features, including: Among them, R1 represents the result of feature grouping, D1 represents the representative feature after preprocessing, that is, the feature of another dimension, μ represents the mean of the representative feature D, V j represents the j-th central feature in the representative feature D1 after preprocessing, σ represents the variance of the representative feature D, D represents the representative feature, ||||||| represents the norm sign, and c represents the mean smoothing factor of the representative feature D.
5. The human resources text information management method based on big data according to claim 1, wherein: The processed representative features are classified and grouped into features, including: Among them, R2 represents the result of feature classification, D1 represents the representative feature after preprocessing, that is, the feature of another dimension, b represents the scaling factor of the representative feature D, ||||||| represents the norm symbol, and w k represents the kth word in the text data T, p i represents the weight factor of the i-th feature in the preprocessed representative feature D1, ε represents the penalty factor of the preprocessed representative feature D1, θ i Represents the scaling factor of the i-th feature in the preprocessed representative feature D1.
6. The human resources text information management method based on big data according to claim 1, wherein: Calculating the feature score of the text data of the object to be analyzed based on the results of the feature grouping and feature classification of the representative features, including: Formula (5) is used to calculate the feature score of the object to be analyzed: Among them, R represents the feature score, R1 represents the result of feature grouping, R2 represents the result of feature classification, X1 represents the weight influence factor of the feature score result R1, X2 represents the weight influence factor of the feature grouping result R2, τ1 represents the feature scaling coefficient of the feature score result R1, τ2 represents the feature scaling coefficient of the feature grouping result R2, and τ3 and τ4 represent feature adjustment factors.
7. The human resources text information management method based on big data according to any one of claims 1 to 6, wherein: After calculating the feature score of the object to be analyzed based on the results of the feature grouping and feature classification of the representative features, the human resources text information management method based on big data further includes: The feature score of the object to be analyzed is pushed to the analysis user.
8. The human resources text information management method based on big data according to any one of claims 1 to 6, wherein: Before the step of extracting representative features of the text information, the human resources text information management method based on big data further includes: The text data is cleaned.
9. A human resources text information management device based on big data, comprising: a memory configured to store instructions; as well as A processor, wherein the processor is configured to call the instruction from the memory to execute the human resources text information management method based on big data according to any one of claims 1-8.
10. A medium having instructions stored thereon, wherein the instructions are used to enable a machine to execute the human resources text information management method based on big data according to any one of claims 1-8.
Citation Information
Cited By
Recruitment information supervision platform and supervision method based on multi-channel integration
CN121788086A