A talent evaluation and management method and system based on artificial intelligence
By using text recognition and structured processing of resume data, a 3D point cloud dataset is constructed, spatial vector angles are calculated, and dynamic correction coefficients are generated. This solves the problem of low efficiency in traditional talent evaluation methods and achieves more accurate and comprehensive talent evaluation.
Patent Information
- Application Number
- CN202511724625.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-11-24
AI Technical Summary
Traditional talent evaluation and management methods rely on manual screening, which is inefficient and makes it difficult to fully analyze the implicit information in a large number of resumes. Furthermore, fixed evaluation indicators cannot be dynamically adapted to the personalized needs of different positions, making it difficult to discover the potential abilities and development potential of candidates.
By acquiring resume data and performing text recognition processing, a structured text dataset is generated. Keyword feature vectors are extracted, a 3D point cloud dataset is constructed, the spatial vector angle boundary is calculated, the sector area range is defined, dynamic correction coefficients are generated, and finally, a talent management score is output.
It achieves a precise match between talent characteristics and evaluation standards, reduces the subjectivity of manual screening, enhances the comprehensiveness and objectivity of evaluation, adapts to the dynamic changes in talent characteristics, and outputs more accurate scores.
Smart Images

Figure CN121190025B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a talent evaluation and management method and system based on artificial intelligence. Background Technology
[0002] Traditional talent evaluation and management methods often rely on manual resume screening, interview questions and answers, and experience-based judgment, which can be inefficient. For example, HR personnel may find it difficult to fully analyze the implicit information in a large number of resumes in a short period of time, and may miss outstanding talents due to personal preferences or cognitive limitations. At the same time, fixed evaluation indicators (such as education and years of work experience) cannot be dynamically adapted to the personalized needs of different positions, making it difficult to discover the potential abilities and development potential of candidates.
[0003] With the development of information technology, some digital tools have been introduced to assist in talent evaluation, improving screening efficiency through keyword matching, simple scoring models, and other methods. However, some of these methods still remain at the level of static data processing, only achieving a superficial comparison of resume information. They may not be able to deeply analyze the correlation and dynamic evolution of talent characteristics. For example, the inherent logic between the skills and project experience accumulated by candidates at different career stages, as well as the deep matching degree between these elements and the requirements of the target position, are difficult to effectively capture through traditional digital means. Summary of the Invention
[0004] The technical problem to be solved by this invention is to provide a talent evaluation and management method and system based on artificial intelligence, so as to improve the comprehensiveness of the evaluation.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0006] Firstly, an artificial intelligence-based talent evaluation and management method, the method comprising:
[0007] Step 1: Obtain the resume data of the candidates to be evaluated, and perform text recognition processing on the resume data to generate a structured text dataset;
[0008] Step 2: Extract keyword feature vectors from the structured text dataset, calculate the similarity value between the keyword feature vectors and the preset talent evaluation keyword database, and generate initial similarity data;
[0009] Step 3: Construct a 3D point cloud dataset based on the initial similarity data, where each data point contains the weight coordinate information of the keyword;
[0010] Step 4: Determine the reference coordinate points based on the spatial density distribution of the point cloud dataset, and generate two spatial vector data based on the reference coordinate points, and calculate the boundary data of the included angle between the spatial vector data.
[0011] Step 5: Define the range of the sector area using the included angle boundary data, select the first spatial coordinate data within the sector area, and symmetrically select the second and third spatial coordinate data outside the range;
[0012] Step 6: Generate arc trajectory data according to the selection order of the first, second, and third spatial coordinate data; analyze the curvature feature value of the arc trajectory data to generate dynamic correction coefficients; integrate the initial similarity data and dynamic correction coefficients to output the final talent management score.
[0013] Secondly, an artificial intelligence-based talent evaluation and management system includes:
[0014] The processing module is used to acquire the resume data of the candidates to be evaluated, perform text recognition processing on the resume data, and generate a structured text dataset.
[0015] The calculation module is used to extract keyword feature vectors from the structured text dataset, calculate the similarity value between the keyword feature vectors and the preset talent evaluation keyword library, and generate initial similarity data.
[0016] The building module is used to construct a 3D point cloud dataset based on the initial similarity data, where each data point contains the weight coordinate information of the keyword;
[0017] The determination module is used to determine the reference coordinate point based on the spatial density distribution of the point cloud dataset, generate two spatial vector data based on the reference coordinate point, and calculate the boundary data of the included angle between the two spatial vector data.
[0018] The selection module is used to define the range of a sector area using the included angle boundary data, select the first spatial coordinate data within the sector area, and symmetrically select the second and third spatial coordinate data outside the range.
[0019] The parsing module is used to generate arc-shaped trajectory data based on the time series relationship of the first, second, and third spatial coordinate data, parse the curvature feature value of the trajectory data to generate dynamic correction coefficients, and finally fuse the initial similarity data and dynamic correction coefficients to output the final talent management score.
[0020] Thirdly, a computing device, comprising:
[0021] One or more processors;
[0022] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.
[0023] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0024] The above-described solution of the present invention has at least the following beneficial effects:
[0025] By transforming unstructured resumes into standardized structured datasets through text recognition, information extraction biases caused by the chaotic format of original resumes are avoided. Core features are extracted from the structured data and compared with a pre-defined keyword database, achieving precise matching between talent characteristics and evaluation standards. This reduces the subjectivity of manual screening, making the initial evaluation results more objective and targeted. Multi-dimensional information such as keyword weight, similarity, and frequency of occurrence are mapped to three-dimensional coordinates, visually presenting the complex relationships of talent characteristics in a spatial manner. This breaks through the limitations of single-dimensional evaluation, achieving a three-dimensional portrayal of talent traits. Determining benchmark coordinates and calculating the angle between spatial vectors captures the core distribution patterns of talent characteristics, making the evaluation more aligned with mainstream talent traits. By selecting coordinates within and outside the sector area, the matching degree between talent and standards (within the area) and the differentiation of traits (outside the area) are considered, ensuring the benchmark nature of the evaluation while identifying the unique advantages of talent, thus improving the comprehensiveness of the evaluation. Dynamic correction coefficients are generated through trajectory curvature analysis, and the initial results are adjusted in conjunction with time series relationships, enabling the evaluation to adapt to the dynamic changes in talent characteristics. Ultimately, the scores output by integrating multi-dimensional data are more accurate and better suited to actual talent management needs. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating an artificial intelligence-based talent evaluation and management method provided by an embodiment of the present invention.
[0027] Figure 2 This is a schematic diagram of an artificial intelligence-based talent evaluation and management system provided by an embodiment of the present invention. Detailed Implementation
[0028] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0029] like Figure 1 As shown in the figure, an embodiment of the present invention proposes a talent evaluation and management method based on artificial intelligence, the method comprising the following steps:
[0030] Step 1: Obtain the resume data of the candidates to be evaluated, and perform text recognition processing on the resume data to generate a structured text dataset;
[0031] Step 2: Extract keyword feature vectors from the structured text dataset, calculate the similarity value between the keyword feature vectors and the preset talent evaluation keyword database, and generate initial similarity data;
[0032] Step 3: Construct a 3D point cloud dataset based on the initial similarity data, where each data point contains the weight coordinate information of the keyword;
[0033] Step 4: Determine the reference coordinate points based on the spatial density distribution of the point cloud dataset, and generate two spatial vector data based on the reference coordinate points, and calculate the boundary data of the included angle between the spatial vector data.
[0034] Step 5: Define the range of the sector area using the included angle boundary data, select the first spatial coordinate data within the sector area, and symmetrically select the second and third spatial coordinate data outside the range;
[0035] Step 6: Generate arc trajectory data according to the selection order of the first, second, and third spatial coordinate data; analyze the curvature feature value of the arc trajectory data to generate dynamic correction coefficients; integrate the initial similarity data and dynamic correction coefficients to output the final talent management score.
[0036] In this embodiment of the invention, unstructured resumes are transformed into standardized structured datasets through text recognition, avoiding information extraction bias caused by the chaotic format of the original resumes. Core features are extracted from the structured data and compared with a preset keyword library, achieving precise matching between talent characteristics and evaluation standards, reducing the subjectivity of manual screening, and making the initial evaluation results more objective and targeted. Multi-dimensional information such as keyword weight, similarity, and frequency of occurrence are mapped to three-dimensional coordinates, intuitively presenting the complex relationships of talent characteristics in a spatial manner, breaking through the limitations of single-dimensional evaluation and achieving a three-dimensional portrayal of talent traits. Determining the benchmark coordinates and calculating the angle between spatial vectors captures the core distribution patterns of talent characteristics, making the evaluation more aligned with the mainstream traits of talent. By selecting coordinates inside and outside the fan-shaped area, the matching degree between talent and standards (within the area) and the differentiation of traits (outside the area) are considered, ensuring the benchmark nature of the evaluation while identifying the unique advantages of talent, thus improving the comprehensiveness of the evaluation. Dynamic correction coefficients are generated through trajectory curvature analysis, and the initial results are adjusted in conjunction with time series relationships, enabling the evaluation to adapt to the dynamic changes in talent characteristics. Finally, the scores output by integrating multi-dimensional data are more accurate and better suited to actual talent management needs.
[0037] In a preferred embodiment of the present invention, step 1 includes:
[0038] Step 100: Perform paragraph semantic segmentation on resume data to identify the first text block containing educational experience, the second text block containing work experience, and the third text block containing skills and qualifications. Specifically, it includes: First, construct and train a semantic segmentation model. Select the pre-trained BERT-base model as the basic architecture, which contains 12 layers of Transformer encoders and a 768-dimensional word embedding dimension, and can effectively capture the deep semantic information of the text. Construct a three-class semantic segmentation model for resume text. The input of the model is a text sequence with a maximum length set to 512, and the output is the probability distribution of the paragraphs belonging to the three categories of educational experience, work experience, and skills and qualifications respectively. Preparation of the dataset before model training, collect resume data from public resume samples on the Internet, the enterprise's internal historical resume library, and resume data from multi-industry recruitment platforms, covering industries such as technology, management, and operation, including various formats such as PDF, Word, and TXT, and the cumulative sample size is not less than 100,000. Three annotators with a professional background in human resources manually annotate the paragraphs containing educational experience, work experience, and skills and qualifications in each resume according to the unified annotation specification. After the annotation is completed, a high-quality annotated dataset is formed through cross-validation (the consistency rate of the three annotators ≥ 95% is regarded as valid annotation). Perform data augmentation on the annotated dataset. For synonym replacement, use the WordNet thesaurus, and only replace non-core professional terms, such as replacing "be responsible for" with "lead" and "participate in" with "assist". The word order adjustment follows the principle that the subject-predicate-object grammar structure remains unchanged, and only adjusts the positions of attributives and adverbials. For text fragment recombination, select non-key information fragments from different resumes in the same category, such as the course learning description in educational experience and the daily affairs description in work experience, and splice them. Finally, expand the sample size to 2 times the original scale. Randomly divide the training set, validation set, and test set according to the ratio of 8:1:1, ensuring that the industry distribution and format type ratio of each dataset are consistent.
[0039] In the model training process, the first step is to preprocess the text paragraphs in the training set. Use the jieba word segmentation tool to perform word segmentation, filter out common stop words such as "de" and "shi", and generate a vocabulary sequence. Convert the vocabulary sequence into a word number sequence through the built-in word table of the model. Add the [CLS] identifier (used to extract paragraph-level features) at the head of the sequence and the [SEP] identifier (used to distinguish different texts) at the tail. For sequences with a length less than 512, fill them with the [PAD] identifier, and truncate sequences exceeding 512 at the tail to generate a standardized input sequence. The second step is to pass the standardized input sequence into the embedding layer of the BERT model, and generate a 768-dimensional initial embedding vector through the collaborative calculation of word embedding, position embedding, and segment embedding: The word embedding maps each word number to the corresponding 768-dimensional semantic vector by querying the pre-trained word table of the model to capture the semantic information of the vocabulary itself. The position embedding uses sine function encoding, where is the identifier of the position embedding vector, and the specific calculation formula is , in Indicates the position number of a word in the sequence. This represents the dimension index of the embedded vector, with values ranging from 0 to 383. =768,2 Corresponding to even-numbered dimensions +1 corresponds to an odd-numbered dimension. The hidden layer dimension of the model is fixed at 768 here, and this formula encodes the sequential positional relationship of words. The segment embedding is a fixed-dimensional vector of 768. Since only single-segment text is processed, all words are assigned the same fixed segment embedding vector to identify text belonging. The word embedding, position embedding, and segment embedding are added one by one according to their corresponding dimension elements to obtain a 768-dimensional word embedding vector that integrates word semantics and positional information. In the third step, the word embedding vector is input into a 12-layer cascaded Transformer encoder for deep feature extraction. Each encoder layer contains a multi-head self-attention mechanism and a feedforward neural network core structure. Residual connections and layer normalization (normalization dimension is 768) are added to the input and output of each layer. In the multi-head self-attention mechanism, Q represents the query vector, K represents the key vector, and V represents the value vector. The 768-dimensional input vector is first split into 12 attention heads with 64 dimensions / head. The attention weight of each head is calculated in parallel. The weight is obtained by normalizing the dot product of Q and K. The attention calculation function is given by the following formula: Here, dk represents the dimension of a single attention head, fixed at 64, and Softmax is the normalized activation function used to map weight values to the [0, 1] interval, capturing semantic associations of words in different dimensions through this calculation. The output vectors of the 12 attention heads are concatenated and mapped back to 768 dimensions through a linear transformation layer to obtain the multi-head attention output. The feedforward neural network is executed according to the linear transformation, GELU activation, and linear transformation. GELU is the Gaussian error linear unit activation function used to introduce non-linear feature expression. The first linear transformation layer maps the 768-dimensional vector to 3072 dimensions. After GELU activation, the second linear transformation layer maps the 3072-dimensional vector back to 768 dimensions, enhancing the non-linear fitting ability of the features. The 12-layer encoder processes the data layer by layer, ultimately processing the input order. The paragraph-level semantic feature representation is extracted from the output vector corresponding to the [CLS] identifier in the column header. This vector aggregates the global semantic information of the entire paragraph. In the fourth step, a two-layer fully connected network is constructed in the model output layer to map the paragraph-level semantic feature representation to the probability distribution of three target categories. The first fully connected layer maps the 768-dimensional semantic feature vector to 256 dimensions and uses ReLU (Linear Rectified Activation Function) to filter effective features and suppress the propagation of invalid information. The second fully connected layer further maps the 256-dimensional vector to 3 dimensions (corresponding to the three categories of education experience, professional experience, and skills and qualifications). Then, the Softmax normalized activation function is used to process the 3-dimensional vector so that the three output values are all in the range [0, 1] and the sum is 1, which correspond to the probability distribution of the paragraph belonging to the three categories.
[0040] The cross-entropy loss function is used to calculate the loss value between the predicted probability distribution and the ground truth labeled category. The ground truth labeled category adopts a one-hot encoding form, such as education experience corresponding to [1, 0, 0], professional experience corresponding to [0, 1, 0], and skills and qualifications corresponding to [0, 0, 1]. The loss value is calculated as follows: for each sample, the predicted probability of each of the three categories is multiplied by the corresponding ground truth label value, and the natural logarithm is taken. The three logarithm results are summed and the negative value is taken. Then, the arithmetic mean of the loss values of all samples in the training set is calculated to obtain the total loss value of a single training round. The AdamW optimizer is used to minimize the loss value, with the weight decay coefficient set to 0.01, the batch size set to 32, and the initial learning rate set to 2e-5. The learning rate is adjusted every two training rounds. The proportional gain decays exponentially by 0.9. An early stopping strategy is implemented: if the loss value on the validation set does not decrease for three consecutive rounds and the fluctuation is less than 1e-4, training is stopped and the current optimal model parameters are saved. After training, accuracy, precision, and recall are used as core evaluation metrics. Accuracy is the ratio of correctly classified paragraphs to the total number of paragraphs; precision is the ratio of correctly classified paragraphs in a category to all paragraphs predicted as positive in that category; and recall is the ratio of correctly classified paragraphs in a category to the actual number of paragraphs in that category. The precision and recall for all three categories are required to be no less than 92%, and the overall accuracy no less than 94%, ensuring the model meets the accuracy requirements for resume paragraph classification.
[0041] After model training, paragraph semantic segmentation is performed. First, the acquired resume data is preprocessed using regular expressions to remove redundant spaces, line breaks, special symbols, and non-textual interference such as images and tables. All resume text is then uniformly converted to UTF-8 encoding to eliminate processing biases caused by differences in resume formats. Next, the resume data is split into several independent paragraphs based on sentence-ending punctuation marks such as periods, semicolons, and exclamation marks, as well as paragraph markers such as line breaks and paragraph separators, constructing an initial paragraph set. Paragraphs shorter than 5 characters are considered invalid and discarded. For each valid paragraph in the initial set, the trained semantic segmentation model is input to extract paragraph-level semantic feature vectors. The calculation process involves first extracting the word embedding vectors of each core word within the paragraph from the model's embedding layer output (core words are selected using the TF-IDF algorithm, taking the top 80% of words by TF-IDF value), and then calculating the word frequency weight of each core word (word frequency weight = number of times the word appears in the paragraph ÷ total number of words in the preprocessed paragraph, where the total number of words is...). The total vocabulary after removing stop words is calculated. The word embedding vector of each core word is multiplied by its corresponding word frequency weight, and then weighted and summed. The sum is then divided by the total number of core words to obtain the average word embedding vector, which serves as the semantic feature vector for the paragraph. Three semantic template vectors are pre-constructed: educational experience, professional experience, and skills / qualifications. These template vectors are obtained by calculating the average of the semantic feature vectors of all labeled sample paragraphs under the corresponding category. The cosine similarity is calculated between the semantic feature vector of the current paragraph and each of the three template vectors. The cosine similarity is calculated as follows: Let vector A... The dot product of the paragraph semantic feature vector and vector B (template vector) is A1×B1+A2×B2+...+An×Bn (n is the vector dimension, i.e., 768). Calculate the magnitude of vector A and the magnitude of vector B. Cosine similarity = dot product ÷ (magnitude of vector A × magnitude of vector B). Take the category with the largest cosine similarity among the three categories as the category to which the paragraph belongs. Integrate all paragraphs belonging to educational experience into the first text block, paragraphs belonging to professional experience into the second text block, and paragraphs belonging to skills and qualifications into the third text block.
[0042] Step 101: Map the first text block to a pre-defined structured field for educational background, the second text block to a structured field for work experience, and the third text block to a structured field for professional skills. This generates a structured text dataset consisting of structured field labels and bound text content. Each structured field label contains a field type identifier, specifically including: three pre-constructed sets of standardized structured fields. The educational background structured field includes six subfields: educational level, graduating institution, major, graduation time, degree type, and whether it is full-time. The work experience structured field includes seven subfields: employer, job title, work period, industry, job responsibilities, project experience, and achievements. The professional skills structured field includes six subfields: skill name, skill proficiency, qualification certificate name, certificate issuing institution, certificate acquisition time, and skill application scenario. For the first text block, the BiLSTM-CRF named entity recognition method is used to extract education-related entity information. The BiLSTM layer captures text context dependencies, and the CRF layer optimizes entity boundary recognition accuracy. The extracted entity information is categorized by semantic attributes: undergraduate, master's, doctoral, etc., corresponding to academic level subfields; XX university, XX college, etc., corresponding to graduating institution subfields; computer science and technology, business administration, etc., corresponding to major name subfields; and time expressions such as June 2020, 2018-2022, etc. The mapping between the first text block and the structured field of education background is completed by using the graduation time subfield. For the second text block, an event extraction algorithm based on trigger words is used to extract career-related event elements. Trigger words include high-frequency words in career scenarios such as holding a position, being responsible for, participating in, and completing tasks. The core of the event is located through trigger words, and then the subject (corresponding to the employing unit and job title), time (corresponding to the work period), content (corresponding to job responsibilities and project experience), and result (corresponding to performance achievements) are extracted and mapped to the corresponding subfields of the work experience structured field. For the third text block, skill and qualification-related information is extracted by combining TF-IDF keyword extraction and entity recognition technology. The TF-IDF algorithm extracts skill keywords such as Python and project management, and entity recognition technology extracts qualification entities such as professional qualification certificates and technical certifications. Skill keywords are mapped to skill proficiency subfields according to descriptions such as proficiency, skilled, and mastery. Qualification entities are split and mapped to subfields such as qualification certificate name and issuing institution. Each structured field is assigned a unique field type identifier, consisting of 6 characters. The first two characters are the category code (01 for education background, 02 for work experience, 03 for skill qualification), and the last four characters are the field's own sequence code (assigned sequentially according to the subfield's order in the category field set, such as 0001, 0002, ...).For example, the identifier for the educational background subfield "Educational Level" is 010001, and the identifier for the work experience subfield "Job Title" is 020002. Finally, all structured fields bound to the corresponding text content and their corresponding field type identifiers are integrated to form a uniformly formatted and semantically clear structured text dataset.
[0043] This embodiment achieves precise separation of key talent information from different dimensions through paragraph semantic segmentation, avoiding extraction interference caused by the mixing of various types of information; it maps unstructured text blocks to standardized structured fields, eliminating information bias caused by the chaotic format and inconsistent expression of the original resume, making the presentation of talent information more standardized; the binding design of structured fields and field type identifiers reduces human intervention in information interpretation throughout the process, reduces the influence of subjective factors, and makes the basic data for talent evaluation more objective.
[0044] In a preferred embodiment of the present invention, step 2 includes:
[0045] Step 200: For each structured field label in the structured text dataset, parse the text content bound to the field and extract the core keywords. Specifically, it includes: For each structured field label in the structured text dataset, first preprocess the text content bound to it, which is specifically divided into two steps to remove useless words. The first step is to remove Chinese general stop words. The stop word list used is the Harbin Institute of Technology Chinese general stop word list, which contains words such as "的", "了", "在", "和", "及" that have no substantial semantic meaning in daily language and only play a grammatical connection role. Removing such words can avoid interference from meaningless information. The second step is to remove the special stop words corresponding to the field type. These stop words are generated by artificial intelligence through special mining of the industry corpus corresponding to the field type identifier, that is, first collect relevant texts in the industry to which the field belongs. For example, for the education background field, it corresponds to resumes in the education industry and text of college introductions. For the work experience field, it corresponds to job description texts for various industry positions. Count the occurrence frequencies of all words in these texts, screen out the top 5% of high-frequency words in terms of frequency, and then through manual sampling verification (the sampling ratio is 10% of the total amount of high-frequency words),剔除其中对字段核心信息,如教育背景的学历、院校,工作经历的职责、业绩无实际语义贡献的词汇,最终形成各字段的专用停用词表,例如教育背景字段的专用停用词包括等、相关、各类,工作经历字段的专用停用词包括进行、开展、、完成、推进;完成停用词去除后,采用jieba分词工具结合行业专用词库(教育背景字段包含双一流、985工程、211工程、硕士研究生等,工作经历字段包含项目负责人、跨部门协作、业绩达成等,专业技能字段包含机器学习、Java开发、SQL优化等)拆分复合词汇为基础词汇,进一步过滤长度小于2个字符的孤立无义词汇,如某、个、本等,这类词汇未被纳入停用词表但无独立语义,生成标准化词汇序列;将该标准化词汇序列输入微调后的BERT模型,输出每个词汇的768维深层语义向量,该向量由模型通过12层Transformer编码器逐层捕捉词汇上下文语义关联后生成,能反映词汇在职业场景中的含义,同时人工智能基于该字段类型标识符对应的标注文本数据,挖掘字段核心语义方向,如教育背景字段聚焦学历、院校、专业、学位相关语义,工作经历字段聚焦任职单位、岗位、职责、业绩相关语义,专业技能字段聚焦技能术语、资质名称、应用场景相关语义,通过计算标注文本中所有核心词汇(经人工标注的字段关键信息词汇)语义向量的算术平均值,生成字段核心语义向量;计算每个词汇的语义向量与字段核心语义向量的余弦相似度,筛选相似度高于0. It should be noted that there seems to be some incomplete or incorrect expressions in the original text (such as "剔除其中对字段核心信息……无实际语义贡献的词汇" and "、相关、各类" etc. which are not clearly presented). The above translation is based on the existing text as accurately as possible.A threshold of 7 words was selected as candidate keywords. This threshold was determined by artificial intelligence through 5-fold cross-validation to ensure that the selected words were highly relevant to the core semantics of the field. Subsequently, the artificial intelligence used an association rule mining algorithm to count the co-occurrence frequency of the candidate keywords in the text content, that is, the number of times they appeared together with other candidate keywords and the association strength with the core semantics of the field (reusing the cosine similarity calculated earlier). Redundant words with a co-occurrence frequency of less than 2 and an association strength of less than 0.75 were eliminated (the co-occurrence frequency threshold was determined by statistically analyzing the expression habits of the core information of the field, and the association strength threshold was increased based on the candidate keyword selection threshold to further ensure the accuracy of the keywords). Determine the core keywords under this field tag. For example, the core keywords for the Education Background field could be Bachelor's, Master's, PhD, 985 University, Double First-Class University, Computer Science and Technology, Business Administration, Engineering Degree, or Science Degree; the core keywords for the Work Experience field could be Project Manager, Technical Supervisor, Product Manager, Cross-departmental Collaboration, Project Coordination, 30% Annual Sales Growth, Team Size Expanding to 20 People, and Core Business Development; and the core keywords for the Professional Skills field could be Python Programming, Deep Learning, Big Data Analysis, SQL Optimization, PMP Certification, CFA Certificate, AWS Solutions Architect, E-commerce Platform Development, and Supply Chain Optimization.
[0046] Step 201: Generate a numerical feature vector based on the frequency weight and position weight of each core keyword in the text content. Specifically, this includes: first, calculating the frequency weight using the formula: Frequency Weight = Number of times the core keyword appears in the text content ÷ Total number of words in the preprocessed text content (total number of words is the total number of words after removing stop words and meaningless words; for example, if the total number of words in the preprocessed field of educational background is 50, and the keyword "Master's degree" appears 3 times, then the frequency weight = 3 ÷ 50 = 0.06); the position weight is calculated through the following steps: First, determine the total number of words in the paragraph containing the core keyword (the number of words after removing stop words and meaningless words, denoted as N), and record the position index of the keyword in the paragraph (from the paragraph...). The process involves four steps: 1) Numbering keywords sequentially from beginning to end, starting with 1 (denoted as P); 2) Calculating the relative position value: Position Index P ÷ Total number of words in the paragraph N. If a keyword appears multiple times, the average of all position indices is used to calculate the relative position value; 3) Using artificial intelligence, the relevance data between the positions of words in the resume and the core requirements of the job is analyzed. Fixed position importance benchmark coefficient tables are established for three fields: Education Background, Work Experience, and Professional Skills (the benchmark coefficient for Education Background is 1.2, for Work Experience is 1.3, and for Professional Skills is 1.4. These coefficients are determined based on the common expression patterns of core information in different fields; for example, the core responsibilities of work experience are often found at the beginning of paragraphs, hence the benchmark coefficient is slightly higher); 4) Combining the keyword with the core requirements of the fields... The semantic association strength, i.e., the cosine similarity calculated in step 200, is denoted as S. The correction coefficient is calculated as: Association Strength S × 0.8 + 0.2. Here, 0.8 is the weighting coefficient for association strength, used to amplify the impact of the keyword's semantic fit on the positional weight (the higher the fit, the closer the correction coefficient is to 1, and the more the positional weight reflects the importance of the keyword). 0.2 is a basic protection coefficient, used to prevent the correction coefficient from approaching 0 due to a low association strength S (e.g., when S = 0.1, 0.1 × 0.8 = 0.08; adding 0.2 results in a correction coefficient of 0.28). This ensures that even if the keyword's association with the core semantics is weak, it still retains a basic contribution to the positional weight and is not directly ignored. This formula allows the correction coefficient to be controlled within a certain range. The value should be between 0.2 and 1.0 to avoid a correction coefficient of 0 due to excessively low association strength. The fifth step, final position weight, is calculated as: relative position value × position importance baseline coefficient × correction coefficient. If the result is greater than 1, take 1; if less than 0, take 0. Ensure the value is controlled between 0 and 1. For example, for a work experience field, the core keyword "project leader" has a total word count N=40, position index P=5, relative position value = 5÷40=0.125, association strength S=0.85, correction coefficient = 0.85×0.8+0.2=0.88, and position weight = 0.125×1.3×0.88=0.143. The basic feature value of a single keyword is calculated as: word frequency weight × 0.6 + position weight × 0.4, where 0.6 and 0...4 represents the weighting coefficients obtained through 5-fold cross-validation optimization using artificial intelligence (the labeled data is split into training and validation sets in an 8:2 ratio, and the coefficients are adjusted to optimize the accuracy of subsequent feature vector matching, ultimately determining this ratio). These coefficients balance the influence of word frequency and position on word importance. For example, in the previous example, the word frequency weight was 0.06, and the position weight was 0.143, so the basic feature value = 0.06 × 0.6 + 0.143 × 0.4 = 0.0932. This basic feature value is then added as a new dimension and concatenated dimensionally with the 768-dimensional deep semantic vector of the keyword output by the fine-tuned BERT model. This means that dimensions 1 to 768 of the semantic vector are combined, and the new basic feature value is added as the 769th dimension, generating a 769-dimensional numerical feature vector. This vector simultaneously encompasses statistical features and deep semantic features extracted by artificial intelligence, achieving a comprehensive representation of the keyword.
[0047] Step 202 involves performing similarity matching between the feature vectors and the standard feature vectors under the same field type identifier in the preset talent evaluation keyword library, and outputting the similarity value of the corresponding keywords. Specifically, the preset talent evaluation keyword library is stored according to field type identifiers. Each standard keyword in the library comes from industry standard documents, job qualification specifications, and high-quality recruitment demand texts. For example, the education background field includes doctoral degree, "Double First-Class" university, computer science major, etc.; the work experience field includes project management, teamwork, exceeding performance targets, etc.; and the professional skills field includes Python programming, deep learning, PMP certification, etc. The standard feature vector generation logic for each standard keyword in the library is consistent with step 201: first, a 768-dimensional semantic vector is extracted using the same fine-tuned BERT model, and then... The AI learns industry importance coefficients from industry job demand data and historical recruitment matching data. (The differences in the importance requirements of keywords for different industries and positions are learned autonomously by the AI. For example, the coefficient for Python programming in the information technology industry is 1.5, and the coefficient for office software operation is 0.8; the coefficient for instructional design in the education industry is 1.4, and the coefficient for research ability is 1.3.) The semantic vector is then weighted and adjusted, that is, the value of each dimension of the semantic vector is multiplied by the corresponding industry importance coefficient, and finally a 769-dimensional standardized feature vector is generated. The cosine similarity between the 769-dimensional numerical feature vector of the core keyword to be evaluated and the corresponding standard feature vector is calculated. This cosine similarity directly reflects the degree of fit between the keyword to be evaluated and the industry standard keyword, and the value ranges from 0 to 1.
[0048] Step 203: Aggregate keywords and corresponding similarity values under all structured field tags to generate initial similarity data containing field type identifiers, keywords, and similarity values. Specifically, this includes: using an AI-weighted aggregation strategy, statistically analyzing historical recruitment matching data for the target position over the past 3 years (selecting resumes of hired personnel and calculating the ratio of the average occurrence frequency of the core keyword in this field to the average occurrence frequency of the core keywords across all fields, denoted as F1), industry job standard documents, calculating the proportion of this field in the document's chapters, denoted as F2 (chapter proportion = number of words in the relevant chapters ÷ total number of words in the document), and fields in the recruitment requirement text. The frequency of mention (the ratio of the number of times related words in the recruitment requirements to the total number of words, denoted as F3) is used to determine the importance weight of different field type identifiers. The specific calculation method is: the importance weight of a field = (F1 × 0.5) + (F2 × 1.0 × 0.3) + (F3 × 0.2), and the sum of the importance weights of all field type identifiers is normalized to 1. For example, for the technical job professional skills field, F1 = 0.6, F2 = 0.4, F3 = 0.5, then the weight = 0.6 × 0.5 + 0.4 × 0.3 + 0.5 × 0.2 = 0.52; for the education background field, F1 = 0.2, F2 = 0.2, F3 = 0.5 ... If 3 = 0.1, then the weight = 0.2 × 0.5 + 0.2 × 0.3 + 0.1 × 0.2 = 0.18, and the weight of the work experience field = 1 - 0.52 - 0.18 = 0.3. For all core keywords under each field label, the basic feature value of each keyword is used as the local weight. The keyword similarity value is multiplied by the corresponding local weight and then summed to obtain the field-level similarity. For example, a professional skills field has 2 core keywords. Keyword 1 has a basic feature value of 0.0932 and a similarity value of 0.85, and Keyword 2 has a basic feature value of 0.07 and a similarity value of 0.9. Then the field-level similarity = 0.0932 × 0.52 = 0.18. 0.85 + 0.07 × 0.9 = 0.14222; Subsequently, the field type identifiers of all structured field tags are integrated, such as Education Background 01, Work Experience 02, Professional Skills 03, and various core keywords, such as Master's Degree, Project Management, Python Programming, corresponding similarity values, such as 0.88, 0.92, 0.85, and field-level similarities, such as 0.18, 0.3, 0.52, and the corresponding field-level similarity of 0.14222, etc., and organized according to the unified format of field type identifier, core keywords, keyword similarity values, and field-level similarities to generate initial similarity data with standardized structure and complete information.
[0049] This embodiment breaks through the limitation of keyword extraction relying solely on surface-level word frequency, capturing the deep association between words and the core semantics of fields, thus improving the targeting and effectiveness of core keywords. Word frequency weight, position weight, and various matching coefficients are all determined through artificial intelligence learning or cross-validation optimization. The position weight is assigned through multi-dimensional quantitative calculation, avoiding the subjectivity of manually set parameters, so that the numerical feature vector can more objectively reflect the importance and semantic value of keywords. The standard feature vector is generated by artificial intelligence combined with industry needs and historical data training, and is based on the same model architecture and generation logic as the feature vector to be evaluated, ensuring the consistency and rationality of similarity matching and reducing the bias caused by surface comparison. The artificial intelligence weighted aggregation strategy determines the field weight based on historical recruitment data, industry standard documents, and recruitment requirement text statistics, taking into account both field priority and the local importance of keywords. This ensures that the initial similarity data not only fully retains key details, but also clearly reflects the differentiated contribution of different fields and keywords to talent evaluation, further improving the objectivity, accuracy, and targeting of talent evaluation.
[0050] In a preferred embodiment of the present invention, step 3 includes:
[0051] Step 300: For each keyword in the initial similarity data, based on the field type identifier, use the preset field semantic weight coefficient as the X-axis coordinate value, and read the similarity value of the corresponding keyword as the Y-axis coordinate value. Specifically, this includes: First, a preset field semantic weight coefficient library needs to be constructed. The core function of this coefficient library is to quantify the semantic importance of different field types to the evaluation of target job talents. The value range of all coefficients is strictly limited to the range of 0 to 1. The generation process of this coefficient library is completed by artificial intelligence. The specific process is as follows: First, artificial intelligence collects three categories of core information, namely, real job requirement texts across industries and multiple positions, authoritative talent evaluation standard documents, and historical recruitment matching successful cases in the past 3 years. This data is cleaned to remove duplicate, invalid, and logically contradictory data. Key information directly related to the field type is extracted from it, including the frequency of mention of each field in the job requirements, the description of the core requirements, the impact record of each field on the recruitment results in historical cases, and the chapter proportion and indicator weight allocation of each field in the authoritative standard document.
[0052] The second step involves AI-based feature quantification and correlation analysis to calculate multi-dimensional talent evaluation contribution indicators for each field. These indicators include: 1) the frequency of mentions of job requirements, calculated as the total number of times a single field is mentioned in all job requirement texts divided by the total number of mentions of all fields; 2) the correlation strength between job requirements, calculated as the number of overlapping words between the field description and the core job responsibility keywords divided by the total number of words in the field description—higher overlap indicates stronger correlation; and 3) the historical case decision contribution coefficient, calculated as (the number of cases where the field information met the criteria and were ultimately hired divided by the total number of hired cases) minus (the number of cases where the field information did not meet the criteria but were still hired divided by the total number of hired cases), multiplied by a normalization adjustment. The coefficient 1.25 is normalized to ensure that the value range is strictly between 0 and 1 (0 for negative results and 1 for greater than 1). The higher the coefficient, the greater the influence of the field on the recruitment decision. Fourth, the authoritative standard importance quantification result is calculated as follows: First, two basic indicators are calculated separately. The first is the total number of words in the chapter corresponding to the field divided by the total number of words in the standard document. The second is the sum of the weights of all indicators included in the field (the weight of a single indicator is between 0 and 1, and the sum of the weights of all indicators in the authoritative standard document is 1), divided by the sum of the weights of all indicators in the document (fixed to 1). Then, the arithmetic mean of the results of the two basic indicators is taken to obtain the final authoritative standard importance quantification result.
[0053] The third step involves AI-driven multi-source data fusion analysis and coefficient optimization. Based on the aforementioned four-dimensional contribution indicators (the proportion of job requirement mentions, the strength of job requirement association, the contribution coefficient of historical case decisions, and the quantitative results of the importance of authoritative standards), a multi-source data fusion analysis framework is constructed. Supervised learning analysis is conducted with the actual impact of fields on talent evaluation as the objective. First, the four-dimensional indicators are weighted and summed according to weights of 22%, 30%, 28%, and 20% (the total weight sum is 1, set based on data distribution characteristics and impact priority) to obtain a preliminary comprehensive contribution. Then, the weight proportions of each dimension are adjusted using a 5-fold cross-validation method. Specifically, the dataset is randomly divided into 5 groups, each group serving as the validation set in turn, and the rest as the training set. The error of the comprehensive contribution under different weight combinations is calculated using mean squared error (MSE). The weight combination with the smallest MSE value is retained to reduce the impact of data bias and ensure the stability and reliability of the output coefficients.
[0054] The fourth step is to divide the comprehensive contribution of each field after multi-source fusion by the sum of the comprehensive contributions of the three fields to complete the normalization process, ensuring that the final sum of coefficients is 1 and that each individual coefficient is in the range of 0 to 1. Finally, the semantic weight coefficients corresponding to each field type are determined, namely, the education background field (01) is 0.3, the work experience field (02) is 0.4, and the professional skills field (03) is 0.3. This allocation ratio is consistent with the comprehensive demand ratio of work ability, professional skills and education background for most positions. Work experience can directly reflect the candidate's actual job performance ability, and its comprehensive contribution is the highest, so its coefficient is the highest. Education background and professional skills are important manifestations of basic qualifications and core abilities, respectively, and their comprehensive contributions are similar, so the coefficients are balanced and matched.
[0055] After obtaining each keyword from the initial similarity data, the field type identifier associated with the keyword is first extracted. The identifier is then precisely matched with a preset field semantic weight coefficient library. The semantic weight coefficient of the successfully matched field is directly assigned as the X-axis coordinate value of the keyword's three-dimensional coordinates. This visually reflects the evaluation priority of the field to which the keyword belongs in space. At the same time, the similarity value of the keyword, which has been calculated in step 202, is directly read and used as the Y-axis coordinate value of the three-dimensional coordinates. This value directly reflects the degree of fit between the keyword and industry standard keywords, achieving fast and accurate assignment of X-axis and Y-axis coordinates.
[0056] Step 301: Based on the original frequency of the corresponding keyword in the text content of the corresponding field in the structured text dataset, calculate the normalized frequency value as the Z-axis coordinate value. Specifically, this includes: First, counting the original frequency of each keyword, using the text content bound to the field label of the keyword as the statistical range. This text must be the standardized text that has been preprocessed in step 200 (removing general Chinese stop words, field-specific stop words, and meaningless words) to avoid redundant information interfering with the statistical results; By traversing the standardized text line by line, the total number of times the target keyword actually appears in the text is accumulated, and this number is defined as the original frequency F; To ensure that the Z-axis coordinate is consistent with the X-axis and Y-axis coordinates (both in the range of 0 to 1), the min-max normalization method is used to process the original frequency F, and the normalized frequency value is calculated as the Z-axis coordinate value. The specific calculation process is divided into three steps. The first step is to determine the extreme value of the original frequency of all core keywords under the same field type, that is, to traverse all core keywords under the field label and select the maximum value of the original frequency. and minimum value The second step is to combine the original frequency F of the target keyword with the frequency of the same field. and Substituting into the normalization formula, the normalized frequency value Z = (original frequency F - minimum value) ) ÷ (maximum value) - Minimum value The formula maps raw frequencies of different magnitudes to a uniform range of 0 to 1. The third step handles special cases where all core keywords within the same field have identical raw frequencies. = At this point, the denominator of the normalization formula is 0, which will cause the calculation to fail. In order to ensure the basic distinguishability of this dimension, the Z-axis coordinate value of all keywords under this field is directly set to 0.5.
[0057] Step 302 involves defining the X, Y, and Z coordinates of each keyword as a three-dimensional spatial point. The spatial point coordinates of all keywords are then integrated to form a three-dimensional point cloud dataset. Specifically, after determining the X, Y, and Z axis coordinates, the three sets of coordinate values corresponding to each keyword are combined to define a three-dimensional spatial point (X, Y, Z). The X-axis coordinate reflects the evaluation importance of the field to which the keyword belongs, the Y-axis coordinate reflects the degree of conformity of the keyword to industry standards, and the Z-axis coordinate reflects the density of the keyword's occurrence in the text. These three coordinates together constitute a three-dimensional representation of the core features of the keyword. Subsequently, all keywords in the initial similarity data are traversed, and a corresponding three-dimensional spatial point is generated for each keyword according to the above rules. After all the three-dimensional spatial points for all keywords have been generated, the coordinate information of all spatial points is systematically integrated according to a unified format of field type identifier, keyword, X-axis coordinate, Y-axis coordinate, and Z-axis coordinate, ensuring the structural standardization and information completeness of each data entry. The final three-dimensional point cloud dataset comprehensively encompasses the multi-dimensional core information of all keywords.
[0058] This embodiment quantifies the differences in importance of different fields to talent evaluation by pre-setting semantic weight coefficients for fields and mapping them to X-axis coordinates, so that the spatial position of keywords can reflect the evaluation priority; using similarity values as Y-axis coordinates, it directly relates the degree of fit between keywords and industry standards, preserving the core results of previous matching; using normalized frequency values as Z-axis coordinates, it reflects the information density of keywords in the text and supplements the regularity of keyword occurrence; the construction of three-dimensional coordinates realizes the spatial integration of multi-dimensional information of keywords, breaks through the limitations of single-dimensional evaluation, and can intuitively present the complex association and distribution patterns of talent characteristics.
[0059] In a preferred embodiment of the present invention, step 4 includes:
[0060] Step 400: Divide the 3D point cloud dataset into spatial grids, count the number of data points in each grid cell, mark grid cells with more than a preset threshold as dense cells, and extract the data points contained in all dense cells to form a candidate point set. Specifically, this includes: first, determining the coordinate range of the 3D point cloud dataset, that is, counting the maximum and minimum values of the X-axis, Y-axis, and Z-axis coordinates of all spatial points to obtain the X-axis range. Y-axis range Z-axis range Based on this coordinate range, a uniform spatial grid is generated. The side length of each grid cell in the X, Y, and Z dimensions is uniformly set to 0.1 (the side length value is based on the distribution range of coordinate values from 0 to 1, ensuring that the grid can accurately capture the distribution density of data points, while avoiding computational redundancy caused by overly fine division). After the division is completed, the number of data points contained in each grid cell is counted. All three-dimensional spatial points are traversed, and the grid cell index to which each point belongs is calculated using the coordinate values. Specifically, for a point with coordinates (X, Y, Z), its grid index in the X-axis direction is calculated. =floor[ ], grid index in the Y-axis direction =floor[ ], Z-axis grid index =floor[ ], where floor is the floor function. All are non-negative integers, and their values correspond to the number of grid cells in the X, Y, and Z axes, respectively. For example, the number of grid cells in the X-axis direction is floor[ +1, The maximum value is this number minus 1, ensuring the index does not exceed the grid division range; according to Determine the unique grid cell to which a point belongs, and accumulate the number of points contained in each grid cell.
[0061] A preset threshold of 3 was set. This threshold was determined by artificial intelligence through multi-step refined analysis. Specifically, the process involved: first, generating 3D point cloud datasets from different industries and job types, covering technical, managerial, and functional roles; then, dividing each dataset into identical spatial grids (each with a 3D side length of 0.1), counting the number of points contained in all grid cells within each dataset, and calculating the mean, median, standard deviation, and 90th quantile of the grid points for each dataset; and finally, analyzing the spatial distribution characteristics of keywords corresponding to core information, revealing that the number of grid cell points in the core information set was generally higher than the 90th quantile of the overall data, and that the point distribution of this type of grid... Relatively stable; based on this rule, artificial intelligence first sets multiple candidate thresholds (2, 3, 4, etc.) to screen dense units in the validation set data, and calculates the proportion of core information points and the filtering rate of redundant information; finally, the minimum threshold with a core information point proportion of not less than 95% and a redundant information filtering rate of not less than 80% is selected as the final preset threshold 3, to ensure that the grids in the core information set can be accurately screened, while avoiding over-filtering of effective information; grid units with more than 3 data points are marked as dense units, and finally all data points contained in all dense units are extracted and integrated to form a candidate point set, realizing the initial separation of core information and redundant information.
[0062] Step 401: Construct a triangular mesh structure in the candidate point set, calculate the area-to-side-length ratio coefficient of each triangular unit, select all triangular units whose side-length ratio coefficients meet the preset conditions, and identify the target triangle with the longest side among all triangular units. Specifically, this includes: constructing a triangular mesh structure on the candidate point set using the Delaunay triangulation algorithm. This algorithm ensures that the generated triangular units are as close as possible to equilateral triangles, avoiding extremely long and narrow triangles. For each triangular unit, perform two key calculations. When calculating the triangle area, let the coordinates of the three vertices of the triangle be P1(X1, Y1, Z1), P2(X2, Y2, Z2), and P3(X3, Y3, Z3). First, calculate the vectors P1P2=(X2-X1, Y2-Y1, Z2-Z1) and P1P3=(X3-X1, Y3-Y1, Z3-Z1). Then calculate the cross product of the two vectors. The magnitude of the cross product is |(Y2-Y1)(Z3-Z1)-(Z2-Z1)(Y3-Y1), (Z2-Z1)(X3-X1)-(X2-X1)(Z3-Z1), (X2-X1)(Y3-Y1)-(Y2-Y1)(X3-X1)|. The area of the triangle is half the magnitude of the cross product. When calculating the side length ratio coefficient, calculate the length of each of the three sides of the triangle and take the ratio of the longest side length to the shortest side length as the side length ratio coefficient. Set the preset condition that the side length ratio coefficient is less than 3. This condition is based on the side length ratio characteristics of equilateral triangles to ensure that regular and stable triangular units are selected. Select all triangular units that meet this condition. Among these valid triangular units, compare the length of all the longest sides and determine the triangle with the longest side length as the target triangle.
[0063] Step 402: Calculate the distance weights from the three vertices of the target triangle on the longest side to the centroid of the point cloud dataset. Based on these distance weights, weightedly fuse the coordinates of the three vertices to obtain the fused coordinates, which are then used as the reference coordinate points. Specifically, this involves: first, calculating the centroid coordinates of the 3D point cloud dataset; traversing all spatial points and accumulating the sums of the X-axis, Y-axis, and Z-axis coordinates of all spatial points; then dividing the sums by the total number of spatial points to obtain the X-axis, Y-axis, and Z-axis coordinates of the centroid. These three coordinates together constitute the centroid coordinates. The calculation logic is as follows: Centroid X-axis coordinate = sum of X-axis coordinates of all spatial points ÷ total number of spatial points; Centroid Y-axis coordinate = sum of Y-axis coordinates of all spatial points ÷ total number of spatial points; Centroid Z-axis coordinate = sum of Z-axis coordinates of all spatial points ÷ total number of spatial points. Next, the distances and distance weights from the three vertices of the target triangle to the centroid are calculated. The three vertices of the target triangle are selected, and the X, Y, and Z coordinates of each vertex are recorded. The distance from each vertex to the centroid is calculated using the three-dimensional Euclidean distance formula, yielding the distances for each of the three vertices. The distance weights are normalized based on the reciprocal of the distance from each vertex to the centroid. To ensure the sum of the three weights is 1, the calculation is as follows: First, calculate the reciprocal of the distance to each vertex. Then, calculate the sum of the three reciprocals: the distance weight of the first vertex = the reciprocal of the first vertex's distance ÷ the sum of the three reciprocals; the distance weight of the second vertex = the reciprocal of the second vertex's distance ÷ the sum of the three reciprocals; the distance weight of the third vertex = the reciprocal of the third vertex's distance ÷ the sum of the three reciprocals. Finally, weighted summation of the three vertex coordinates based on the distance weights yields the baseline coordinate point. The X-axis coordinate of the baseline coordinate point = the X-axis coordinate of the first vertex × its distance weight + the X-axis coordinate of the second vertex. The X-axis coordinate is multiplied by its distance weight, and the X-axis coordinate of the third vertex is multiplied by its distance weight. The Y-axis coordinate of the reference point is equal to the Y-axis coordinate of the first vertex multiplied by its distance weight, the Y-axis coordinate of the second vertex multiplied by its distance weight, and the Y-axis coordinate of the third vertex multiplied by its distance weight. The Z-axis coordinate of the reference point is equal to the Z-axis coordinate of the first vertex multiplied by its distance weight, the Z-axis coordinate of the second vertex multiplied by its distance weight, and the Z-axis coordinate of the third vertex multiplied by its distance weight. This reference point integrates the spatial position characteristics of the vertices of the target triangle and the distance characteristics of each vertex from the centroid, and has strong core representativeness.
[0064] Step 403: Generate a first spatial vector by extending from the reference coordinate point in the positive X-axis direction of the point cloud dataset, and a second spatial vector by extending in the positive Y-axis direction; calculate the three-dimensional spatial angle between the first and second spatial vectors as the angle boundary data. Specifically, this includes: generating the first spatial vector by extending from the reference coordinate point in the positive X-axis direction. To facilitate angle calculation and eliminate the influence of vector length differences, the vector length is uniformly set to 1, so the coordinates of the first spatial vector are (1, 0, 0), that is, moving 1 unit from the reference coordinate point along the positive X-axis direction, while keeping the Y-axis and Z-axis coordinates consistent with the reference coordinate point; generating... The second spatial vector, also uniformly set to a length of 1, has coordinates (0, 1, 0), meaning it moves 1 unit along the positive Y-axis from the reference point, while maintaining its X and Z coordinates with the reference point. To calculate the three-dimensional spatial angle between the two spatial vectors, the first step is to calculate the dot product of the two vectors: multiply the X-axis component of the first spatial vector by the X-axis component of the second spatial vector, add the Y-axis component of the first spatial vector multiplied by the Y-axis component of the second spatial vector, and add the Z-axis component of the first spatial vector multiplied by the Z-axis component of the second spatial vector. The second step is to calculate the magnitude of the two vectors. Since both vectors have a uniform length of 1, their magnitudes are both 1 / 2. The result is 1 for all three vectors. The third step is to substitute the angle into the three-dimensional space formula, the angle θ = inverse cosine function [(dot product of two vectors) ÷ (magnitude of the first vector × magnitude of the second vector)], and finally obtain the angle value θ, which is used as the angle boundary data.
[0065] This embodiment, through spatial grid division and dense unit screening, can focus on the spatial points corresponding to core keywords, filter redundant and sparse information, and highlight the core characteristics of key dimensions of talent evaluation. Based on Delaunay triangulation, a triangular network is constructed and regular triangular units are selected to ensure the stability and representativeness of the target triangles, providing a reliable foundation for the calculation of benchmark coordinate points. Combined with a weighted fusion method based on centroid distance weights, the benchmark coordinate points integrate the location information of core spatial points while also taking into account the distribution characteristics of the overall point cloud, resulting in stronger comprehensive representativeness. By generating spatial vectors and calculating the three-dimensional angle, the correlation between the semantic importance of fields and the standard matching degree is quantified into a spatial angle index, further improving the scientific rigor and accuracy of talent evaluation.
[0066] In a preferred embodiment of the present invention, step 5 includes:
[0067] Step 500: Using the reference coordinate point as the vertex and the included angle boundary data as the expansion angle, define a sector-shaped spatial region in the XY plane. Specifically, this includes: using the reference coordinate point determined in step 402 as the vertex of the sector-shaped spatial region, the projection of this vertex onto the XY plane is the core position of the vertex of the sector. The Z-axis direction does not limit the range of the sector-shaped region (only focusing on the angular distribution of the XY plane). The included angle boundary data obtained in step 403 is used as the expansion angle of the sector. The central bisector of the sector is set as the angle bisector of the angle between the first spatial vector and the second spatial vector in step 403, that is, the direction angle of the central bisector is half of the included angle boundary data, ensuring that the sector-shaped region is symmetrically distributed on both sides of the angle bisector. Through the three core parameters of vertex, expansion angle and central bisector, the sector-shaped spatial region is clearly defined in the XY plane. This region covers the core spatial range with high correlation between field semantic importance and standard matching degree.
[0068] Step 501: Select a data point within the sector-shaped spatial region as the first spatial coordinate data. Using the central bisector of the sector-shaped spatial region as the axis of symmetry, select data points at symmetrical positions outside the region as the second and third spatial coordinate data, respectively. Specifically, this includes: First, selecting the first spatial coordinate data within the sector-shaped spatial region defined in step 500, clarifying the specific range of the sector-shaped region in the XY plane, using the XY plane projection of the reference coordinate point as the vertex, and the central bisector as the symmetrical midline (the direction is determined by the included angle boundary data in step 403, i.e., the midline is the middle position of the included angle), and expanding the left and right boundaries towards both sides of the midline, with the total expansion angle being the included angle boundary data in step 403. Only the XY plane projection falls on the vertex, the left boundary, and the right boundary. Points within the bounded area are spatial points within the fan-shaped region (Z-axis coordinates are unrestricted). All points meeting the criteria are iterated through, and the Euclidean distance between each point and the reference coordinate point's XY plane projection is calculated. The arithmetic mean of the distances between all points within the region is calculated, with a selection range set to 90% to 110% of this average. Simultaneously, the average Z-axis coordinate of the points within the region is calculated. Points whose Z-axis coordinate difference from this average is ≤0.05 are selected from the distance range. These points are located in the core area of the fan-shaped region, approximately 1 / 2 the distance from the reference point in the XY plane and close to the average height of the region in the Z-axis direction. One of these points is randomly selected as the first spatial coordinate data, ensuring it fits the center of the fan and represents the spatial distribution characteristics of most points in the region. Then, using the central bisector as the axis of symmetry, points are analyzed outside the fan-shaped region... Select symmetrical second and third spatial coordinate data. The specific steps are as follows: First, calculate the perpendicular distance from the first spatial coordinate data to the central bisector (based only on the XY plane). The central bisector is a straight line on the horizontal plane, with one end passing through the horizontal projection of the reference coordinate point (let the projected coordinates be x0, y0), and the other end extending along the symmetrical midline of the sector. The inclination is consistent with the symmetrical midline. Its linear equation can be expressed as ax + by + c = 0, where a and b are the direction coefficients of the line (together determining the inclination angle of the line; for example, when the line is horizontal, b = 0 and a is a non-zero constant; when the line is vertical, a = 0 and b is a non-zero constant; when the line is inclined, both a and b are non-zero constants), and c is the position constant (determining the specific position of the line on the horizontal plane). (Related to the projection of the reference point); the method for obtaining parameters a, b, and c is as follows: First, determine a and b based on the inclination direction of the symmetrical centerline of the sector. For example, if the inclination angle of the symmetrical centerline is known, its slope can be calculated, and then converted into the form ax + by = 0 (for example, when the slope is 2, we can set a = 2 and b = -1, satisfying the slope of the line as -a / b = 2); then, since the line must pass through the horizontal projection (x0, y0) of the reference point, substitute x0 and y0 into ax + by + c = 0 to calculate c, that is, lock the specific position of the line through the position of the reference point; when calculating the distance, first take the horizontal abscissa x1 and ordinate y1 of the first spatial coordinate data, substitute them into the above line equation, calculate the result of ax1 + by1 + c, and take the absolute value; then divide this absolute value by The final result is the reference distance, which accurately reflects the distance of the first spatial coordinate data from the central bisector on the horizontal plane. The second step is to screen candidate points on the other side of the central bisector (opposite to the first point, and whose XY projection is not within the sector). Four conditions must be met simultaneously: First, the planar range is limited, and the XY projection of the candidate point is outside the sector boundary and located on the other side of the axis of symmetry. Second, the planar distances are matched, and the difference between the Euclidean distance of the candidate point to the reference point and the first point is ≤5% (e.g., if the first point is 10 units away from the reference point, the candidate point must be between 9.5 and 10.5 units). Third, the vertical distances are symmetrical, and the vertical distance of the candidate point to the central bisector is exactly equal to the reference distance. Fourth, the Z-axis heights are consistent, and the difference between the Z-axis coordinate of the candidate point and the first point is ≤0.1, ensuring that the heights of the three points are in the same horizontal range.
[0069] This embodiment defines a sector-shaped spatial region using reference coordinate points and included boundary data. This allows for precise focusing on the core spatial range where the semantic importance of fields and the standard matching degree are highly correlated, filtering out redundant data from irrelevant spatial regions and improving the targeting of feature point selection. Using the central bisector as the axis of symmetry, three spatial coordinate data are symmetrically selected within and outside the region. This ensures that the first data point is representative of the core region's features, while also obtaining comparative features through symmetrical data points outside the region. This provides a comprehensive and balanced coordinate foundation for the subsequent construction of a three-dimensional arc-shaped trajectory, ensuring that the trajectory fully reflects the spatial correlation between core features and comparative features, and laying a reliable data foundation for subsequent curvature analysis.
[0070] In a preferred embodiment of the present invention, step 6 includes:
[0071] Step 600: Using the first spatial coordinate data as the starting point, the second spatial coordinate data as the intermediate point, and the third spatial coordinate data as the ending point, connect them in the selected order to generate a three-dimensional spatial arc trajectory. Specifically, this includes: using the first spatial coordinate data selected in step 501 as the trajectory starting point, the second spatial coordinate data as the trajectory intermediate point, and the third spatial coordinate data as the trajectory ending point, and using a three-point spherical arc fitting algorithm to generate a continuous and smooth three-dimensional spatial arc trajectory in the spatial order of starting point, intermediate point, and ending point. The specific process is as follows: First, extract the complete three-dimensional coordinates of the three points, denoted as the starting point (x1, y1, z1), intermediate point (x2, y2, z2), and ending point (x3, y3, z3), respectively. Calculate the coordinates of the center of the sphere (xc, yc, zc) and the radius R of the sphere where these three points are located. The core logic of constructing the equation system is that the straight-line distance from any point on the sphere to the center of the sphere is equal to the radius. Based on this, the following three equations are listed to form the equation system. (x2 - xc) 2 +(y2-yc) 2 +(z2-zc) 2 =R2 (x3 - xc) 2 +(y3-yc) 2 +(z3-zc) 2 =R 2 By solving this system of equations, the following can be eliminated: Then, a system of linear equations for xc, yc, and zc is obtained. The specific values of xc, yc, and zc are solved, and then substituted into any one of the equations to calculate R, ensuring that all three points are on the sphere. The second step is to determine the arc trajectory connecting the three points on the sphere. Based on the latitude and longitude positions of the three points on the sphere, the shortest spherical arc segment is selected as the main body of the trajectory, following the order from the starting point to the middle point and then to the ending point, ensuring that the trajectory is continuous, smooth, and of reasonable length. The third step is to calculate the spherical arc length from the starting point to the middle point and from the middle point to the ending point, respectively. The calculation logic is as follows: first, calculate the straight-line distance between the two points (three-dimensional straight-line distance), then, combined with the sphere radius R, use the cosine theorem to calculate the central angle at the center of the sphere. The formula is: central angle = arccos[(2R)] 2 - Straight-line distance between two points 2 )÷(2R 2 The length of a single spherical arc is calculated as radius R × corresponding central angle. Finally, the two arc lengths are added together to obtain the total arc length. Through this fitting process, the generated three-dimensional spatial arc trajectory is continuous and smooth, and accurately passes through the spatial positions corresponding to the three coordinate data, fully reflecting the spatial correlation characteristics of the three points.
[0072] Step 601: Calculate the instantaneous curvature feature value of the three-dimensional arc trajectory at the second spatial coordinate data position. The instantaneous curvature feature value is determined by the rate of change of the direction angle of the trajectory tangent at the second spatial coordinate data position and the offset of the trajectory normal vector. Specifically, the instantaneous curvature feature value is focused on the position of the second spatial coordinate data (midpoint) and calculated by weighted fusion of the rate of change of the trajectory tangent direction angle and the offset of the trajectory normal vector. The specific process is as follows: Calculate the trajectory tangent direction angle. First, determine the two tangent vectors before and after the midpoint. The front tangent vector is the tangent vector of the trajectory arc segment from the starting point to the midpoint at the midpoint (based on the definition of the tangent for a spherical arc, this tangent is perpendicular to the line connecting the center of the sphere and the midpoint, and points to...). (Towards the termination point), the tangent vector is the tangent vector at the midpoint of the trajectory arc segment from the midpoint to the termination point (similarly, this tangent is perpendicular to the line connecting the center of the sphere and the midpoint, and deviates from the direction of the starting point); then calculate the angle between each tangent vector and the X-axis, Y-axis, and Z-axis. Taking the angle with the X-axis as an example, take the ratio of the X-axis component of the tangent vector to the magnitude of the tangent vector (the square root of the sum of the squares of the components), and calculate the angle using the inverse cosine function. Similarly, obtain the angles between the two tangents and the Y-axis and Z-axis respectively, ultimately forming the three-dimensional direction angles of the two tangents; the direction angle change rate calculation, firstly, calculate the difference between the corresponding direction angles of the two tangents, that is, calculate the absolute value of the difference in the X-axis direction angle, the absolute value of the difference in the Y-axis direction angle, and the absolute value of the difference in the Z-axis direction angle respectively. The first step is to calculate the absolute value of the difference in direction angles. The second step is to calculate the total difference in direction angles by adding the squares of the three direction angle differences and taking the square root of the sum. The third step is to substitute the total arc length calculated in step 600. The rate of change of direction angles is equal to the total difference in direction angles divided by the total arc length. This value reflects the rate of change of the tangent direction at the midpoint. The third step is to calculate the trajectory normal vector offset. The first step is to calculate the trajectory normal vector based on the three-point fitted spherical arc (ideal trajectory). The normal vector of the midpoint is the vector connecting the center of the sphere to the midpoint (direction from the midpoint to the center of the sphere). Divide each component of this vector by the vector magnitude (square root of the square of the component) to obtain the ideal unit normal vector. The second step is to perform local curve fitting on the generated three-dimensional spatial arc trajectory (actual trajectory). First, take five trajectory sampling points before and after the midpoint. Then, calculate the normal vector of the actual trajectory at the midpoint using the second derivative of the curve. Similarly, divide each component of this vector by the vector magnitude to obtain the actual unit normal vector. Second, calculate the angle between the two unit normal vectors using the three-dimensional vector dot product formula. The angle is equal to the result of the inverse cosine function. The parameters of the inverse cosine function are (X component of the ideal unit normal vector multiplied by X component of the actual unit normal vector, plus Y component of the ideal unit normal vector multiplied by Y component of the actual unit normal vector, plus Z component of the ideal unit normal vector multiplied by Z component of the actual unit normal vector). If the angle is negative, take the absolute value, which reflects the degree of deviation between the actual trajectory and the ideal trajectory.Considering that the rate of change of the orientation angle reflects the curvature of the trajectory, and the normal vector offset reflects the stability of the trajectory, the two are weighted and summed in a 4:6 ratio. That is, the instantaneous curvature characteristic value is equal to the rate of change of the orientation angle multiplied by 0.4, plus the normal vector offset multiplied by 0.6. This value comprehensively reflects the morphological characteristics and stability of the trajectory at the intermediate point.
[0073] Step 602: Input the instantaneous curvature feature value into a preset normalization function. The normalization function is a conversion function that linearly maps the curvature extreme value range to the dynamic correction coefficient range and outputs the dynamic correction coefficient. Specifically, this includes: First, determining the curvature extreme value range by collecting more than 1000 sets of three-dimensional arc trajectory data from different industries (manufacturing, internet, finance, etc.) and different job types (technical, management, functional, etc.). Calculate the instantaneous curvature feature value for each set of data, remove outliers (data exceeding plus or minus 3 standard deviations of the mean of all data), and the minimum value among the remaining data is the minimum curvature value, and the maximum value is the maximum curvature value, forming the curvature extreme value range; Second, determining the dynamic correction coefficient range. Combining the accuracy requirements and error tolerance of talent evaluation, the range is set to 0.8 to 1.2. This range can both correct the influence of amplified core matching features and avoid over-correction leading to evaluation distortion; the normalization function is a conversion function that linearly maps the curvature extreme value range to the dynamic correction coefficient range. The conversion function for the number interval is calculated as follows: Based on linear normalization logic, the calculation method is that the normalized curvature value is equal to (instantaneous curvature characteristic value minus curvature minimum value) divided by (maximum curvature minus curvature minimum value); if the instantaneous curvature characteristic value is less than the curvature minimum value, it is directly taken as 0; if it is greater than the curvature maximum value, it is directly taken as 1, ensuring that the normalized value is strictly in the interval between 0 and 1, eliminating the interference of extreme values; the normalized value is converted into a dynamic correction coefficient through linear mapping, and the calculation method is that the dynamic correction coefficient is equal to the minimum dynamic correction coefficient (0.8) plus the normalized curvature value multiplied by (maximum dynamic correction coefficient (1.2) minus minimum dynamic correction coefficient (0.8)); for example, when the normalized curvature value is 0, the dynamic correction coefficient is 0.8; when the normalized value is 1, it is 1.2; when the normalized value is 0.5, the dynamic correction coefficient is 1.0; the intermediate values are mapped proportionally to achieve a precise correspondence between curvature characteristics and correction coefficients, so that the correction coefficients can dynamically reflect the trajectory morphology characteristics.
[0074] Step 603: Extract the similarity value of each keyword in the initial similarity data, and perform weighted aggregation using the X-axis coordinate value of the corresponding keyword in the 3D point cloud dataset to generate a comprehensive similarity benchmark value; integrate the comprehensive similarity benchmark value with the dynamic correction coefficient to obtain the final talent management score. Specifically, this includes: First, extracting the similarity value of each keyword in the initial similarity data (calculated in step 202, with a value range of 0 to 1, where 1 indicates a complete match with industry standard keywords and 0 indicates a complete mismatch, reflecting the degree of fit between the keyword and industry standard keywords); simultaneously extracting the similarity value from the 3D point cloud dataset... The first step is to centrally obtain the X-axis coordinate value of the corresponding keyword from the cloud dataset. This value is the semantic weight coefficient of the field (ranging from 0 to 1, as determined in step 300; 0 indicates that the corresponding field has no importance to the evaluation, and 1 indicates that the corresponding field is the core of the evaluation; the larger the value, the higher the importance, reflecting the evaluation importance of the field to which the keyword belongs). The second step is to calculate the weighted similarity value of each keyword. The weighted similarity value is equal to the keyword similarity value multiplied by the semantic weight coefficient of the corresponding field. For example, if a keyword has a similarity value of 0.8 and its corresponding field weight coefficient is 0.9, then the weighted similarity value = 0.8 × 0.9 = 0. 0.72, this calculation achieves the quantification of association between matching degree and importance; the third step is to calculate the comprehensive similarity benchmark value. First, the weighted similarity values of all keywords are summed to obtain the weighted similarity sum; then, the semantic weight coefficients of all keywords are summed to obtain the weighted sum (the weighted sum ranges from 0 to the total number of keywords, since the weight of a single field is 0 to 1); if the weighted sum is not 0, the comprehensive similarity benchmark value is equal to the weighted similarity sum divided by the weighted sum (the result is still within the range of 0 to 1); if the weighted sum is 0 (this rarely occurs in actual scenarios, and is mostly due to data anomalies causing all keywords to belong to different categories), ...). If all field weight coefficients are 0, then the arithmetic mean of all keyword similarity values (the sum of similarity values divided by the total number of keywords, with a result between 0 and 1) is taken as the comprehensive similarity benchmark value. This benchmark value comprehensively reflects the correlation effect between the basic matching level of all keywords and the field importance weight (values between 0 and 1). The comprehensive similarity benchmark value (0 to 1) is then fused with the dynamic correction coefficient (value range between 0.8 and 1.2) output in step 602. The fusion method is to directly multiply the two, that is, the final talent management score = comprehensive similarity benchmark value × dynamic correction coefficient (result value range between 0 and 1).2) The core of this fusion logic is that the dynamic correction coefficient is generated based on spatial trajectory features. These features reflect the spatial correlation stability between the core matching points and the comparison points. If the trajectory is stable (small curvature feature value), the correction coefficient approaches 1, and the baseline value remains essentially unchanged. If the trajectory exhibits significant curvature or deviation (large curvature feature value), the correction coefficient is adjusted accordingly (below 1 reduces the influence of the basic matching degree, above 1 amplifies its influence). Ultimately, this achieves dynamic optimization of the talent evaluation score, ensuring that the evaluation results not only conform to the core matching features but also adapt to the stability differences in spatial distribution, thereby improving the accuracy of quantitative evaluation.
[0075] This embodiment, by constructing a three-dimensional arc-shaped trajectory, transforms discrete spatial coordinate data into continuous trajectory features, enabling intuitive capture of the spatial variation patterns of core and contrasting features. The calculation of instantaneous curvature feature values accurately extracts the morphological change features at key positions of the trajectory, reflecting the abrupt changes or stable states of talent-related features. The normalized dynamic correction coefficient achieves a reasonable conversion of curvature features into evaluation coefficients, allowing the evaluation results to dynamically adapt to feature changes. The weighted aggregation of comprehensive similarity benchmark values fully considers the semantic importance weights of different fields. The final talent management score obtained after integrating the dynamic correction coefficients integrates multi-dimensional information such as basic matching degree and spatial feature changes, effectively improving the accuracy, dynamic adaptability, and scientific nature of talent evaluation, and reducing the bias caused by single-dimensional evaluation.
[0076] like Figure 2 As shown, embodiments of the present invention also provide an artificial intelligence-based talent evaluation and management system, comprising:
[0077] The processing module is used to acquire the resume data of the candidates to be evaluated, perform text recognition processing on the resume data, and generate a structured text dataset.
[0078] The calculation module is used to extract keyword feature vectors from the structured text dataset, calculate the similarity value between the keyword feature vectors and the preset talent evaluation keyword library, and generate initial similarity data.
[0079] The building module is used to construct a 3D point cloud dataset based on the initial similarity data, where each data point contains the weight coordinate information of the keyword;
[0080] The determination module is used to determine the reference coordinate point based on the spatial density distribution of the point cloud dataset, generate two spatial vector data based on the reference coordinate point, and calculate the boundary data of the included angle between the two spatial vector data.
[0081] The selection module is used to define the range of a sector area using the included angle boundary data, select the first spatial coordinate data within the sector area, and symmetrically select the second and third spatial coordinate data outside the range.
[0082] The parsing module is used to generate arc-shaped trajectory data based on the time series relationship of the first, second, and third spatial coordinate data, parse the curvature feature value of the trajectory data to generate dynamic correction coefficients, and finally fuse the initial similarity data and dynamic correction coefficients to output the final talent management score.
[0083] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0084] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0085] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0086] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A talent evaluation management method based on artificial intelligence, characterized by, The method comprises: Step 1, obtaining resume data of the talent to be evaluated, and performing text recognition processing on the resume data to generate a structured text data set; Step 2, extracting a keyword feature vector from the structured text data set, calculating the similarity value of the keyword feature vector and the preset talent evaluation keyword library, and generating initial similarity data; Step 3, constructing a three-dimensional point cloud data set based on the initial similarity data, wherein each data point contains weight coordinate information of the keyword; Step 4, determining a reference coordinate point according to the spatial density distribution of the point cloud data set, and generating two spatial vector data based on the reference coordinate point, and calculating the included angle boundary data between the spatial vector data; Step 5, defining a sector area range using the included angle boundary data, selecting a first spatial coordinate data within the sector area range, and selecting a second spatial coordinate data and a third spatial coordinate data outside the range, comprising: defining a sector space area in the XY plane with the reference coordinate point as the vertex and the included angle boundary data as the expansion angle; selecting a data point within the sector space area as the first spatial coordinate data, and selecting a data point outside the area as the second spatial coordinate data and the third spatial coordinate data with the center bisector of the sector space area as the symmetry axis; Step 6, generating arc trajectory data according to the selection order of the first, second and third spatial coordinate data, analyzing the curvature characteristic value of the arc trajectory data, and generating a dynamic correction coefficient; fusing the initial similarity data and the dynamic correction coefficient to output the final talent management score, comprising: connecting the first spatial coordinate data as the starting point, the second spatial coordinate data as the intermediate point, and the third spatial coordinate data as the termination point to generate a three-dimensional space arc trajectory according to the selection order; calculating the instantaneous curvature characteristic value of the three-dimensional space arc trajectory at the second spatial coordinate data position, wherein the instantaneous curvature characteristic value is determined by the change rate of the direction angle of the trajectory tangent at the second spatial coordinate data position and the offset of the trajectory normal vector; inputting the instantaneous curvature characteristic value into a preset normalization function, wherein the normalization function is a conversion function that linearly maps the curvature extreme range to the dynamic correction coefficient interval, and outputting the dynamic correction coefficient; extracting the similarity values of each keyword in the initial similarity data, and weighting and aggregating the X-axis coordinate values of the corresponding keywords in the three-dimensional point cloud data set to generate a comprehensive similarity reference value; fusing the comprehensive similarity reference value and the dynamic correction coefficient to obtain the final talent management score. 2.The artificial intelligence-based talent evaluation management method of claim 1, wherein The step 1 comprises: performing paragraph semantic segmentation on the resume data to identify a first text block containing education experience, a second text block containing work experience, and a third text block containing skill qualification; mapping the first text block to a preset education background structured field, mapping the second text block to a work experience structured field, and mapping the third text block to a professional skill structured field to generate a structured text data set composed of structured field labels and bound text content, wherein each structured field label contains a field type identifier. 3.The AI-based talent evaluation management method of claim 2, wherein, The step 2 comprises: Parsing the text content of the field binding to extract core keywords for each structured field label in the structured text dataset; Generating a numerical feature vector according to the term frequency weight and position weight of each core keyword in the text content; Matching the feature vector with a standard feature vector under the same field type identifier in the preset talent evaluation keyword library to output a similarity value of the corresponding keyword; Aggregating the keywords and corresponding similarity values under all structured field labels to generate initial similarity data containing the field type identifier, keyword, and similarity value. 4.The AI-based talent evaluation management method of claim 3, wherein, The step 3 includes: For each keyword in the initial similarity data, using the preset field semantic weight coefficient as the X-axis coordinate value according to the field type identifier, and reading the similarity value of the corresponding keyword as the Y-axis coordinate value; According to the original occurrence frequency of the corresponding keyword in the structured text dataset corresponding field text content, calculating the normalized frequency value as the Z-axis coordinate value; Defining the X, Y, and Z coordinate values of each keyword as a three-dimensional space point, and integrating the space point coordinates of all keywords to form a three-dimensional point cloud dataset. 5.The artificial intelligence-based talent evaluation management method of claim 4, wherein The step 4 includes: Dividing the three-dimensional point cloud dataset into a spatial grid, counting the number of data points in each grid cell, marking the grid cells with a data point number greater than a preset threshold as dense cells, and extracting the data points contained in all dense cells to form a candidate point set; Constructing a triangular mesh structure in the candidate point set, calculating the area and edge length ratio coefficient of each triangular cell, selecting all triangular cells with edge length ratio coefficients meeting the preset conditions, and identifying the target triangle with the longest side in all triangular cells; Calculating the distance weight of the three vertices of the target triangle with the longest side to the centroid of the point cloud dataset, weighting and fusing the coordinates of the three vertices based on the distance weight to obtain fused coordinates, and determining the fused coordinates as the reference coordinate point; Generating a first spatial vector from the reference coordinate point to the positive direction of the X-axis of the point cloud dataset, and generating a second spatial vector to the positive direction of the Y-axis; calculating the three-dimensional space included angle value between the first spatial vector and the second spatial vector as the included angle boundary data.
6. An artificial intelligence-based talent evaluation management system, the system implementing the method of any one of claims 1 to 5, characterized by, It includes: The processing module is used for acquiring the resume data of the talent to be evaluated, performing text recognition processing on the resume data, and generating a structured text dataset; The calculation module is used for extracting a keyword feature vector from the structured text dataset, calculating the similarity value of the keyword feature vector with a preset talent evaluation keyword library, and generating initial similarity data; The construction module is used for constructing a three-dimensional point cloud dataset based on the initial similarity data, wherein each data point contains weight coordinate information of a keyword; The determination module is used for determining a reference coordinate point according to the spatial density distribution of the point cloud dataset, generating two spatial vector data based on the reference coordinate point, and calculating the included angle boundary data between the two spatial vector data; The selecting module is configured to define a sector region range by using the included angle boundary data, select first spatial coordinate data in the sector region range, and symmetrically select second spatial coordinate data and third spatial coordinate data outside the range, and includes: defining a sector space region in the XY plane with a reference coordinate point as a vertex and the included angle boundary data as an opening angle; selecting a data point in the sector space region as the first spatial coordinate data, and selecting a data point at a symmetric position outside the region as the second spatial coordinate data and the third spatial coordinate data respectively with a bisector of the sector space region as a symmetric axis; The analyzing module is configured to generate arc trajectory data according to a time sequence relationship of the first, second, and third spatial coordinate data, analyze a curvature characteristic value of the trajectory data to generate a dynamic correction coefficient, and finally fuse the initial similarity data and the dynamic correction coefficient to output a final talent management score, and includes: connecting the first spatial coordinate data as a starting point, the second spatial coordinate data as an intermediate point, and the third spatial coordinate data as a terminal point in the selected order to generate a three-dimensional space arc trajectory; calculating an instantaneous curvature characteristic value of the three-dimensional space arc trajectory at the position of the second spatial coordinate data, wherein the instantaneous curvature characteristic value is determined by a change rate of a direction angle of a tangent line of the trajectory at the position of the second spatial coordinate data and an offset of a normal vector of the trajectory; inputting the instantaneous curvature characteristic value into a preset normalization function, wherein the normalization function is a conversion function based on linear mapping of a curvature extreme value range to a dynamic correction coefficient interval, and outputting the dynamic correction coefficient; extracting similarity values of each keyword in the initial similarity data, and performing weighted aggregation on the X-axis coordinate values of the corresponding keywords in the three-dimensional point cloud data set to generate a comprehensive similarity reference value; and fusing the comprehensive similarity reference value and the dynamic correction coefficient to obtain the final talent management score.
7. A computing device, comprising: The method comprises: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a program, and the program is executed by the processor to implement the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Intelligent talent information analysis method based on deep learning
CN120069823A
Talent evaluation management method and system based on AI intelligence
CN120338741A