Intelligent talent tag portrait analysis system based on big data

By designing a talent tag portrait intelligent analysis system based on big data, the problems of inaccurate tag extraction, rigid portrait structure and neglected semantic fusion in the existing technology are solved, and the accuracy of high-accurate tag normalization and human positions are improved.

CN120087927AActive Publication Date: 2025-06-03LUOKE (XIAMEN) NETWORK TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510538338.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-06-03
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

When constructing and analyzing talent tag portraits, there are problems such as inaccurate label extraction, rigid portrait structure, and neglecting synonyms and semantic fusion in multi-source corpus, resulting in redundancy in portrait dimension construction.

Method used

An intelligent analysis system for talent tag portraits based on big data is designed. The vector module is used to construct candidate tag association vectors, and combined with the dynamic real-time update of tag semantics, the graph construction module analyzes the ability scores of candidate talents and builds a tag map. The matching module optimizes the portrait through cosine similarity calculation and label offset analysis. It is recommended that the module generates a job adaptation sorting table.

Benefits of technology

It improves the accuracy of label normalization and portrait structure consistency, improves the accuracy of person-position matching, enhances the ability of system self-learning and label compensation, and improves the personalization and effectiveness of job recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087927A_ABST
    Figure CN120087927A_ABST
Patent Text Reader

Abstract

The invention discloses a talent tag portrait intelligent analysis system based on big data, relates to the technical field of intelligent analysis, realizes unified access and processing of multi-source heterogeneous talent data through a vector module, and effectively solves the problems of tag expression inconsistency and semantic drift in combination with BERT embedding and tag semantic evolution mechanisms. Therefore, the accuracy of label normalization and portrait structure consistency is improved. According to the system, the label structure relation and the capability score are linked and fused through a graph construction module, a structurable original portrait vector is generated, a semantic collaboration network between labels is constructed, and deep semantic support is provided for portrait calculation. The matching module improves the man-post matching precision through vector alignment and similarity calculation of post portraits and candidate portraits, triggers a label offset analysis and reconstruction scoring mechanism when the matching value is insufficient, realizes dynamic optimization of the portraits, and enhances the self-learning and label compensation capabilities of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent analysis technology, and particularly relates to an intelligent analysis system for talent label portraits based on big data. Background Art

[0002] Currently, with the construction of information-based and digital recruitment management platforms, talent feature modeling, skill assessment, and job portrait matching have become key requirements for recruitment intelligence. In this context, talent portrait technology has gradually evolved into a key implementation method of tagging, structuring, and semanticization. For the construction and analysis of talent label portraits in the above fields, the current mainstream methods mainly rely on rule-based label extraction or shallow analysis methods based on keyword extraction, which may lead to inaccuracies in talent portraits, ultimately affecting the evolution process of the map structure and resulting in a rigid portrait structure. Moreover, for synonymous words in multi-source corpora, such semantic labels, existing systems ignore dynamic recognition and semantic fusion, resulting in redundancy in the construction of portrait dimensions. Summary of the Invention

[0003] In view of the above problems existing in the prior art, the present application provides an intelligent analysis system for talent label portraits based on big data.

[0004] An embodiment of the present disclosure provides an intelligent analysis system for talent label portraits based on big data, including:

[0005] A vector module, which obtains talent data of multiple candidate talents through multiple sources, performs dimensionality reduction processing, constructs a candidate label association vector and, in combination with the dynamics of label semantics, updates the candidate label association vector in real time ;

[0006] A map construction module, which will, according to the updated candidate label association vector analyze the ability scores of each candidate talent in the corresponding label dimension to generate an original portrait vector of each candidate talent and construct a label map structure body;

[0007] A matching module, which, by pre-obtaining a portrait vector of a job to be matched calculates the cosine similarity with the original portrait vectors of each candidate talent to obtain a matching value SIM. If the matching value SIM does not exceed a preset threshold, a label offset analysis instruction is executed, and unqualified labels of the corresponding candidate talent are extracted and optimized to generate a new portrait vector of each candidate talent ;

[0008] A suggestion module, based on the new portrait vectors of each candidate talent , construct a corresponding candidate talent portrait adaptation ranking table for all positions to be matched to achieve the job of position recommendation and prompt.

[0009] Optionally, the vector module includes a vector construction unit and a semantic update unit;

[0010] The vector construction unit uses big data technology to collect the talent data of each candidate from multiple sources. The multiple sources include but are not limited to resume systems, education platforms, recruitment platforms, project records, behavior logs, and social media content. It performs text parsing, field regularization, missing value filling, and noise elimination on the talent data, uses the BERT embedding model to extract label words from the talent data, obtains the word vectors corresponding to each label word, and matches the word vectors corresponding to each label word with the standard label vectors of the positions in the position word library through BERT-CLS, so as to map the word vectors corresponding to each label word to the standard labels of the corresponding positions in the position word library respectively, and perform dimensionality reduction processing on the word vectors corresponding to each label word to construct candidate label association vectors .

[0011] Optionally, the semantic update unit updates the standard labels of the positions in the position word library in real time based on the dynamics of label semantics. Specifically: statistically analyze the semantic change trends of each standard label of the positions in the position word library in the past period to construct a label semantic drift entropy coefficient Ht. Based on the numerical size of the label semantic drift entropy coefficient Ht, determine the semantic stability of the corresponding label in each time period. If it is determined that the semantics of the corresponding label is unstable in each time period, mark the corresponding label as a pseudo-label, and determine the set of labels to be replaced. By analyzing the label direction offset between the pseudo-label and each standard label of the positions in the position word library, obtain the label direction offset value Dj to complete the judgment of the label transformation ability, where the set of labels to be replaced refers to the standard label vectors of the positions in the position word library except for the pseudo-labels; specifically: ;

[0012] Among them, is the word vector of the g-th pseudo-label, are the standard label vectors of the positions in the set of labels to be replaced; is the cosine similarity of the angle between the pseudo-label and the standard label vectors of the positions in the set of labels to be replaced in the vector space;

[0013] Extract the standard label of the position corresponding to the smallest label direction offset value Dj, and fuse it with the corresponding pseudo-label to form a parallel extended label, and return the parallel extended label to the vector construction unit to re-match with the word vectors corresponding to each label word in the talent data through BERT-CLS to update the candidate label association vectors .

[0014] Optionally, the graph construction module includes a feature unit, a scoring unit, and a construction unit;

[0015] The feature unit, based on the updated candidate label association vector and talent data, identifies the feature items of all candidate talents, analyzes the response intensity of each feature item under different label dimensions, and projects each feature item under all label dimensions using semantic alignment to obtain the relationship matrix V.

[0016] Optionally, the scoring unit analyzes the ability scores of each candidate talent under the corresponding label dimension according to the relationship matrix V and using the TF-IDF algorithm to obtain the total ability score tss of each candidate talent under the corresponding label dimension. Specifically:

[0017] where, is the total ability score of the corresponding candidate talent under the i-th label dimension, is the response value of the r-th feature item of the corresponding candidate talent under the i-th label dimension, is the importance weight of the r-th feature item, r is the feature item number, and m is the total number of feature items;

[0018] The total ability scores tss of each candidate talent under all label dimensions are statistically analyzed to generate the original portrait vectors of each candidate talent .

[0019] Optionally, co-occurrence statistics are performed on the updated candidate label association vector output by the semantic update unit to obtain a co-occurrence matrix. Each element in the co-occurrence matrix represents the co-occurrence times Yc of two labels among all candidates. Based on the co-occurrence matrix, the association strength factor RAF between labels is calculated;

[0020] The construction unit uses the labels as nodes and the association strength factor RAF as the edges between nodes, and the original portrait vectors of each candidate talent as the numerical representation form of the label graph structure to draw the label graph structure.

[0021] Optionally, the matching module includes a matching unit and an optimization unit;

[0022] The matching unit pre-obtains a set of portrait vectors of the positions to be matched and combines them with the original portrait vectors of each candidate talent to obtain a matching value SIM. The matching value SIM is specifically obtained according to the following formula: ;

[0023] where, SIM is the matching value of the a-th candidate talent relative to the portrait vector of the b-th position to be matched, where a is the candidate talent number and b is the position number to be matched. is the original portrait vector of the a-th candidate talent. is the portrait vector of the b-th position to be matched.

[0024] Optionally, the optimization unit compares the matching value SIM with a preset threshold. If the matching value SIM exceeds the preset threshold, it indicates that the current candidate talent's matching with the corresponding portrait vector of the position to be matched is in a qualified state. At this time, the current candidate is included in the candidate talent portrait adaptation ranking table corresponding to the position to be matched that matches it; if the matching value SIM does not exceed the preset threshold, the label offset analysis instruction is executed.

[0025] Optionally, receive the label offset analysis instruction, analyze the deviation degree of each label dimension of the corresponding candidate talent on the corresponding position to be matched to obtain the difference , if the difference >0 for the label dimension, then execute the reconstruction score update: ;

[0026] where is the total ability score adjusted for the i-th label dimension in the original portrait vector of the a-th candidate talent relative to the b-th position to be matched, is the optimization fusion coefficient;

[0027] And after executing the reconstruction score update, regenerate the original portrait vector of the corresponding candidate talent , and mark it as the new portrait vector of the corresponding candidate talent .

[0028] Optionally, the recommendation module returns the new portrait vector of the corresponding candidate talent to the matching unit to regenerate the candidate talent portrait adaptation ranking table, and extract the candidate talent with the largest SIM value in the regenerated candidate talent portrait adaptation ranking table, and use this candidate talent as the recommended candidate for the corresponding position to be matched.

[0029] The beneficial effects of the present invention are as follows:

[0030] The vector module realizes unified access and processing of multi-source heterogeneous talent data. Combined with BERT embedding and label semantic evolution mechanism, it effectively solves the problems of inconsistent label expression and semantic drift, thereby improving the accuracy of label normalization and portrait structure consistency. The system integrates the label structure relationship with the ability score through the graph construction module to generate a structured original portrait vector, and builds a semantic collaborative network between labels to provide deep semantic support for portrait calculation. The matching module improves the accuracy of person-job matching through vector alignment and similarity calculation of job portraits and candidate portraits, and triggers label offset analysis and reconstruction scoring mechanism when the matching value is insufficient, realizes dynamic optimization of portraits, and enhances the system's self-learning and label compensation capabilities. The recommendation module generates job adaptation ranking results through new portrait vectors to improve the personalization and effectiveness of job recommendations. The overall system has the technical advantages of comprehensive portrait semantic expression, timely dynamic label update, high person-job matching accuracy and strong adaptability of job recommendations, which further improves the efficiency of talent selection and the intelligence level of organizational recruitment decisions.

[0031] Through B, the system normalizes labels with the same semantics but different expressions from different sources into the standard label library for positions, which enhances the comparability of the portraits and the basis for position adaptation. The semantic update unit introduces a dynamic calculation mechanism for the label semantic drift entropy coefficient Ht, which can accurately monitor the changing trend of label semantics in different periods. When the label semantics is unstable, it is automatically identified as a pseudo-label, and label replacement recommendations and fusion generation are performed based on the directional offset value Dj to ensure the continuous consistency of the label system and the semantic structure of the position. Overall, this module improves the accuracy and timeliness of label semantic recognition, making the portrait construction more dynamically adaptable, and providing a label foundation with stable semantics, clear dimensions, and strong structural evolution capabilities for the subsequent establishment of label graph structure and position matching.

[0032] By performing full co-occurrence statistics on the updated candidate tag association vectors, constructing a co-occurrence matrix and calculating the association intensity factor RAF between tags based on it, it is possible to quantify the combination frequency and synergy of different tags in talent portraits, accurately reflect the semantic coupling relationship between tags, and overcome the problem of unclear label edge weights or static structure in traditional graph construction. Secondly, the construction unit uses labels as nodes and RAF as edges to achieve dynamic construction and continuous evolution of the label graph structure, so that the system can adjust the graph structure in real time according to changes in label semantic relationships, improving the true reflection of label relationship modeling and system interpretability.

[0033] In traditional job matching systems, when candidates do not fully match the job profile, they are often directly eliminated because the matching value does not meet the standard. This ignores the situation that some candidates have individual ability dimension deviations but their overall ability is close to the job requirements, causing excellent talents to be misjudged as unsuitable, thus affecting the matching accuracy and talent utilization efficiency. To solve this problem, the system introduces a label offset analysis and reconstruction scoring mechanism through the optimization unit. When the matching value SIM does not reach the threshold, the label dimension in the portrait with a large deviation from the job portrait is identified, and its score value is weighted and adjusted to construct a new portrait vector. Subsequently, the recommendation module regenerates the ranking table based on the optimized portrait results to ensure that potential talents that are relatively suitable for the position are discovered within the acceptable adjustment range. This mechanism effectively improves the system's ability to identify and recommend suboptimal but malleable candidates, enhances the intelligent controllability of the portrait system and the robustness of job adaptation, and realizes the upgrade from static matching to dynamic optimization matching, ultimately improving the accuracy of job recommendations and the efficiency of talent utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the present application or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application.

[0035] Figure 1 It is a system block diagram of the present invention. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solution and advantages of this application clearer, the technical solution of this application will be described clearly and completely in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, but not all of them. Based on the embodiments in this application, ordinary technicians in this field can make the following inventions without creative work.

[0037] All other embodiments obtained under the premise of the above-mentioned embodiment belong to the scope of protection of this application.

[0038] Example 1

[0039] like Figure 1 As shown, the present invention provides a talent label portrait intelligent analysis system based on big data, including:

[0040] The vector module obtains the talent data of multiple candidates based on multiple sources, and constructs the candidate label association vector after dimensionality reduction processing. , and combined with the dynamic nature of label semantics, update the candidate label association vector in real time ;

[0041] The graph construction module associates the updated candidate label vectors , analyze the ability scores of each candidate talent under the corresponding label dimensions to generate the original portrait vectors of each candidate talent , and construct a label graph structure

[0042] Matching module, by pre-obtaining the portrait vectors of the positions to be matched , compare them with the original portrait vectors of each candidate talent to calculate the cosine similarity to obtain the matching value SIM. If the matching value SIM does not exceed the preset threshold, execute the label offset analysis instruction, extract the unqualified labels of the corresponding candidate talent and optimize them to generate the new portrait vectors of each candidate talent ;

[0043] Recommendation module, based on the new portrait vectors of each candidate talent , construct a corresponding candidate talent portrait adaptation ranking table for all positions to be matched to implement the job recommendation prompt operation.

[0044] In the embodiments of the present invention, the multi-source candidate data is uniformly modeled through the vector module, and the reduced-dimensional candidate label correlation vectors are constructed, effectively solving the problems of multi-channel data heterogeneity, label expression differences, and semantic redundancy;

[0045] Introduce a label semantic dynamic update mechanism, which can identify and correct the label drift phenomenon, ensuring the timeliness and industry adaptability of the portrait dimensions;

[0046] Through the graph construction module, the ability labels of the candidates are structurally modeled to form a label graph and the original portrait vectors, providing a structural basis for subsequent position matching; the matching module realizes the high-precision matching of candidates and position portraits through cosine similarity, and at the same time introduces a portrait optimization mechanism, which can automatically identify and optimize the unqualified label dimensions when the SIM value is lower than the threshold, improving the adaptability between the portrait and the position; the recommendation module can generate a position adaptation ranking table based on the new portrait vectors, realizing the dynamic matching recommendation of candidates and positions, and further improving the accuracy and intelligent level of position recommendation.

[0047] Embodiment 2

[0048] As Figure 1 shown, specifically: the vector module includes a vector construction unit and a semantic update unit;

[0049] Vector construction unit, using big data technology, collects talent data of each candidate from multiple sources. The multiple sources include but are not limited to resume systems, education platforms, recruitment platforms, project records, behavior logs, and social media content. Then, it performs text parsing, field regularization, missing value filling, and noise removal on the talent data. It uses the BERT embedding model to extract label words from the talent data and obtains the word vectors corresponding to each label word. The word vectors corresponding to each label word are matched with the job standard label vectors in the job word library using BERT-CLS (equivalent to identifying and merging field items with the same semantics but different representations from different sources), so as to map the word vectors corresponding to each label word to the corresponding job standard labels in the job word library respectively, and perform dimensionality reduction processing on the word vectors corresponding to each label word to construct candidate label association vectors 。

[0050] The above BERT embedding model is a bidirectional Transformer structure language model proposed by Google, which can model the semantics of words or sentences from the context and is used to transform label words or candidate phrases into high-dimensional semantic vector representations;

[0051] Using BERT-CLS matching is one of the semantic matching methods;

[0052] The vector module is used to solve the problems of chaotic data sources and invisible talent characteristics, turning text information, behavior data, etc. into digital forms (vectors), and this vectorization process can also discover hidden capabilities. For example, a person does not explicitly write "project management", but the behavior data and language structure show that he has done it.

[0053] Specifically, a label word refers to a keyword used to express a certain ability, job characteristic, technical direction, or behavior attribute, and it constitutes the semantic unit for scoring each dimension in the talent portrait.

[0054] A word vector is to transform a natural language unit such as a label word into a computable dense vector representation through mathematical modeling, so as to perform computational operations such as comparison, clustering, and matching in the model.

[0055] Semantic update unit, based on the dynamics of label semantics, updates the job standard labels in the job word library in real time. Specifically: it statistically analyzes the semantic change trends of each job standard label in the job word library over a past period to construct a label semantic drift entropy coefficient Ht. The label semantic drift entropy coefficient Ht is obtained through the following formula: ;

[0056] Among them, is the label semantic drift entropy coefficient of the i-th label, is the proportion of label semantic changes in the kth period (i.e., the proportion of label semantic changes in this period), is a small positive number to prevent log(0), usually ; n is the total duration of the past period, k is the number of each period in the past period, and i is the number of the label in the candidate label association vector; is the entropy calculation term with base 2, indicating the degree of uncertainty of semantic change;

[0057] The tag semantic drift entropy coefficient Ht is used to measure the fluctuation degree of a tag semantics over a period of time. This formula is essentially an information entropy formula, which reflects the "uncertainty of the tag semantic change trend";

[0058] Specifically, if the tag semantic drift entropy coefficient Ht is large, it means that the tag semantics changes frequently in the time dimension, the direction is scattered, and the trend is unstable; if the tag semantic drift entropy coefficient Ht is close to 0, it means that the tag semantics remains basically unchanged in each time period and the semantics tends to be stable.

[0059] By collecting statistics on the semantic change trends of standard labels for various positions in the past period, we can automatically identify the semantic change trends of labels and improve the long-term accuracy of the system. Label updates can be propagated in the graph to ensure the overall consistency of the system, and the label evolution path and version management records are retained to improve the interpretability of the system.

[0060] Based on the value of the label semantic drift entropy coefficient Ht, the semantic stability of the corresponding label in each time period is determined. If the semantics of the corresponding label is determined to be unstable in each time period, the corresponding label is marked as a pseudo-label, and the label set to be replaced is determined. By analyzing the label direction offset between the pseudo-label and the standard labels of each position in the position vocabulary, the label direction offset value Dj is obtained to complete the judgment of the label transformation ability, where the label set to be replaced refers to the standard label vector of the position in the position vocabulary except the pseudo-label; specifically: ;

[0061] in, is the word vector of the g-th pseudo-label, indicating the semantic difference, is the standard label vector of each position in the label set to be replaced; is the cosine similarity of the angle between the pseudo label and the standard label vector space of each position in the label set to be replaced. The larger the value, the closer it is (the maximum is 1). It means converting similarity into difference, so that the smaller the value, the closer the direction; g is the number of the pseudo label;

[0062] This formula is used to measure "which candidate label is closest to the direction of the original label" in the semantic vector space; essentially, it is a "substitution ability judgment" of label semantics; the smaller the label direction offset value Dj, the more suitable the corresponding job standard label in the label set to be replaced is as the new semantic representative of the current pseudo-label.

[0063] Extract the job standard label corresponding to the smallest label direction offset value Dj and fuse it with the corresponding pseudo-label to form a parallel extended label, thereby realizing the generation of a new label;

[0064] It should be noted that the original pseudo-label and the original job standard label to be fused are retained, and only a new label is generated on the basis of the fusion;

[0065] For the dynamic nature of label semantics, taking "Full Stack Engineer" as an example: in the early stage, this label may represent the ability to be competent in both front-end and back-end programming work; but in the later stage, it may be necessary to master new capabilities such as cloud native technology, automated deployment, container orchestration, and security architecture at the same time; if the definition of this label in the original portrait system fails to be updated accordingly, it will cause label mapping distortion or ability evaluation lag. Therefore, considering the impact of the dynamic nature of label semantics on later labels, this part avoids the rigidity of the portrait structure, so as to better adapt to the actual needs such as job updates, label expansion, and talent label migration;

[0066] Return the parallel extended label to the vector construction unit and use BERT-CLS to match it again with the word vectors corresponding to each label word in the talent data to update the candidate label association vector 。

[0067] In the embodiment of the present invention, first, the vector construction unit can extract candidate data from multiple sources and use the BERT embedding model to obtain the word vectors of label words, and automatically match and normalize labels with the same semantics in different expression forms (such as "Full Stack Engineer" and "Fullstack Developer") in the semantic space, avoiding label redundancy and semantic distortion problems.

[0068] Secondly, the semantic update unit introduces the label semantic drift entropy coefficient Ht to dynamically monitor the semantic change trend of labels in the job vocabulary at different time periods. It can automatically identify pseudo-labels and recommend relatively better alternative expressions when semantic drift occurs. For example, if the label "AI architect" tends to be "AI platform architecture design" under the background of new technology development, the system will automatically analyze its direction offset value Dj and fuse it to form a new parallel extended label "AI architect - platform direction", thus realizing the evolutionary update of the label map. Through the above mechanism, the system can continuously optimize the candidate label association vector, improve the accuracy, update ability and system adaptability of the talent label portrait, and provide a solid foundation for job matching and recommendation.

[0069] Embodiment 3

[0070] As Figure 1 shown, specifically: The map construction module includes a feature unit, a scoring unit and a construction unit;

[0071] The feature unit, according to the updated candidate label association vector and talent data, identifies the feature items of all candidate talents, analyzes the response intensity of each feature item under different label dimensions, and projects each feature item under all label dimensions using the semantic alignment method to obtain the relationship matrix V. The expression of the relationship matrix V is: ;

[0072] where M is the number of label dimensions and K is the number of feature items.

[0073] The semantic alignment method is used for multi-label projection and label-feature mapping. Semantic alignment means projecting input items such as feature items / keywords / label words into the multi-label semantic space so that they have response representations under multiple label dimensions;

[0074] The scoring unit, according to the relationship matrix V and using the TF-IDF algorithm, analyzes the ability scores of each candidate talent under the corresponding label dimension to obtain the total ability score tss of each candidate talent under the corresponding label dimension. Specifically: ;

[0075] where is the total ability score of the corresponding candidate talent under the i-th label dimension, is the response value of the r-th feature item of the corresponding candidate talent under the i-th label dimension, that is, the supportiveness of this feature item to this label, is the importance weight of the r-th feature item, r is the feature item number, and m is the total number of feature items;

[0076] The total ability score tss is used to form the "tag strength vector" in the subsequent talent portrait. It is the core numerical expression of the tag portrait construction. The expression strengths of multiple feature dimensions on the current tag are weighted and summed, and the result is the score of the candidate under this tag semantics. The higher the value, the stronger the ability score.

[0077] Statistically analyze the total ability scores tss of each candidate talent under all tag dimensions to generate the original portrait vector of each candidate talent. ;

[0078] Among them, the importance weight of the r-th feature item is obtained through the following formula: ;

[0079] Among them, is the frequency of occurrence of the r-th feature item in the talent data of the corresponding candidate talent, is the total number of candidate talents, is the number of talent data containing the r-th feature item, that is, the document frequency. is used to suppress the weight of high-frequency tags so that common words do not interfere with the true discrimination.

[0080] In the embodiments of the present invention, the graph construction module in the present invention realizes the structured, semantic and quantifiable expression of talent portrait construction through the collaborative design of the feature unit, the scoring unit and the construction unit, and has significant technical beneficial effects. First, the feature unit jointly analyzes the candidate tag association vector and the talent data, automatically identifies the feature items of the candidate, and projects them in different tag dimensions through the semantic alignment method to construct a response matrix V reflecting the tag-feature relationship, enhancing the semantic relevance of feature expression.

[0081] Secondly, the scoring unit assigns importance weights to each feature item using the TF-IDF method, and combines its response values in different tag dimensions to calculate the ability score TSS in each tag dimension, thereby generating a complete original portrait vector, ensuring the quantifiable and differential expression of the portrait. This design effectively improves the expression accuracy and semantic integrity of the talent portrait for the job tags, providing accurate and structured portrait input for subsequent graph construction and job matching.

[0082] Embodiment 4

[0083] As Figure 1 shown, specifically: perform co-occurrence statistics on the updated candidate tag association vector output in the semantic update unit to obtain a co-occurrence matrix. Each element in the co-occurrence matrix represents the co-occurrence times Yc of two tags among all candidates. Based on the co-occurrence matrix, calculate the association strength factor RAF between tags, specifically: ;

[0084] Among them, is the co-occurrence times of the i-th label and the j-th label among all candidates, is the correlation strength factor between the i-th label and the j-th label, is the independent occurrence frequency of the i-th label, is the independent occurrence frequency of the j-th label, where both i and j are the numbers of labels within the candidate label association vector;

[0085] Since the label itself does not exist in isolation, its semantic utility must be understood in combination with the "structure". Only through individual scoring and the structure of the group can the true semantic combination effect of the label be restored. Therefore, when constructing the label graph structure body, the correlation strength factor RAF between labels needs to be considered;

[0086] The correlation strength factor RAF between labels reflects the correlation strength between labels. The higher the value, the more frequently they appear in combination; it is the basis for the edge weight of the label graph structure body. The more relevant the label edges, the thicker they are, and the system will consider that the labels may combine to form a composite ability, reflecting the semantic cooperation and collocation frequency between labels;

[0087] Construction unit, taking the label as a node and the correlation strength factor RAF as the edge between nodes, and the original portrait vectors of each candidate talent as the numerical representation form of the label graph structure body to draw the label graph structure body, where the original portrait vectors

[0088] The matching module includes a matching unit and an optimization unit;

[0089] Matching unit, pre-obtaining the set of portrait vectors of the positions to be matched, and combining with the original portrait vectors of each candidate talent ;

[0090] Among them, is the matching value of the a-th candidate talent relative to the b-th portrait vector of the position to be matched, a is the number of the candidate talent, b is the number of the position to be matched, is the original portrait vector of the a-th candidate talent, is the b-th portrait vector of the position to be matched; is the modulus of the original portrait vector of the a-th candidate talent, is the norm of the b-th candidate position portrait vector.

[0091] In the embodiments of the present invention, global co-occurrence statistics are performed on the candidate label association vectors updated by semantics, a co-occurrence matrix is constructed, and the association strength factor RAF between labels is calculated accordingly, realizing the quantitative expression of the semantic collaboration relationship between labels. The higher the RAF value, the more frequently the labels co-occur, effectively revealing the combination preferences between labels and the semantic matching patterns of positions. The construction unit takes labels as nodes, RAF as edge weights, and portrait vectors as numerical expressions of labels to draw a label map structure, thus providing support for the structured relationship between labels.

[0092] Secondly, the matching module measures the degree of ability adaptation by calculating the cosine similarity SIM between the candidate and the position's original portrait vector. The matching result is numericalized and standardized to support ranking and screening. This mechanism not only realizes the quantitative alignment of multi-dimensional semantics but also provides a structural and interpretable basis for position recommendation and person-position matching, improving the transparency and credibility of the intelligent recommendation system.

[0093] Embodiment 5

[0094] Such as Figure 1 shown, specifically: The optimization unit compares the matching value SIM with a preset threshold. If the matching value SIM exceeds the preset threshold, it indicates that the current candidate talent's matching with the corresponding candidate position portrait vector is in a qualified state. At this time, the current candidate is included in the candidate talent portrait adaptation ranking table corresponding to the candidate position that matches it; if the matching value SIM does not exceed the preset threshold, the label deviation analysis instruction is executed.

[0095] Since the portrait itself may have insufficient expression, the label ability score of the candidate is not absolutely accurate. There may be information missing in the feature modeling stage, insufficient semantic coverage in the label projection stage, or ambiguity / redundancy in the candidate's description, which will result in incomplete coverage or incomplete expression of the candidate's ability portrait for the position portrait. Therefore, when the matching value SIM does not exceed the preset threshold, the label deviation analysis instruction is considered to be executed;

[0096] Receiving the label deviation analysis instruction, analyzing the deviation degree of each label dimension of the corresponding candidate talent on the corresponding candidate position to obtain the difference : ; is the total ability score of the i-th label dimension in the original portrait vector of the a-th candidate talent, is the total ability score of the i-th label dimension in the b-th candidate position portrait vector, is the difference of the i-th label dimension in the original portrait vector of the a-th candidate talent for the b-th position to be matched; if the difference > 0 for the label dimension, then perform the reconstruction score update: ;

[0097] where is the total ability score adjusted for the i-th label dimension in the original portrait vector of the a-th candidate talent relative to the b-th position to be matched, is the optimization fusion coefficient, which is a regulation factor for performing the reconstruction score, controlling the adjustment range to prevent the label ability score from being overly corrected, and the numerical range is in the interval from 0 to 1;

[0098] Candidates may have potential abilities in the corresponding label dimensions, but they are not fully expressed in the portrait generation stage. At this time, we need to try to guide the portrait to better align with the semantic structure of the position portrait to achieve the fitting and optimization of the semantic dimension. In other words, the score update does not change the reality, but strengthens the expression, which is similar to the principle of "semantic embedding optimization" in natural language processing. Instead of directly covering the old value (to avoid losing the original features), the difference is used as incremental information for fusion correction, controlling the amplitude of the correction. Its essence is to perform a "weighted score sliding update" to ensure the smooth evolution of the portrait rather than mutation, because the purpose of optimization is not to cover the position portrait, but to "make up for the ability gap", making the portrait converge to the target portrait without destroying its ontology structure;

[0099] And after performing the reconstruction score update, regenerate the original portrait vector of the corresponding candidate talent and mark it as the new portrait vector of the corresponding candidate talent .

[0100] The suggestion module returns the new portrait vector of the corresponding candidate talent to the matching unit to regenerate the candidate talent portrait adaptation ranking table, and extract the candidate talent with the largest SIM value in the regenerated candidate talent portrait adaptation ranking table. This candidate talent is used as the recommended candidate for the corresponding position to be matched.

[0101] In an embodiment of the present invention, the system automatically determines whether a candidate matches the job profile by comparing the matching value SIM with a preset threshold, ensuring that the recommended candidate has the actual job competence and improving the reliability of job recommendation. When the matching value is insufficient, the system can trigger a tag deviation analysis mechanism to calculate the ability difference between the candidate and the job in each tag dimension, identify the specific ability tags with deficiencies, and then perform a reconstruction score update on the deviated dimension to generate a new tag ability score, realizing the "local reinforcement" and "ability repair" of the candidate profile and enhancing the adaptive adjustment ability of the system. Secondly, after updating the profile, the system can re-rank based on the new profile and output relatively better recommended candidates, forming a recommendation feedback loop.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A talent label portrait intelligent analysis system based on big data, characterized by: include: The vector module obtains the talent data of multiple candidates based on multiple sources, and constructs the candidate label association vector after dimensionality reduction processing. , and combined with the dynamic nature of label semantics, update the candidate label association vector in real time ; The graph construction module associates the updated candidate label vectors , analyze the ability scores of each candidate under the corresponding label dimension to generate the original portrait vector of each candidate , and construct a label graph structure; The matching module obtains the job profile vector to be matched in advance , and compare it with the original portrait vector of each candidate talent Perform cosine similarity calculation to obtain the matching value SIM. If the matching value SIM does not exceed the preset threshold, execute the label offset analysis instruction to extract the unqualified labels of the corresponding candidate talents and optimize them to generate a new portrait vector for each candidate talent. ; Recommendation module, based on new profile vectors of each candidate , build a corresponding candidate talent profile adaptation ranking table for all positions to be matched to realize the job of job suggestion prompts.

2. According to the big data-based talent label portrait intelligent analysis system of claim 1, it is characterized by: The vector module includes a vector construction unit and a semantic update unit; The vector construction unit uses big data technology to obtain the talent data of each candidate from multiple sources, including resume systems, education platforms, recruitment platforms, project records, behavior logs and social media content, and performs text parsing, field regularization, missing value filling and noise removal on the talent data. The BERT embedding model is used to extract label words from the talent data, and the word vector corresponding to each label word is obtained. The word vector corresponding to each label word is matched with the job standard label vector in the job vocabulary by using BERT-CLS to map the word vector corresponding to each label word to the corresponding job standard label in the job vocabulary, and the word vector corresponding to each label word is reduced in dimension to construct the candidate label association vector. .

3. According to the big data-based talent label portrait intelligent analysis system of claim 2, it is characterized by: The semantic update unit updates the job standard labels in the job vocabulary in real time based on the dynamic nature of label semantics. Specifically, the semantic change trends of each job standard label in the job vocabulary in the past period are statistically analyzed to construct the label semantic drift entropy coefficient Ht. Based on the value of the label semantic drift entropy coefficient Ht, the semantic stability of the corresponding label in each time period is determined. If the semantics of the corresponding label is determined to be unstable in each time period, the corresponding label is marked as a pseudo-label, and the label set to be replaced is determined. By analyzing the label direction offset between the pseudo-label and each job standard label in the job vocabulary, the label direction offset value Dj is obtained to complete the judgment of the label transformation capability. The label set to be replaced refers to the job standard label vector other than the pseudo-label in the job vocabulary; specifically: ; in, is the word vector of the g-th pseudo label, is the standard label vector of each position in the label set to be replaced; is the cosine similarity of the angle between the pseudo label and the standard label vector space of each position in the label set to be replaced; Extract the job standard label corresponding to the minimum label direction offset value Dj, and merge it with the corresponding pseudo label to form a parallel extended label. Return the parallel extended label to the vector construction unit, and re-match it with the word vector corresponding to each label word in the talent data using BERT-CLS to update the candidate label association vector .

4. The talent label portrait intelligent analysis system based on big data according to claim 3 is characterized by: The graph construction module includes a feature unit, a scoring unit, and a construction unit; Feature unit, according to the updated candidate label association vector and talent data, identify the feature items of all candidate talents, analyze the response strength of each feature item under different label dimensions, and use semantic alignment to project each feature item under all label dimensions to obtain the relationship matrix V.

5. According to the big data-based talent label portrait intelligent analysis system of claim 4, it is characterized by: The scoring unit analyzes the ability score of each candidate in the corresponding label dimension based on the relationship matrix V and the TF-IDF algorithm to obtain the total ability score tss of each candidate in the corresponding label dimension, specifically: ; in, Score the overall ability of the corresponding candidate in the i-th label dimension. is the response value of the rth feature item of the corresponding candidate talent under the i-th label dimension, is the importance weight of the rth feature item, r is the feature item number, and m is the total number of feature items; The total ability score tss of each candidate talent under all label dimensions is counted to generate the original portrait vector of each candidate talent .

6. The talent label portrait intelligent analysis system based on big data according to claim 5 is characterized by: Updated candidate label association vector output from the semantic update unit Perform co-occurrence statistics to obtain a co-occurrence matrix. Each element in the co-occurrence matrix represents the number of co-occurrences Yc of two labels in all candidates. Based on the co-occurrence matrix, calculate the association intensity factor RAF between labels; Construct a unit, use the label as the node, the association intensity factor RAF as the edge between the nodes, and the original portrait vector of each candidate talent As a numerical representation of a label map structure, to draw the label map structure.

7. The talent label portrait intelligent analysis system based on big data according to claim 6 is characterized by: The matching module includes a matching unit and an optimization unit; Matching unit, pre-acquires the job portrait vector to be matched and combined with the original portrait vector of each candidate talent , obtain the matching value SIM, the matching value SIM is obtained according to the following formula: ; in, is the matching value of the a-th candidate talent relative to the b-th position portrait vector to be matched, a is the number of the candidate talent, b is the number of the position to be matched, is the original portrait vector of the a-th candidate talent, is the portrait vector of the bth position to be matched.

8. The talent label portrait intelligent analysis system based on big data according to claim 7 is characterized by: The optimization unit compares the matching value SIM with the preset threshold. If the matching value SIM exceeds the preset threshold, it indicates that the current candidate talent is relatively poor relative to the corresponding job profile vector to be matched. The match is qualified, and the current candidate is included in the candidate talent portrait adaptation sorting table corresponding to the qualified matching position; if the matching value SIM does not exceed the preset threshold, the label offset analysis instruction is executed.

9. The talent label portrait intelligent analysis system based on big data according to claim 8 is characterized by: Receive label deviation analysis instructions, analyze the deviation of each label dimension of the corresponding candidate talent in the corresponding matching position, and obtain the difference , if the difference is satisfied >0, then perform reconstruction score update: ; in, is the total ability score adjusted by the i-th label dimension in the original portrait vector of the a-th candidate relative to the b-th position to be matched, To optimize the fusion coefficient; is the difference between the i-th label dimension in the original portrait vector of the a-th candidate talent and the b-th position to be matched; After performing the reconstruction score update, the original portrait vector of the corresponding candidate talent is regenerated , and marked as the new portrait vector of the corresponding candidate talent .

10. The talent label portrait intelligent analysis system based on big data according to claim 9 is characterized by: The recommendation module converts the new profile vector of the corresponding candidate talent into Return to the matching unit to regenerate the candidate talent portrait adaptation ranking table, and extract the candidate talent with the largest corresponding matching value SIM value in the regenerated candidate talent portrait adaptation ranking table, and use the candidate talent as the recommended candidate for the corresponding position to be matched.

Citation Information

Patent Citations

  • Resume screening method and device, terminal and computer readable storage medium

    CN110263818A

  • Resume document matching method and device, computing equipment and storage medium

    CN114117222A

  • People and post matching analysis method and device based on big data portraits

    CN116523268A

  • Talent management method and system based on objective data evaluation

    CN119809583A

Cited By

  • Talent evaluation management method and system based on AI intelligence

    CN120338741A

  • Talent evaluation management method and system based on AI intelligence

    CN120338741B

  • Human resource matching method and system based on machine learning

    CN120851821A

  • A human resource matching method and system based on machine learning

    CN120851821B

  • Knowledge graph-based post demand dynamic portrait construction method and system

    CN120874833A