Talent evaluation management method and system based on AI intelligence
Through OCR, speech transcription and domain knowledge graph technology, unstructured data is converted into standardized text, combined with context semantic coding and dynamic semantic matching, the problems of information omission and ambiguity in the existing technology are solved, the multi-dimensional accuracy and rationality of talent assessment are achieved, and the accuracy and reliability of evaluation are improved.
Patent Information
- Application Number
- CN202510757971.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-09
AI Technical Summary
In the prior art, enterprises rely on manual screening of paper resumes and manual recording of interview content during talent assessment, resulting in inefficient information extraction and easy to miss key content. Traditional OCR technology has low accuracy in recognition of complex handwritten characters, lacks contextual semantic correlation analysis of speech transcription, and text fragmentation often results in format differences. The application of existing knowledge graphs is mostly limited to static term matching, and lacks the ability to align dynamic disambiguation and cross-domain nodes.
The handwritten resume is converted into structured text through OCR image recognition, and the speech transfer converts the interview recording into dialogue text with timing marks, and performs format analysis to generate a standardized text data set containing semantic labels; the domain knowledge graph is used for context semantic coding and semantic disambiguation, and generates text feature vectors containing entity association relationships; dynamic semantic matching text feature vectors and job requirements vectors, analyzes the emotional tendency characteristics of dialogue text, generates personality trait evaluation parameters, and constructs quadrilaterals through four orthogonal detection points for dimension calibration.
It realizes standardized processing of unstructured data, accurately aligns professional terms and eliminates ambiguity, ensures the multi-dimensional accuracy and rationality of talent assessment, improves the accuracy and reliability of evaluation, and provides interactive evaluation reports to support decision-making.
Smart Images

Figure CN120338741B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence analysis technology, and in particular to a talent evaluation management method and system based on AI intelligence. Background Art
[0002] In existing technologies, most companies rely on manual screening of paper resumes, manual recording of interview content, or electronic tools based on keyword matching for talent assessment. These technologies have the following significant drawbacks:
[0003] Unstructured data such as handwritten resumes and interview recordings are difficult to process in a standardized manner, resulting in inefficient information extraction and the easy omission of key content. Traditional OCR technology has low recognition accuracy for complex handwriting, speech transcription lacks contextual semantic analysis, and electronic document parsing often leads to text fragmentation due to format differences. Polysemous words, industry terms and abbreviations are common in resumes and interview texts. Traditional methods rely on manual experience for semantic judgment, which can easily lead to misjudgment of job suitability due to contextual understanding deviations. Existing knowledge graph applications are mostly limited to static term matching and lack dynamic disambiguation and cross-domain node alignment capabilities. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a talent evaluation management method and system based on AI intelligence, which can realize accurate talent evaluation through dynamic semantic matching and multi-dimensional detection point calibration.
[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0006] First, a talent evaluation and management method based on AI intelligence, the method comprising:
[0007] Step S1: Convert the handwritten resume image into structured text data through OCR image recognition; convert the interview recording into conversation text with time sequence tags through speech transcription; parse the electronic document format to extract the original text content; encode and standardize the structured text data, conversation text, and original text content to generate a standardized text dataset containing semantic tags;
[0008] Step S2: contextual semantic encoding is performed on the standardized text dataset, and node alignment and semantic disambiguation processing of professional terms are performed through the domain knowledge graph to generate a text feature vector containing entity association relationships;
[0009] Step S3: Dynamically semantically match the text feature vector with the job requirement vector, optimize the vector representation of polysemous words in the job context, analyze the emotional tendency characteristics of the dialogue text, and generate personality trait assessment parameters;
[0010] Step S4: Set four orthogonal detection points in the semantic matching result, construct a quadrilateral based on the semantic coverage, context dependency, term accuracy, and sentiment consistency index corresponding to the detection points, generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, and perform dimension calibration on the text feature vector to obtain a calibrated text feature vector;
[0011] Step S5: Based on the calibrated text feature vector and personality trait assessment parameters, combined with the job requirement vector, an interactive evaluation report including a dynamic skill topology map, fitness heat distribution, and career development prediction curve is generated.
[0012] Furthermore, handwritten resume images are converted into structured text data through OCR image recognition; interview recordings are converted into conversational text with time series tags through speech transcription; electronic documents are formatted and the original text content is extracted; the structured text data, conversational text, and original text content are coded and standardized to generate a standardized text dataset containing semantic tags, including:
[0013] When performing OCR image recognition on handwritten resume images, an adaptive threshold segmentation algorithm is used to perform illumination equalization on the image, segment the candidate text area, and parse the text area. The recognition results are mapped into structured field data according to the resume field classification. Among them, the education background field is parsed into the name of the school, major and degree level, the work experience field is parsed into the company, position and job description, and the skill certificate field is parsed into the certificate name and certification body. When performing voice transcription on interview recordings, different speakers are identified based on voiceprint features. During the recognition process, the mixed audio is separated into independent audio tracks and combined with the voice end Point detection transcribes each audio track in segments, generating a time-stamped dialogue text. The dialogue text is annotated with speech emotional fundamental frequency features, including intonation fluctuation frequency, semantic pause intervals, and speech rate change rate. These emotional fundamental frequency features are extracted using acoustic parameters and bound to text paragraphs. When formatting electronic documents, a preset template engine is used based on the document type to identify the hierarchical structure of tables, paragraphs, and headings in the document, extracting the unformatted original text content and mapping the text content to rows and columns. The association between merged cells across columns is achieved through contextual semantic reconstruction.
[0014] The structured field data, conversation text with time tags and original text content are input into the coding standardization module. Semantic tags are added to skill names and project experience keywords through the preset industry terminology dictionary, and the time format is unified into a standard format and the place names are normalized into the full names of administrative divisions. A standardized text dataset containing semantic tags is generated. The dataset is stored in a structured document format according to the field type, and each field is associated with a semantic tag and a normalized attribute value.
[0015] Furthermore, contextual semantic encoding is performed on the standardized text dataset, and node alignment and semantic disambiguation of professional terms are performed through the domain knowledge graph to generate text feature vectors containing entity association relationships, including:
[0016] Based on the semantic tags and structured document formats in the standardized text dataset, contextual semantic encoding technology is used to perform semantic analysis on the text and generate initial semantic vectors;
[0017] The initial semantic vector is input into the domain structured semantic network for term alignment. By traversing the professional term nodes in the semantic network, the semantic similarity between the terms in the text and the nodes is calculated. Terms with similarity above a preset threshold are mapped to corresponding nodes. The contextual semantics of the terms are then expanded based on the association paths between the nodes. The expanded contextual semantics of the terms include upstream and downstream skill associations, industry scenario constraints, and job capability dependencies.
[0018] Disambiguation weights are constructed using the co-occurrence relationships and hierarchical classification information of nodes in the semantic network. The probability distribution of ambiguous terms is calculated based on the context window of the terms in the text, and the vector representation of the terms is dynamically adjusted to eliminate semantic ambiguity and obtain disambiguation results.
[0019] The aligned term node features, extended semantics and disambiguation results are integrated, and the attribute features of the nodes in the semantic network are concatenated with the text semantic vector to generate a text feature vector containing entity association relationships.
[0020] Furthermore, dynamic semantic matching is performed between the text feature vector and the job requirement vector, the vector representation of polysemous words in the job context is optimized, the emotional tendency characteristics of the dialogue text are analyzed, and personality trait assessment parameters are generated, including:
[0021] Based on the job description, skill requirements, years of experience, and ability keywords are extracted to construct a job requirement vector. The skill requirement keywords are semantically enhanced based on the node attributes in the domain structured semantic network. Semantic enhancement extracts upstream and downstream skill dependencies and industry scenario constraints by traversing the associated paths of the nodes in the semantic network, generating a job requirement vector consistent with the dimensions of the text feature vector.
[0022] When expanding the co-occurrence distribution of polysemous words through the association paths in the semantic network, the co-occurrence word set related to job requirements is dynamically loaded based on the upstream and downstream skill dependencies between nodes and industry scenario constraints;
[0023] Count the frequency and distribution density of co-occurring words in job requirements descriptions to generate coverage indicators;
[0024] Adjust the vector weights of polysemous words based on the coverage index to generate an optimized semantic matching score;
[0025] The optimized semantic matching scores are input into the multimodal fusion layer, and the conversation text is segmented by speaker based on time sequence markers. The acoustic feature fundamental frequency parameters of each text segment are extracted. The acoustic features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate.
[0026] The emotional keywords in each text are matched through the emotional dictionary, and the emotional tendency intensity value is calculated by combining the acoustic feature fundamental frequency parameters;
[0027] The sentiment intensity value and text content are input into the sentiment calculation module, which fuses the acoustic features and text semantics through time series alignment technology to output a smoothed sentiment polarity score sequence.
[0028] Calculate the variance of the emotional stability index based on the fluctuation amplitude and frequency of the emotional polarity score in the time series;
[0029] In the multimodal fusion layer, the semantic matching score is fused with the sentiment polarity score and the emotional stability index at the feature level to generate a joint vector. The data distribution pattern of the joint vector is then analyzed to generate personality trait assessment parameters, including decision-making tendency, teamwork degree, and stress tolerance dimensions.
[0030] Furthermore, the optimized semantic matching score is input into the multimodal fusion layer, and the conversation text is segmented by speaker based on the time sequence markers. The acoustic feature fundamental frequency parameters of each text segment are extracted. The acoustic features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate, including:
[0031] Based on the time sequence mark of the dialogue text, the mixed audio is cut into independent speech segments according to the audio track segment identifier of the speaker, and the fundamental frequency of each speech segment is extracted to generate fundamental frequency change data;
[0032] According to the first-order difference value of the fundamental frequency change data, the number of sudden changes in the direction and amplitude of the fundamental frequency change in the speech segment is counted, and the frequency of intonation fluctuation per unit time is calculated based on the segment duration;
[0033] Perform speech endpoint detection on the speech segment, identify the starting position of the semantic demarcation point, divide the speech segment into semantic unit segments based on the demarcation point, and calculate the average of the silence intervals between adjacent semantic unit segments as the semantic pause interval;
[0034] Extract the total duration of the speech segment and the number of characters in the corresponding text segment, calculate the initial speaking rate based on the ratio of the number of characters to the speech duration, and perform linear normalization on the initial speaking rate based on the preset industry standard speaking rate range to generate the speaking rate change rate;
[0035] The intonation fluctuation frequency, semantic pause interval and speech rate change rate are bound to the timing markers of the corresponding text segments to generate the fundamental frequency parameters of structured acoustic features.
[0036] Furthermore, the sentiment intensity value and text content are input into the sentiment calculation module, and the acoustic features and text semantics are integrated through time series alignment technology to output a smoothed sentiment polarity score sequence, including:
[0037] Based on the time sequence markings of the conversation text and the timestamps of the acoustic feature fundamental frequency parameters, a time sequence mapping relationship between text paragraphs and speech segments is established to generate a time-aligned multimodal data index;
[0038] Dynamically segment text paragraphs based on temporal mapping relationships, extract the speech segmentation time window corresponding to each sentence, and match the intensity weights of emotional keywords in the corresponding sentences based on the emotional dictionary;
[0039] The acoustic feature fundamental frequency parameters and emotional keyword intensity weights within the same time window are concatenated for multimodal data. The fused features on the time series are scanned by the convolution kernel to extract the cross-modal contextual association feature vector.
[0040] The emotional tendency offset is calculated based on the distribution density of the cross-modal contextual association feature vector, and the initial emotional polarity score is generated according to the projection direction and modulus of the offset in the preset emotional dimension coordinate system;
[0041] The initial sentiment polarity scores of the continuous time window are subjected to sliding window mean filtering, and a smoothed sentiment polarity score sequence is output.
[0042] Furthermore, four orthogonal detection points are set in the semantic matching results. A quadrilateral is constructed based on the semantic coverage, context dependency, term accuracy, and sentiment consistency indicators corresponding to the detection points. A dynamic correction coefficient is generated by calculating the vertex distribution eigenvalues and side length correlation coefficients of the quadrilateral, and the text feature vector is dimensionally calibrated to obtain the calibrated text feature vector, including:
[0043] Four orthogonal detection points are set in the semantic matching results, corresponding to semantic coverage, context dependency, term accuracy and sentiment consistency indicators respectively;
[0044] The coordinates of the quadrilateral vertices are constructed based on the index values of the four orthogonal detection points, and the dynamic correction coefficient is generated by calculating the vertex distribution characteristic value and the side length correlation coefficient of the quadrilateral;
[0045] The dynamic correction coefficient is used to calibrate the dimension of the text feature vector to generate a calibrated text feature vector.
[0046] Furthermore, the coordinates of the quadrilateral vertices are constructed based on the index values of the four detection points. By calculating the vertex distribution characteristic values and the side length correlation coefficient of the quadrilateral, a dynamic correction coefficient is generated, including:
[0047] Map the semantic coverage, context dependency, term accuracy, and sentiment consistency indicators to four vertices in a two-dimensional coordinate system, with semantic coverage and term accuracy as the positive and negative components of the horizontal axis, and context dependency and sentiment consistency as the positive and negative components of the vertical axis, to generate an initial quadrilateral vertex coordinate set.
[0048] Calculate the geometric center of gravity coordinates of the quadrilateral based on the vertex coordinate set, and perform relative coordinate transformation on the four vertex coordinates with the center of gravity coordinates as the origin to generate normalized vertex offsets based on the center of gravity;
[0049] The standard deviation is calculated based on the distribution dispersion of the normalized vertex offsets to generate vertex distribution feature values that represent the degree of spatial dispersion of the vertices. Based on the initial quadrilateral vertex coordinate set, the Euclidean distance ratio between adjacent vertices is calculated as the side length proportional coefficient, and the cosine value of the angle between each side vector is extracted. The cosine values are arithmetic averaged to generate the side length correlation coefficient.
[0050] The vertex distribution eigenvalues and the edge length correlation coefficients are linearly superimposed according to the preset weight coefficients to generate the initial correction factor, which is then interval-normalized to output the dynamic correction coefficient.
[0051] Secondly, the AI-based talent evaluation and management system includes:
[0052] The data collection module is used to extract multi-source resume information through OCR recognition, speech transcription, and format parsing, and standardize it into a text dataset with semantic labels;
[0053] The semantic encoding module is used to perform contextual semantic encoding and domain knowledge graph disambiguation on standardized text datasets to generate text feature vectors containing entity association relationships;
[0054] Dynamic semantic matching module, used to dynamically match text feature vectors with job requirement vectors and analyze conversation sentiment to generate personality trait assessment parameters;
[0055] A dynamic correction module is used to construct an evaluation quadrilateral by setting orthogonal detection points, calculate eigenvalues and correlation coefficients to generate correction coefficients, and generate calibrated eigenvectors;
[0056] The evaluation report module is used to generate an interactive evaluation report based on the calibrated feature vectors, personality parameters and job requirement vectors.
[0057] According to a third aspect, a computer-readable storage medium stores a program, which implements the method described above when executed by a processor.
[0058] The above solution of the present invention includes at least the following beneficial effects:
[0059] Leveraging technologies like OCR image recognition and speech transcription, data in various formats, such as handwritten resumes, interview recordings, and electronic documents, is uniformly converted into standardized text, breaking down data barriers and preventing information omissions. Data is also structured through semantic tagging. Utilizing contextual semantic encoding and domain knowledge graphs, the system deeply mines text semantic information, precisely aligns professional terminology, and eliminates ambiguity. This constructs a text feature vector containing entity relationships, accurately capturing the candidate's skill content and relationships, and avoiding misjudgments caused by superficial keyword matching. Dynamic semantic matching is achieved between text features and job requirements, with contextual adaptation for polysemous words. Furthermore, sentiment analysis of conversational text is combined to generate personality trait parameters. Talent is comprehensively assessed across multiple dimensions, including professional skills, semantic understanding, and emotional traits, to create a comprehensive candidate profile.
[0060] A quadrilateral is constructed from four orthogonal checkpoints, and the balance of matching results is tested from dimensions such as semantic coverage and contextual dependency. Dynamic correction coefficients are generated by calculating the quadrilateral's geometric features, and the text feature vector is calibrated to ensure that the assessment results are more closely aligned with the actual job requirements, improving evaluation accuracy and reliability. The calibrated assessment results are presented in visual formats such as dynamic skill topology maps, compatibility heat maps, and career development prediction curves. Interactive operations are also supported, helping decision makers quickly identify candidates' strengths and weaknesses. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 This is a flow chart of the talent evaluation and management method based on AI intelligence provided by an embodiment of the present invention.
[0062] Figure 2 This is a schematic diagram of an AI-based talent evaluation and management system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0063] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0064] like Figure 1As shown, an embodiment of the present invention proposes a talent evaluation management method based on AI intelligence, which includes the following steps:
[0065] Step S1: Convert the handwritten resume image into structured text data through OCR image recognition; convert the interview recording into conversation text with time sequence tags through speech transcription; parse the electronic document format to extract the original text content; encode and standardize the structured text data, conversation text, and original text content to generate a standardized text dataset containing semantic tags;
[0066] Step S2: contextual semantic encoding is performed on the standardized text dataset, and node alignment and semantic disambiguation processing of professional terms are performed through the domain knowledge graph to generate a text feature vector containing entity association relationships;
[0067] Step S3: Dynamically semantically match the text feature vector with the job requirement vector, optimize the vector representation of polysemous words in the job context, analyze the emotional tendency characteristics of the dialogue text, and generate personality trait assessment parameters;
[0068] Step S4: Set four orthogonal detection points in the semantic matching result, construct a quadrilateral based on the semantic coverage, context dependency, term accuracy, and sentiment consistency index corresponding to the detection points, generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, and perform dimension calibration on the text feature vector to obtain a calibrated text feature vector;
[0069] Step S5: Based on the calibrated text feature vector and personality trait assessment parameters, combined with the job requirement vector, an interactive evaluation report including a dynamic skill topology map, fitness heat distribution, and career development prediction curve is generated.
[0070] In an embodiment of the present invention, by leveraging contextual semantic encoding and domain knowledge graphs, the implicit semantics behind the text are deeply explored, professional terms are precisely aligned and ambiguity is eliminated, and a text feature vector containing skill associations is constructed. This accurately captures the connotation and extension of a candidate's skills, avoiding misjudgments caused by relying solely on surface keywords, and making talent capability assessment more in-depth and accurate. Dynamic semantic matching is achieved between text feature vectors and job requirement vectors, and contextual adaptation is performed for polysemous words to ensure the accuracy of skill matching. At the same time, by analyzing the emotional tendencies and acoustic characteristics of the conversation text, the personality trait parameters of the candidate are quantified, filling the gap in the objective assessment of soft skills in traditional talent evaluation, and comprehensively portraying the candidate's portrait from multiple dimensions such as professional skills, semantic understanding, and emotional traits.
[0071] By setting four orthogonal checkpoints—semantic coverage, contextual dependency, terminology accuracy, and sentiment consistency—a quadrilateral is constructed to examine the balance of semantic matching results from multiple perspectives, effectively avoiding evaluation bias caused by unilaterally over- or under-performing indicators. Dynamic correction coefficients are generated based on the geometric characteristics of the quadrilateral to calibrate the dimensions of text feature vectors and automatically adjust the weights of each dimension, making the evaluation more tailored to the actual job requirements and enhancing the rationality and reliability of the evaluation. Decision makers can quickly identify a candidate's strengths and weaknesses. Furthermore, the career development prediction curve, combining industry trends with candidate characteristics, provides a forward-looking reference for companies developing talent plans and providing candidates with career development directions.
[0072] In a preferred embodiment of the present invention, the above step S1, converting the handwritten resume image into structured text data through OCR image recognition; converting the interview recording into a time-stamped dialogue text through speech transcription; performing format parsing on the electronic document to extract the original text content; and encoding and standardizing the structured text data, dialogue text, and original text content to generate a standardized text dataset containing semantic tags, may include:
[0073] Step S100: When performing OCR image recognition on a handwritten resume image, an adaptive threshold segmentation algorithm is used to perform illumination equalization on the image, segment the candidate text area, and parse the text area. The recognition results are mapped into structured field data according to the resume field classification. Among them, the education background field is parsed into the name of the school, major and degree level, the work experience field is parsed into the company, position and job description, and the skill certificate field is parsed into the certificate name and certification body. When performing voice transcription on the interview recording, different speakers are identified based on voiceprint features. During the recognition process, the mixed audio is separated into independent audio tracks and combined with Voice endpoint detection transcribes each audio track in segments, generating a time-stamped dialogue text. The dialogue text is annotated with speech emotional fundamental frequency features, including intonation fluctuation frequency, semantic pause intervals, and speech rate change rate. These emotional fundamental frequency features are extracted using acoustic parameters and bound to text paragraphs. When formatting electronic documents, a preset template engine is used based on the document type to identify the hierarchical structure of tables, paragraphs, and headings in the document, extracting the unformatted original text content and mapping the text content to rows and columns. The association between merged cells across columns is achieved through contextual semantic reconstruction.
[0074] Step S101: Input the structured field data, the conversation text with time stamps, and the original text content into the coding standardization module, add semantic tags to the skill names and project experience keywords through the preset industry term dictionary, unify the time format into a standard format, and normalize the place names into the full names of administrative divisions, thereby generating a standardized text dataset containing semantic tags. The dataset is stored in a structured document format according to the field type, and each field is associated with a semantic tag and a normalized attribute value.
[0075] In the embodiment of the present invention, the specific logic of the adaptive threshold segmentation algorithm is:
[0076] The image is divided into multiple local subregions (e.g., an 8x8 pixel grid). The mean and variance of the pixel grayscale values are calculated for each subregion. The threshold is dynamically adjusted based on the brightness and darkness of the subregion (for example, lowering the threshold in dark areas to highlight text, and raising it in bright areas to suppress background noise). This subregion-by-subregion processing eliminates the effects of global illumination unevenness (such as localized shadows or reflections on resume scans), significantly improving the contrast between text and background.
[0077] After equalization, edge detection (such as the Canny operator) is used to identify text outlines in the image. Morphological operations (dilation and erosion) are then used to merge adjacent edges to form continuous rectangular candidate regions. Regions that are too small (e.g., less than 10x10 pixels) or have unusual aspect ratios (e.g., height much greater than width, possibly representing table lines) are filtered out. Only regions that meet the text arrangement requirements (e.g., with a moderate height and width sufficient to accommodate multiple characters) are retained.
[0078] Field parsing rules:
[0079] Educational Background:
[0080] When keywords such as "university," "college," "undergraduate," and "master's" are identified, the educational background analysis logic is triggered. For example, if the text contains "University A | Computer Science and Technology | Master of Engineering," the system uses vertical bar delimiters or fixed sentence structures (such as "Graduated from XX University, XX Major") to extract the institution name (University A), major (Computer Science and Technology), and degree level (Master of Engineering).
[0081] Work experience:
[0082] Use timeline keywords (e.g., "2020.01-2023.12" or "Present") or responsibility-guided keywords (e.g., "responsible for," "Main duties include") to locate work experience sections. For example, "XX Company | Software Engineer | Responsible for back-end system development, optimizing database query efficiency by 30%" will be parsed as Company (XX Company), Position (Software Engineer), and Responsibilities (Responsible for back-end system development, optimizing database query efficiency by 30%).
[0083] Voiceprint samples of interviewers and candidates are collected in advance (such as the self-introduction clip at the beginning of the interview), and voiceprint feature vectors (such as MFCC Mel-Frequency Cepstral Coefficients) are extracted using a Gaussian Mixture Model (GMM). During the transcription process, the cosine similarity between the voiceprint features of the audio stream and the preset samples is calculated in real time. Segments with a similarity above 90% are assigned to the corresponding main audio tracks (such as the candidate track and the interviewer track), achieving separation of multi-person conversations.
[0084] The specific logic of voice endpoint detection (VAD):
[0085] A dual-threshold method is used to detect the start and end of speech: a speech segment begins when the audio energy exceeds a "high threshold," and ends when the energy remains below a "low threshold" for at least 0.5 seconds. For example, when a candidate answers a question, a pause of more than 0.5 seconds is considered the end of a semantic unit, generating a separate text segment with a starting timestamp (e.g., "00:02:15-00:03:40").
[0086] For each speech segment, the fundamental frequency curve (reflecting pitch changes) is extracted. The number of rising and falling inflection points in the curve is calculated and divided by the segment duration (in seconds) to obtain the number of intonation fluctuations per minute (e.g., if a 20-second segment of speech has five inflection points, the intonation fluctuation frequency is 15 times / minute). At the semantic unit boundaries obtained by speech endpoint detection, the duration of silence between two adjacent units is measured (e.g., if unit A ends at 00:02:30 and unit B begins at 00:02:32, the pause interval is 2 seconds). The average of all pause intervals in the audio track is taken as the feature value.
[0087] Calculate the original speaking rate (number of text characters / speech duration, unit: words / second), then compare it with the industry standard speaking rate range (such as 120-150 words / minute). Use linear transformation to map the original speaking rate to the range of 0-1 (for example, if the original speaking rate is 180 words / minute, which exceeds the standard upper limit by 30%, it will be normalized to 1.3).
[0088] 1. Document type adaptation of template engine:
[0089] Word document: By parsing XML tags (such as <w:p>paragraph, <w:tbl>table) to identify the header level (e.g. <w:outlinelvl>The header level of the mark) and table structure are automatically ignored when extracting paragraph text.
[0090] PDF documents: Use the PDF parsing library to identify text flow and coordinate positions, and determine the paragraph level based on the Y-axis value of the text coordinates (paragraphs with the same or similar Y-axis values are considered the same paragraph). For scanned PDFs, the OCR engine is used for secondary recognition.
[0091] Excel tables: Read cell coordinates and merge properties. For cross-column merged cells (such as A1:C1 merge), the actual value of the merged cell is inferred by traversing the contextual logic of the same business data (such as the "Department" column content in adjacent rows) (for example, "Sales Department" is displayed across columns in A1:C1).
[0092] When a cross-column merge is detected, the system scans the data in the same column of adjacent rows upwards and downwards, looking for repeated keywords (such as "Project Name" and "Person in Charge") and infers the merged cell's attributes through pattern matching. For example, if A1:C1 is merged in a table, and the row below A2 contains "Project 1," B2 contains "Zhang San," and C2 contains "2023," then A1:C1 is inferred to be the "Project Information" header row, avoiding content fragmentation due to formatting issues.
[0093] The dictionary predefines standard terminology mappings across different industries (e.g., "JAVA" is uniformly mapped to "Java programming language," and "PM" is mapped to "project manager" or "product manager" depending on the context). Using a forward maximum matching algorithm, the dictionary scans text and adds semantic tags to skill names (e.g., "AI development" is matched to "artificial intelligence development") and project keywords (e.g., "big data platform" is matched to "big data analysis platform") to ensure terminology consistency.
[0094] Normalization of time and place:
[0095] Time format is unified:
[0096] Recognizes multiple time expressions (such as "May 2023", "2023.05", and "2023-05") and converts them uniformly into the standard "YYYY-MM" format. For ambiguous time (such as "the past three years"), it combines the current time (such as 2025) and converts them into the "2022-2025" interval.
[0097] Through technologies such as adaptive threshold segmentation, voiceprint recognition, and template engines, data errors caused by factors such as image blur, confusion between multiple voices, and complex document formats are reduced. Semantic tagging and format normalization ensure uniform data standards. Operations such as the annotation of voice emotional fundamental frequency features and cross-column cell semantic reconstruction add dimensional information such as emotional expression and content logic to the data, fully preserving key content from resumes and interviews and preventing information omissions. Standardized datasets can adapt to a variety of data analysis models and algorithms, reducing the time and cost of data preprocessing. Structured storage also facilitates rapid retrieval and access to data.
[0098] In a preferred embodiment of the present invention, step S2, performing contextual semantic encoding on the standardized text dataset, performing node alignment and semantic disambiguation processing on professional terms through the domain knowledge graph, and generating a text feature vector containing entity association relationships, may include:
[0099] Step S200 , based on the semantic tags and structured document format in the standardized text dataset, a contextual semantic coding technique is used to perform semantic analysis on the text to generate an initial semantic vector;
[0100] Step S201: Input the initial semantic vector into the domain structured semantic network for term alignment. By traversing the professional term nodes in the semantic network, the semantic similarity between the terms in the text and the nodes is calculated. Terms with similarity above a preset threshold are mapped to corresponding nodes. The contextual semantics of the terms are then expanded based on the association paths between the nodes. The expanded contextual semantics of the terms include upstream and downstream skill associations, industry scenario constraints, and job capability dependencies.
[0101] Step S202: constructing disambiguation weights using the co-occurrence relationships and hierarchical classification information of nodes in the semantic network, calculating the probability distribution of ambiguous terms based on the context window of the terms in the text, and dynamically adjusting the vector representation of the terms to eliminate semantic ambiguity and obtain a disambiguation result;
[0102] Step S203 , integrating the aligned term node features, extended semantics, and disambiguation results, and performing feature splicing on the attribute features of the nodes in the semantic network and the text semantic vector to generate a text feature vector containing entity association relationships.
[0103] In an embodiment of the present invention, the standardized text data is split into independent text blocks according to fields (such as "educational background" and "work experience"), and each text block is segmented (such as splitting "responsible for the back-end development of the e-commerce platform, using Java and Spring framework" into "responsible for", "e-commerce platform", "back-end development", "using", "Java", and "Spring framework"). A pre-trained language model (such as a BERT-type model) is used to analyze the contextual dependencies of each word. For example, "Java" is strongly associated with "back-end development" in "Using Java to develop back-end systems", while it focuses more on basic concepts in "Introduction to Java Programming Language". The model captures this semantic difference through a multi-layer neural network and generates an embedding vector (such as a 300-dimensional vector) containing contextual information for each word. The word vectors of the same text block are aggregated through average pooling or an attention mechanism to generate the initial semantic vector of the text block. For example, the vector of the "work experience" field needs to integrate the semantic information of all keywords in the job description.
[0104] In step S201, a semantic network is constructed based on the industry knowledge graph. Nodes are professional terms (e.g., "Python," "cloud computing," "agile development"), and edges are relationships between terms (e.g., "Python → programming language," "cloud computing → AWS / Azure," "agile development → Scrum framework"). Node attributes are also annotated (e.g., "programming language," "tool platform"). For example, if the term "SpringBoot" appears in a text, the system traverses the semantic network nodes and calculates their semantic similarity with the "Spring framework" node (using cosine similarity or edit distance). If the similarity exceeds a threshold (e.g., 0.8), "SpringBoot" is mapped to the "Spring framework" node. After mapping, the semantics are expanded along the associated paths of the semantic network nodes:
[0105] Upstream and downstream skill associations: The "Spring Framework" node is associated with child nodes such as "Dependency Injection" and "Aspect-Oriented Programming." The expanded semantics include these technical details.
[0106] Industry scenario constraints: If the position is "Backend Development in the Financial Industry," the "Spring Framework" node will be associated with industry-specific constraints such as "Financial-Grade Transaction Processing."
[0107] Job capability dependency: If the job requirement includes "microservice architecture," the "Spring Framework" node is further associated with capabilities such as "Spring Cloud" and "Service Registration and Discovery."
[0108] In step S202, disambiguation weights are assigned to each term node based on the co-occurrence relationships of nodes in the semantic network (e.g., "algorithm" often co-occurs with "data structure" and "machine learning") and hierarchical classifications (e.g., "algorithm" belongs to the "technical capability" hierarchy). For example, "algorithm" co-occurs frequently with the "machine learning" node in the context of "machine learning algorithm optimization," so the disambiguation weight favors "machine learning-related algorithms."
[0109] Probability calculation of ambiguous terms:
[0110] Centered on a term, we take the five preceding and succeeding words as a context window (e.g., "algorithm" in "optimization algorithm performance") and calculate the strength of association between the words within the window and each node in the semantic network. For example, "performance" has a strong association with the "algorithm optimization" node but a weak association with the "data structure" node. Therefore, "algorithm" in this context is more likely to refer to "optimization algorithm" rather than "basic algorithm."
[0111] Vector dynamic adjustment:
[0112] Adjust the initial term vector based on the probability distribution of the context window. For example, if "algorithm" has a 70% probability of belonging to the "Optimization Algorithm" node and a 30% probability of belonging to the "Basic Algorithm" node, its vector is the weighted sum of the two node vectors (weights 0.7:0.3). This eliminates ambiguity and more accurately reflects the semantics.
[0113] Step S203: Concatenate the following features by dimension into a final text feature vector:
[0114] Aligned node features: attribute vectors of term mapping nodes (e.g., the "Programming Language" and "Enterprise Development" attributes of the "Spring Framework" node);
[0115] Extended semantic features: Feature vectors such as upstream and downstream skills and industry scenarios extracted from association paths (e.g., "microservice architecture" and "financial compliance");
[0116] Disambiguated vector: dynamically adjusted term vector that eliminates ambiguous interference;
[0117] Original text semantic vector: the initial vector generated in step S200, which retains unstructured context information.
[0118] Vector dimension example:
[0119] Assuming that the initial vector is 300-dimensional, the node attribute vector is 100-dimensional, the extended semantic vector is 200-dimensional, and the disambiguation vector is 300-dimensional, the total dimension of the concatenated text feature vector is 300+100+200+300=900-dimensional, which fully covers the term semantics, association relationships, and contextual information.
[0120] Contextual semantic encoding captures subtle differences in terminology across different contexts (e.g., the different meanings of "product manager" in internet and traditional industries), preventing misunderstandings caused by polysemy. The domain knowledge graph's term alignment mechanism ensures that specialized terminology aligns with industry standards (e.g., accurately mapping "AI development" to the "artificial intelligence development" node), reducing the impact of non-standardized expressions. The semantic network's association path expansion (e.g., "Java → backend development → microservice architecture") ensures that text feature vectors incorporate deep skill dependencies rather than isolated terminology. After disambiguation, the vectors accurately reflect the specific context of the terms (e.g., distinguishing "deep learning algorithm" from "data structure algorithm"), enhancing the sophistication of semantic representation. Feature vectors incorporating entity associations can be directly used to dynamically match job requirements with candidate capabilities (e.g., detecting the skill chain compatibility between "Java development" and "SpringCloud microservices"). Structured feature vectors are adaptable to advanced analytical models such as graph neural networks (GNNs), providing a rich semantic foundation for subsequent personality trait assessment and career development prediction. The disambiguation mechanism effectively handles ambiguous statements in resumes (such as "responsible for system optimization"), infers specific technical directions (such as "database optimization" or "algorithm optimization") through context, and reduces the impact of data noise.
[0121] In a preferred embodiment of the present invention, step S3, which dynamically semantically matches the text feature vector with the job requirement vector, optimizes the vector representation of polysemous words in the job context, analyzes the emotional tendency characteristics of the dialogue text, and generates personality trait assessment parameters, may include:
[0122] Step S300: Extract skill requirements, years of experience, and ability keywords from the job description to construct a job requirement vector. The skill requirement keywords are semantically enhanced based on node attributes in the domain structured semantic network. Semantic enhancement extracts upstream and downstream skill dependencies and industry scenario constraints by traversing the associated paths of nodes in the semantic network, generating a job requirement vector consistent with the dimensions of the text feature vector.
[0123] Step S301: When expanding the co-occurrence distribution of polysemous words through the association path in the semantic network, a co-occurrence word set related to job requirements is dynamically loaded based on the upstream and downstream skill dependencies between nodes and industry scenario constraints;
[0124] Step S302: Count the frequency and distribution density of co-occurring words in the job requirement description to generate a coverage index;
[0125] Step S303: adjusting the vector weight of the polysemous word according to the coverage index to generate an optimized semantic matching score;
[0126] Step S304: Input the optimized semantic matching score into the multimodal fusion layer, segment the conversation text by speaker based on the time sequence markers, and extract the acoustic feature fundamental frequency parameters of each text segment. The acoustic features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate, specifically including:
[0127] Step S3040 , based on the time sequence markers of the dialogue text, the mixed audio is divided into independent speech segments according to the audio track segment identifier of the speaker, and the fundamental frequency of each speech segment is extracted to generate fundamental frequency change data;
[0128] Step S3041: Count the number of sudden changes in the direction and amplitude of the fundamental frequency change within the speech segment based on the first-order difference value of the fundamental frequency change data, and calculate the intonation fluctuation frequency per unit time in combination with the segment duration;
[0129] Step S3042: performing speech endpoint detection on the speech segment to identify the starting position of the semantic demarcation point, dividing the speech segment into semantic unit segments based on the demarcation point, and calculating the average of the silence intervals between adjacent semantic unit segments as the semantic pause interval;
[0130] Step S3043: extracting the total duration of the speech segment and the number of characters in the corresponding text segment, calculating the initial speaking rate based on the ratio of the number of characters to the speech duration, and performing linear normalization processing on the initial speaking rate based on a preset industry standard speaking rate range to generate a speaking rate change rate;
[0131] Step S3044: Bind the intonation fluctuation frequency, semantic pause interval, and speech rate change rate to the time sequence markers of the corresponding text segments to generate structured acoustic feature fundamental frequency parameters;
[0132] Step S305, matching the emotional keywords in each text segment through the emotional dictionary, and calculating the emotional tendency intensity value in combination with the acoustic feature fundamental frequency parameter;
[0133] Step S306: Input the sentiment intensity value and text content into the sentiment calculation module, fuse the acoustic features and text semantics through time alignment technology, and output a smoothed sentiment polarity score sequence, which specifically includes:
[0134] Step S3060: Based on the time sequence mark of the dialogue text and the timestamp of the acoustic feature fundamental frequency parameter, a time sequence mapping relationship between the text paragraphs and the speech segments is established to generate a time-aligned multimodal data index;
[0135] Step S3061: Dynamically segment the text paragraph according to the temporal mapping relationship, extract the speech segmentation time window corresponding to each sentence, and match the intensity weights of the emotional keywords in the corresponding sentence based on the emotional dictionary;
[0136] Step S3062: Multimodal data concatenation is performed on the acoustic feature fundamental frequency parameters and the emotional keyword intensity weights within the same time window, and the fusion features on the time series are scanned by a convolution kernel to extract a cross-modal contextual association feature vector.
[0137] Step S3063: Calculate the sentiment tendency offset based on the distribution density of the cross-modal contextual association feature vector, and generate an initial sentiment polarity score according to the projection direction and modulus of the offset in the preset sentiment dimension coordinate system;
[0138] Step S3063: performing sliding window mean filtering on the initial sentiment polarity scores of the continuous time window, and outputting a smoothed sentiment polarity score sequence;
[0139] Step S307, calculating the variance of the emotional stability index based on the fluctuation amplitude and frequency of the emotional polarity score in the time series;
[0140] Step S308: In the multimodal fusion layer, the semantic matching score is fused with the sentiment polarity score and the emotional stability index at the feature level to generate a joint vector, and the data distribution pattern of the joint vector is analyzed to generate personality trait assessment parameters, which include decision-making tendency, teamwork degree and stress resistance dimensions.
[0141] In an embodiment of the present invention, the job requirement description is analyzed hierarchically:
[0142] Keyword extraction: This method uses regular expressions to match domain dictionaries to extract key elements from text. For example, from the post "Hiring AI algorithm engineers with 5+ years of experience, requiring proficiency in deep learning frameworks (such as TensorFlow), natural language processing, and distributed training optimization capabilities," we extract the skill keywords "deep learning framework" and "natural language processing," the experience requirement "5+ years," and the capability keyword "distributed training optimization."
[0143] Semantic Enhancement: Contextually expand skill keywords based on domain knowledge graphs (such as the IT industry graph). For example, in the graph, this node in the deep learning framework is associated with upstream and downstream nodes such as the "TensorFlow / PyTorch ecosystem," "model training process," and "hardware acceleration adaptation." By traversing these paths, scenario constraints such as "TensorFlow framework optimization experience supporting distributed training" are incorporated into the keyword semantics.
[0144] Vector construction: Map the enhanced keywords to the same dimensional space as the text feature vector (e.g., 500 dimensions), assign initial weights based on keyword importance (core skills weight 0.6-0.8, auxiliary skills weight 0.2-0.4), and generate a job requirement vector.
[0145] Step S301, taking the polysemous word "model" as an example (such as the job requirement "need to optimize the performance of the recommendation model"):
[0146] Semantic network traversal: Locate the "recommendation model" node in the knowledge graph. Its associated paths include upstream and downstream skills such as "collaborative filtering algorithm," "user feature engineering," and "online learning mechanism," as well as industry constraints such as "real-time requirements for e-commerce scenarios" and "cold start strategies for content platforms." Based on path weights (directly associated node weight > indirectly associated node weight), dynamically load the top 5-10 highly relevant co-occurring terms, such as "collaborative filtering," "cold start," and "real-time recommendation." Exclude terms unrelated to the job context (such as "physical model").
[0147] In step S302, assume that the co-occurrence word set contains 8 words, among which "collaborative filtering", "cold start" and "real-time recommendation" appear in the job description, with a coverage of =37.5%. We also analyze the distribution density: if all three words appear in the "Key Skills" section, the density index is 100%. If they are scattered across different sections, the density index is weighted based on the section weights (e.g., 0.7 for the "Key Skills" section and 0.3 for the "Other Requirements" section), which is 0.7× +0.3× ≈0.63.
[0148] In step S303, based on the coverage (37.5%) and density (63%), the vector weight of the polysemous word "model" is adjusted according to preset rules: if the coverage is less than 50% and the density is less than 70%, the weight is reduced by 0.1-0.2 (for example, from the initial 0.5 to 0.4); the cosine similarity (range [-1, 1]) between the job requirement vector and the candidate text feature vector is combined to generate a semantic matching score (for example, 0.72).
[0149] Step S304: Extracting acoustic feature fundamental frequency parameters:
[0150] Speech segmentation and fundamental frequency extraction: Interview recordings are cut into short segments of 1-5 seconds based on the main body of the speech. The fundamental frequency curve is extracted using tools such as Praat. For example, the fundamental frequency range of the candidate's speech segment is 120-180Hz.
[0151] Intonation fluctuation frequency: Calculate the number of zero-crossing points of the first-order difference of the fundamental frequency curve. For example, if it fluctuates 4 times within 5 seconds, the frequency is 0.8 times / second.
[0152] Semantic pause interval: The VAD algorithm is used to detect silent segments (e.g., >200ms) and calculate the average interval between adjacent semantic units. For example, if the three pauses are 300ms, 500ms, and 400ms respectively, the average is 400ms.
[0153] Speech rate change rate: 10 seconds of speech corresponds to 60 words of text, the initial speech rate is 6 words / second, and it is normalized to the industry standard range (4-8 words / second) =0.5.
[0154] Step S305: For the text "The click-through rate of the recommendation system I led increased by 20% after it went online", the sentiment dictionary matches "dominate" (positive +0.6) and "improve" (positive +0.8), combined with the acoustic features (the speech rate is 0.5 and the intonation fluctuation is 0.8 times / second, both higher than the average, and each is weighted 0.1), and the intensity value is ×(1+0.1+0.1)=0.84.
[0155] In step S306, the text sentences and speech segments are aligned by timestamp (e.g., 00:01:00-00:01:10) to create a multimodal index. Acoustic parameters (intonation 0.8, pause 0.4 seconds, speaking rate 0.5) and the emotional keyword intensity (0.84) are concatenated into a feature vector. A 1D convolution is used to extract temporal features (e.g., increasing positive scores over three consecutive windows). A three-window mean filter is applied to the initial score sequence [0.7, 0.8, 0.9, 0.85] to obtain [0.77, 0.85, 0.85].
[0156] Step S307: For the sentiment polarity sequence [0.6, 0.9, 0.7, 0.8], the variance is calculated as =0.0125, reflecting smaller emotional fluctuations.
[0157] Step S308: The semantic matching score (0.72), the sentiment polarity mean (0.8), and the emotional stability variance (0.0125) are fused into a joint vector according to the weights (0.6:0.3:0.1), and generated using the preset mapping rules:
[0158] Decision-making tendency: high match + positive emotions → tendency to make decisive decisions (0.8 points);
[0159] Stress tolerance: High match + low variance → strong stress tolerance (0.9 points).
[0160] By enhancing job requirement vectors through domain semantic networks and optimizing polysemous word representations, we can gain a deeper understanding of job skill requirements, avoid misjudgments caused by superficial keyword matching, and improve the accuracy of talent-to-job matching. Sentiment analysis, combining acoustic features and semantic information from conversational text, not only understands a candidate's professional capabilities but also assesses their personality traits. This comprehensive assessment of candidates from multiple perspectives provides a richer basis for recruitment decisions. Dynamically loading co-occurring word sets and adjusting weights based on coverage allow the system to adapt to changing requirements across different positions and industries, enhancing its versatility and flexibility. Translating abstract factors such as sentiment and personality traits into concrete quantitative indicators makes evaluation results more objective and comparable, reducing human interference.
[0161] In a preferred embodiment of the present invention, the above step S4 sets four orthogonal detection points in the semantic matching result, constructs a quadrilateral according to the semantic coverage, context dependency, term accuracy, and sentiment consistency index corresponding to the detection points, generates a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, and performs dimension calibration on the text feature vector to obtain the calibrated text feature vector, which may include:
[0162] Step S400: setting four orthogonal detection points in the semantic matching results, corresponding to semantic coverage, context dependency, term accuracy, and sentiment consistency index respectively;
[0163] Step S401: constructing the coordinates of the vertices of the quadrilateral according to the index values of the four orthogonal detection points, and generating a dynamic correction coefficient by calculating the vertex distribution characteristic value and the side length correlation coefficient of the quadrilateral, specifically including:
[0164] Step S4010: Map the semantic coverage, context dependency, term accuracy, and sentiment consistency indicators to four vertices in a two-dimensional coordinate system, with semantic coverage and term accuracy serving as the positive and negative components of the horizontal axis, and context dependency and sentiment consistency serving as the positive and negative components of the vertical axis, to generate an initial quadrilateral vertex coordinate set.
[0165] Step S4011, calculating the geometric center coordinates of the quadrilateral based on the vertex coordinate set, and performing relative coordinate transformation on the four vertex coordinates with the center coordinates as the origin to generate normalized vertex offsets based on the center of gravity;
[0166] Step S4012: Calculate the standard deviation based on the distribution dispersion of the normalized vertex offsets to generate a vertex distribution feature value that represents the degree of spatial dispersion of the vertices; calculate the Euclidean distance ratio between adjacent vertices based on the initial quadrilateral vertex coordinate set as the side length proportional coefficient, extract the cosine value of the angle between each side vector, and perform arithmetic averaging on the cosine values to generate a side length correlation coefficient;
[0167] Step S4013: linearly superimpose the vertex distribution characteristic value and the edge length correlation coefficient according to a preset weight coefficient to generate an initial correction factor, and perform interval normalization processing on the initial correction factor to output a dynamic correction coefficient;
[0168] Step S402 , calibrating the dimension of the text feature vector using the dynamic correction coefficient to generate a calibrated text feature vector.
[0169] In this embodiment of the present invention, four orthogonal detection points are defined in the semantic matching results, each detection point corresponding to an evaluation indicator:
[0170] Semantic coverage: measures the semantic match between a candidate's skills and the job requirements (e.g., if a job requirement includes five skills and the candidate covers four, then the match is 80%).
[0171] Contextual dependency: This evaluates the strength of the logical association between a term and its context (e.g., "microservice" has a higher dependency in "SpringCloud microservice architecture" than when it appears in isolation).
[0172] Terminology accuracy: Determine the standardization of the use of professional terms (e.g., whether "AI" accurately corresponds to "artificial intelligence" and not other meanings);
[0173] Emotional consistency: Detects the degree of fit between the emotional tendency of the conversation text and the job requirements (for example, overly emotional expressions in technical job interviews may have low consistency).
[0174] For example, a candidate's semantic coverage is 75%, contextual dependence is 60%, terminology accuracy is 85%, and sentiment consistency is 70%.
[0175] Step S4010: Map the four index values to the four vertices of a two-dimensional coordinate system. The positive and negative directions on the horizontal axis are semantic coverage (positive) and term accuracy (negative), and the positive and negative directions on the vertical axis are context dependency (positive) and sentiment consistency (negative). Assuming the index value range is [0, 1], then:
[0176] Semantic coverage (75%) corresponds to vertex A: (0.75, 0);
[0177] Context dependency (60%) corresponds to vertex B: (0, 0.6);
[0178] The term accuracy (85%) corresponds to vertex C: (-0.85, 0);
[0179] Emotional consistency (70%) corresponds to vertex D: (0, -0.7);
[0180] Generate an initial quadrilateral vertex coordinate set: {A(0.75, 0), B(0, 0.6), C(-0.85, 0), D(0, -0.7)}.
[0181] Step S4011, calculate the geometric center of gravity:
[0182] The barycentric coordinates G(x̄,ȳ) are the mean of the vertex coordinates:
[0183] x̄= =-0.025,ȳ= =-0.025
[0184] Subtract the center of gravity coordinates from each vertex to get the normalized offset:
[0185] A': (0.75+0.025, 0+0.025)=(0.775, 0.025);
[0186] B': (0+0.025, 0.6+0.025)=(0.025, 0.625);
[0187] C': (-0.85+0.025, 0+0.025)=(-0.825, 0.025);
[0188] D': (0+0.025, -0.7+0.025)=(0.025, -0.675).
[0189] Step S4012, vertex distribution eigenvalue (standard deviation):
[0190] Calculate the standard deviation of the normalized offset on the x and y axes to reflect the degree of vertex discreteness:
[0191] x-axis offset: 0.775, 0.025, -0.825, 0.025 → standard deviation ≈ 0.65;
[0192] Y-axis offset: 0.025, 0.625, 0.025, -0.675 → standard deviation ≈ 0.58;
[0193] Take the average value as the vertex distribution eigenvalue: =0.615 (the larger the value, the more dispersed the distribution).
[0194] Side length correlation coefficient:
[0195] Calculate the Euclidean distance between adjacent vertices (e.g., the distance from A' to B' is approximately 0.774, the distance from B' to C' is approximately 1.45, the distance from C' to D' is approximately 0.827, and the distance from D' to A' is approximately 1.45), and obtain the side length proportional coefficient ( ≈ ≈1.87);
[0196] Calculate the cosine of the angle between the side vectors (for example, the cosine of the angle between A'B' and B'C' is approximately -0.98, close to a right angle). After arithmetic averaging, the side length correlation coefficient is approximately 0.92 (the closer the value is to 1, the more regular the quadrilateral).
[0197] In step S4013, the weight of the vertex distribution eigenvalue is preset to 0.6, the weight of the edge length correlation coefficient is preset to 0.4, and the initial correction factor = 0.6×0.615+0.4×0.92=0.731; the correction factor is mapped to the interval [0, 1] (e.g., through linear transformation) to obtain a dynamic correction coefficient approximately equal to 0.78.
[0198] In step S402, a dynamic correction coefficient (0.78) is used to perform weighted adjustment on each dimension of the text feature vector: if a dimension corresponds to a feature related to semantic coverage, its weight is multiplied by the correction coefficient (e.g., original weight 0.3 → 0.3 × 0.78 = 0.234); for dimensions related to term accuracy, the weight adjustment direction is opposite (e.g., original weight 0.2 → 0.2 × (2-0.78) = 0.244), and finally a calibrated vector is generated.
[0199] The quadrilateral model converts abstract indicators into geometric features, which intuitively reflect the balance of semantic matching (for example, the more regular the quadrilateral, the more coordinated the four indicators), avoiding evaluation bias caused by a single indicator being too high or too low. Correction coefficients are generated based on real-time matching results, and the weights of vector dimensions can be automatically adjusted. For example, if the semantic coverage is high but the term accuracy is low (the horizontal axis of the quadrilateral is greatly offset), the coverage weight is reduced and the accuracy weight is increased to make the vector more suitable for job requirements. By calculating the standard deviation and side length ratio of the vertex distribution, outliers (such as a sudden deviation of one indicator from other indicators) can be effectively identified, data fluctuations can be smoothed, and the stability of the feature vector can be improved. The quadrilateral model provides explainability for the evaluation process. Recruiters can quickly locate weak links in the match by observing the vertex distribution (such as low emotional consistency leading to an imbalance in the vertical axis of the quadrilateral), assisting in manual review.
[0200] In a preferred embodiment of the present invention, the above step S5 generates an interactive evaluation report including a dynamic skill topology map, a fitness heat distribution, and a career development prediction curve based on the calibrated text feature vector and personality trait assessment parameters in combination with the job requirement vector, which may include:
[0201] In an embodiment of the present invention, candidate skill nodes (such as "Java", "microservice architecture", "Python") are extracted from the calibrated text feature vector, and the degree of skill mastery is determined based on the vector weight (the higher the weight, the more proficient the mastery); target skill nodes (such as "SpringCloud" "distributed system") are extracted from the job requirement vector, and the upstream and downstream skills in the domain knowledge graph are associated (such as "SpringCloud→Service Registration and Discovery→Eureka").
[0202] Graph structure construction:
[0203] Candidate skill nodes are represented by circles, with their size proportional to their weight (e.g., a "Java" node with a weight of 0.8 has a diameter of 20px, and a "Python" node with a weight of 0.5 has a diameter of 12px);
[0204] Position requirement nodes are represented by blue squares. Requirement nodes that candidates have mastered are outlined with gold (for example, if the "SpringCloud" node is covered, its edge is outlined with gold).
[0205] Edge Rendering:
[0206] Skill-to-skill connections (e.g., "Java → Spring Framework") are represented by dashed lines, with transparency adjusted based on co-occurrence frequency (high-frequency connections have a transparency of 0.8, while low-frequency connections have a transparency of 0.3). Matching edges between candidate skills and job requirements are represented by solid red lines (e.g., "Java" connecting "SpringCloud"). The line width is proportional to the matching score.
[0207] Hovering the mouse over a node displays detailed information (such as skill mastery time and project application examples). Clicking a job requirement node expands the sub-skill tree (e.g., "Distributed Systems → Load Balancing → Nginx") and highlights the sub-nodes that the candidate has mastered. The calibrated text feature vector and the job requirement vector are projected into a two-dimensional space (e.g., using principal component analysis (PCA) for dimensionality reduction). The horizontal axis represents "technical ability match" and the vertical axis represents "soft skill match." Each data point represents a skill or trait dimension (e.g., "algorithmic ability" or "communication ability"), with coordinates determined by the corresponding dimension value in the vector (e.g., if the algorithmic ability match value is 0.7 and the communication ability match value is 0.6, the coordinates are (0.7, 0.6)). A red-yellow-blue gradient is used, with the red area indicating high fit (matching value > 0.8), yellow indicating medium fit (0.5-0.8), and blue indicating low fit (< 0.5). Personality trait parameters (such as decision-making tendency and teamwork) serve as the third dimension and are represented by bubble size (e.g., a teamwork degree of 0.9 corresponds to a bubble radius of 15px, and 0.5 corresponds to 8px).
[0208] Use dotted lines to divide the quadrants:
[0209] The first quadrant (high technology + high soft skills): core advantage area; the fourth quadrant (high technology + low soft skills): area where potential and risk coexist.
[0210] Drag the slider to switch the display dimension (for example, from "Technical Ability" to "Industry Experience"); click a hot area to view a list of specific skills or traits included in that area and matching details.
[0211] Generation of career development prediction curve:
[0212] Candidate data: calibrated skill vectors, years of work experience, and project experience complexity (extracted from descriptions such as "responsible for a system with XX million users" in the text);
[0213] Job trend data: Obtain the changing trends in skill requirements for target positions through industry reports (e.g., the annual growth rate of "large model training" skill requirements for "AI algorithm engineer" positions is 30%).
[0214] Personality trait parameters: stress tolerance, learning ability, etc. are used to adjust the prediction slope (for example, candidates with strong stress tolerance can accelerate their skill improvement by 10%-20%).
[0215] Skill growth curve:
[0216] Base growth rate: set based on the industry average (e.g., the "distributed systems" skill will naturally grow by 5% per year);
[0217] Acceleration factor: For every 1 unit increase in a candidate's learning ability (e.g., from 0.6 to 0.7), the growth rate increases by 2%;
[0218] Generate a skill mastery curve for the next 1-3 years (e.g., if the current "large model training" mastery is 0.4, it is predicted to be 0.6 in 1 year and 0.8 in 2 years).
[0219] Adaptability fluctuation curve:
[0220] By combining job demand trends (for example, when the demand for a certain skill decreases, the fitness curve drops) and candidate skill growth, a dynamic fitness curve is generated (for example, if the current total fitness is 0.7, it is predicted to drop to 0.65 in one year due to changes in job requirements, but it will rebound to 0.78 as the candidate's skills improve).
[0221] The orange curve represents the skill growth trend, and the blue curve represents the change in fitness. The shaded area represents the predicted confidence interval (e.g., skill mastery fluctuates by ±5% at a 95% confidence level). Key time points are marked (e.g., "XX new skill needs to be mastered in 6 months") and associated training resource recommendations are provided.
[0222] like Figure 2 As shown, the embodiment of the present invention also provides an AI-based talent evaluation and management system, including:
[0223] The data collection module is used to extract multi-source resume information through OCR recognition, speech transcription, and format parsing, and standardize it into a text dataset with semantic labels;
[0224] The semantic encoding module is used to perform contextual semantic encoding and domain knowledge graph disambiguation on standardized text datasets to generate text feature vectors containing entity association relationships;
[0225] Dynamic semantic matching module, used to dynamically match text feature vectors with job requirement vectors and analyze conversation sentiment to generate personality trait assessment parameters;
[0226] A dynamic correction module is used to construct an evaluation quadrilateral by setting orthogonal detection points, calculate eigenvalues and correlation coefficients to generate correction coefficients, and generate calibrated eigenvectors;
[0227] The evaluation report module is used to generate an interactive evaluation report based on the calibrated feature vectors, personality parameters and job requirement vectors.
[0228] It should be noted that this system is a system corresponding to the above method, and all implementation methods in the above method embodiment are applicable to this embodiment and can achieve the same technical effects.
[0229] The embodiment of the present invention further provides a computer-readable storage medium storing instructions, which, when executed on a computer, causes the computer to execute the above-described method. All implementations in the above-described method embodiment are applicable to this embodiment and can achieve the same technical effects.
[0230] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.< / w:outlinelvl> < / w:tbl> < / w:p>
Claims
1. The talent evaluation and management method based on AI intelligence is characterized by: The method comprises: Step S1: Convert the handwritten resume image into structured text data through OCR image recognition; convert the interview recording into conversation text with time sequence tags through speech transcription; parse the electronic document format to extract the original text content; encode and standardize the structured text data, conversation text, and original text content to generate a standardized text dataset containing semantic tags; Step S2: contextual semantic encoding is performed on the standardized text dataset, and node alignment and semantic disambiguation processing of professional terms are performed through the domain knowledge graph to generate a text feature vector containing entity association relationships; Step S3: Dynamically semantically match the text feature vector with the job requirement vector, optimize the vector representation of polysemous words in the job context, analyze the emotional tendency characteristics of the dialogue text, and generate personality trait assessment parameters; Step S4: Set four orthogonal detection points in the semantic matching result, construct a quadrilateral based on the semantic coverage, context dependency, term accuracy, and sentiment consistency index corresponding to the detection points, generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, and perform dimension calibration on the text feature vector to obtain a calibrated text feature vector; Step S5: Based on the calibrated text feature vector and personality trait assessment parameters, combined with the job requirement vector, an interactive evaluation report including a dynamic skill topology map, fitness heat distribution, and career development prediction curve is generated.
2. The talent evaluation and management method based on AI intelligence according to claim 1 is characterized in that: Convert handwritten resume images into structured text data through OCR image recognition; convert interview recordings into time-stamped conversation text through speech transcription; and parse electronic documents to extract original text content. The structured text data, conversation text, and original text content are encoded and standardized to generate a standardized text dataset with semantic labels, including: When performing OCR image recognition on handwritten resume images, an adaptive threshold segmentation algorithm is used to perform illumination equalization on the image, segment the candidate text area, and parse the text area. The recognition results are mapped into structured field data according to the resume field classification. Among them, the education background field is parsed into the name of the school, major and degree level, the work experience field is parsed into the company, position and job description, and the skill certificate field is parsed into the certificate name and certification body. When performing voice transcription on interview recordings, different speakers are identified based on voiceprint features. During the recognition process, the mixed audio is separated into independent audio tracks and combined with the voice end Point detection transcribes each audio track in segments, generating a time-stamped dialogue text. The dialogue text is annotated with speech emotional fundamental frequency features, including intonation fluctuation frequency, semantic pause intervals, and speech rate change rate. These emotional fundamental frequency features are extracted using acoustic parameters and bound to text paragraphs. When formatting electronic documents, a preset template engine is used based on the document type to identify the hierarchical structure of tables, paragraphs, and headings in the document, extracting the unformatted original text content and mapping the text content to rows and columns. The association between merged cells across columns is achieved through contextual semantic reconstruction. The structured field data, conversation text with time tags and original text content are input into the coding standardization module. Semantic tags are added to skill names and project experience keywords through the preset industry terminology dictionary, and the time format is unified into a standard format and the place names are normalized into the full names of administrative divisions. A standardized text dataset containing semantic tags is generated. The dataset is stored in a structured document format according to the field type, and each field is associated with a semantic tag and a normalized attribute value.
3. The talent evaluation and management method based on AI intelligence according to claim 2 is characterized in that: Perform contextual semantic encoding on standardized text datasets, perform node alignment and semantic disambiguation on professional terms through domain knowledge graphs, and generate text feature vectors containing entity association relationships, including: Based on the semantic tags and structured document formats in the standardized text dataset, contextual semantic encoding technology is used to perform semantic analysis on the text and generate initial semantic vectors; The initial semantic vector is input into the domain structured semantic network for term alignment. By traversing the professional term nodes in the semantic network, the semantic similarity between the terms in the text and the nodes is calculated. Terms with similarity above a preset threshold are mapped to corresponding nodes. The contextual semantics of the terms are then expanded based on the association paths between the nodes. The expanded contextual semantics of the terms include upstream and downstream skill associations, industry scenario constraints, and job capability dependencies. Disambiguation weights are constructed using the co-occurrence relationships and hierarchical classification information of nodes in the semantic network. The probability distribution of ambiguous terms is calculated based on the context window of the terms in the text, and the vector representation of the terms is dynamically adjusted to eliminate semantic ambiguity and obtain disambiguation results. The aligned term node features, extended semantics and disambiguation results are integrated, and the attribute features of the nodes in the semantic network are concatenated with the text semantic vector to generate a text feature vector containing entity association relationships.
4. The talent evaluation and management method based on AI intelligence according to claim 3 is characterized in that: Dynamically match text feature vectors with job requirement vectors, optimize the vector representation of polysemous words in the job context, analyze the emotional tendency characteristics of the conversation text, and generate personality trait assessment parameters, including: Based on the job description, skill requirements, years of experience, and ability keywords are extracted to construct a job requirement vector. The skill requirement keywords are semantically enhanced based on the node attributes in the domain structured semantic network. Semantic enhancement extracts upstream and downstream skill dependencies and industry scenario constraints by traversing the associated paths of the nodes in the semantic network, generating a job requirement vector consistent with the dimensions of the text feature vector. When expanding the co-occurrence distribution of polysemous words through the association paths in the semantic network, the co-occurrence word set related to job requirements is dynamically loaded based on the upstream and downstream skill dependencies between nodes and industry scenario constraints; Count the frequency and distribution density of co-occurring words in job requirements descriptions to generate coverage indicators; Adjust the vector weights of polysemous words based on the coverage index to generate an optimized semantic matching score; The optimized semantic matching scores are input into the multimodal fusion layer, and the conversation text is segmented by speaker based on time sequence markers. The acoustic feature fundamental frequency parameters of each text segment are extracted. The acoustic features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate. The emotional keywords in each text are matched through the emotional dictionary, and the emotional tendency intensity value is calculated by combining the acoustic feature fundamental frequency parameters; The sentiment intensity value and text content are input into the sentiment calculation module, which fuses the acoustic features and text semantics through time series alignment technology to output a smoothed sentiment polarity score sequence. Calculate the variance of the emotional stability index based on the fluctuation amplitude and frequency of the emotional polarity score in the time series; In the multimodal fusion layer, the semantic matching score is fused with the sentiment polarity score and the emotional stability index at the feature level to generate a joint vector. The data distribution pattern of the joint vector is then analyzed to generate personality trait assessment parameters, including decision-making tendency, teamwork degree, and stress tolerance dimensions.
5. The talent evaluation and management method based on AI intelligence according to claim 4 is characterized in that: The optimized semantic matching score is input into the multimodal fusion layer, and the conversation text is segmented by speaker based on the time sequence markers. The acoustic feature fundamental frequency parameters of each text segment are extracted. The acoustic features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate, including: Based on the time sequence mark of the dialogue text, the mixed audio is cut into independent speech segments according to the audio track segment identifier of the speaker, and the fundamental frequency of each speech segment is extracted to generate fundamental frequency change data; According to the first-order difference value of the fundamental frequency change data, the number of sudden changes in the direction and amplitude of the fundamental frequency change in the speech segment is counted, and the frequency of intonation fluctuation per unit time is calculated based on the segment duration; Perform speech endpoint detection on the speech segment, identify the starting position of the semantic demarcation point, divide the speech segment into semantic unit segments based on the demarcation point, and calculate the average of the silence intervals between adjacent semantic unit segments as the semantic pause interval; Extract the total duration of the speech segment and the number of characters in the corresponding text segment, calculate the initial speaking rate based on the ratio of the number of characters to the speech duration, and perform linear normalization on the initial speaking rate based on the preset industry standard speaking rate range to generate the speaking rate change rate; The intonation fluctuation frequency, semantic pause interval and speech rate change rate are bound to the timing markers of the corresponding text segments to generate the fundamental frequency parameters of structured acoustic features.
6. The talent evaluation and management method based on AI intelligence according to claim 4 is characterized in that: The sentiment intensity value and text content are input into the sentiment calculation module. The acoustic features and text semantics are integrated through time series alignment technology to output a smoothed sentiment polarity score sequence, including: Based on the time sequence markings of the conversation text and the timestamps of the acoustic feature fundamental frequency parameters, a time sequence mapping relationship between text paragraphs and speech segments is established to generate a time-aligned multimodal data index; Dynamically segment text paragraphs based on temporal mapping relationships, extract the speech segmentation time window corresponding to each sentence, and match the intensity weights of emotional keywords in the corresponding sentences based on the emotional dictionary; The acoustic feature fundamental frequency parameters and emotional keyword intensity weights within the same time window are concatenated for multimodal data. The fused features on the time series are scanned by the convolution kernel to extract the cross-modal contextual association feature vector. The emotional tendency offset is calculated based on the distribution density of the cross-modal contextual association feature vector, and the initial emotional polarity score is generated according to the projection direction and modulus of the offset in the preset emotional dimension coordinate system; The initial sentiment polarity scores of the continuous time window are subjected to sliding window mean filtering, and a smoothed sentiment polarity score sequence is output.
7. The talent evaluation and management method based on AI intelligence according to claim 6 is characterized in that: Four orthogonal detection points are set in the semantic matching results. A quadrilateral is constructed based on the semantic coverage, context dependency, term accuracy, and sentiment consistency indicators corresponding to the detection points. A dynamic correction coefficient is generated by calculating the vertex distribution eigenvalues and side length correlation coefficients of the quadrilateral. The dimension of the text feature vector is calibrated to obtain the calibrated text feature vector, including: Four orthogonal detection points are set in the semantic matching results, corresponding to semantic coverage, context dependency, term accuracy and sentiment consistency indicators respectively; The coordinates of the quadrilateral vertices are constructed based on the index values of the four orthogonal detection points, and the dynamic correction coefficient is generated by calculating the vertex distribution characteristic value and the side length correlation coefficient of the quadrilateral; The dynamic correction coefficient is used to calibrate the dimension of the text feature vector to generate a calibrated text feature vector.
8. The talent evaluation and management method based on AI intelligence according to claim 7 is characterized in that: The coordinates of the quadrilateral vertices are constructed based on the index values of the four detection points. The dynamic correction coefficients are generated by calculating the vertex distribution characteristic values and side length correlation coefficients of the quadrilateral, including: Map the semantic coverage, context dependency, term accuracy, and sentiment consistency indicators to four vertices in a two-dimensional coordinate system, with semantic coverage and term accuracy as the positive and negative components of the horizontal axis, and context dependency and sentiment consistency as the positive and negative components of the vertical axis, to generate an initial quadrilateral vertex coordinate set. Calculate the geometric center of gravity coordinates of the quadrilateral based on the vertex coordinate set, and perform relative coordinate transformation on the four vertex coordinates with the center of gravity coordinates as the origin to generate normalized vertex offsets based on the center of gravity; The standard deviation is calculated based on the distribution dispersion of the normalized vertex offsets to generate vertex distribution feature values that represent the degree of spatial dispersion of the vertices. Based on the initial quadrilateral vertex coordinate set, the Euclidean distance ratio between adjacent vertices is calculated as the side length proportional coefficient, and the cosine value of the angle between each side vector is extracted. The cosine values are arithmetic averaged to generate the side length correlation coefficient. The vertex distribution eigenvalues and the edge length correlation coefficients are linearly superimposed according to the preset weight coefficients to generate the initial correction factor, which is then interval-normalized to output the dynamic correction coefficient.
9. An AI-based talent evaluation and management system, the system implementing the method according to any one of claims 1 to 8, characterized in that: include: The data collection module is used to extract multi-source resume information through OCR recognition, speech transcription, and format parsing, and standardize it into a text dataset with semantic labels; The semantic encoding module is used to perform contextual semantic encoding and domain knowledge graph disambiguation on standardized text datasets to generate text feature vectors containing entity association relationships; Dynamic semantic matching module, used to dynamically match text feature vectors with job requirement vectors and analyze conversation sentiment to generate personality trait assessment parameters; A dynamic correction module is used to construct an evaluation quadrilateral by setting orthogonal detection points, calculate eigenvalues and correlation coefficients to generate correction coefficients, and generate calibrated eigenvectors; The evaluation report module is used to generate an interactive evaluation report based on the calibrated feature vectors, personality parameters and job requirement vectors.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, which implements the method according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Patent Citations
Management method and system based on man-post intelligent matching algorithm model
CN119204834A
Intelligent talent tag portrait analysis system based on big data
CN120087927A