Talent evaluation management method and system based on AI intelligence
Through OCR and speech transfer technology, unstructured data is converted into structured text, and the domain knowledge graph is used for semantic coding and disambiguation processing, solving the problems of information omissions and misjudgment in talent assessment by enterprises, and achieving accurate candidate portraits and evaluations.
Patent Information
- Application Number
- CN202510757971.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-09
AI Technical Summary
In the prior art, it is difficult for enterprises to effectively process unstructured data in talent assessment, such as handwritten resumes and interview recordings, resulting in inefficient information extraction and easy to miss key content. Traditional methods lack correlation analysis of context semantics, resulting in misjudgment.
Unstructured data is converted into structured text through OCR image recognition and speech transcription, semantic coding and disambiguation are combined with domain knowledge graphs, text feature vectors containing entity association relationships are generated, dynamic semantic matching and sentiment analysis are performed, and personality trait evaluation parameters are generated.
Accurate assessment of candidate skills and personality traits is achieved, misjudgment is avoided, the accuracy and reliability of the assessment is improved, and interactive evaluation reports are provided to support decision-making.
Smart Images

Figure CN120338741A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence analysis, and particularly to a talent evaluation management method and system based on AI intelligence. Background Art
[0002] In the prior art, most enterprises rely on manual screening of paper resumes, manual recording of interview contents, or electronic tools based on keyword matching for talent evaluation, which have the following significant defects: Unstructured data such as handwritten resumes and interview recordings are difficult to be standardized, resulting in low information extraction efficiency and easy omission of key content. The traditional OCR technology has a low recognition accuracy for complex handwritten characters, the speech-to-text conversion lacks the associated analysis of context semantics, and the electronic document parsing often leads to text fragmentation due to format differences; there are common polysemous words, industry terms and abbreviations in resume and interview texts. The traditional method relies on manual experience for semantic judgment, and is prone to misjudgment of job suitability due to context understanding deviation. The existing knowledge graph applications are mostly limited to static term matching and lack the ability of dynamic disambiguation and cross-domain node alignment. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a talent evaluation management method and system based on AI intelligence, and realize accurate talent evaluation through dynamic semantic matching and multi-dimensional detection point calibration.
[0004] To solve the above technical problem, the technical solution of the present invention is as follows: In the first aspect, a talent evaluation management method based on AI intelligence, the method includes: Step S1, converting the handwritten resume image into structured text data through OCR image recognition; converting the interview recording into dialogue text with time series marks through speech-to-text conversion; parsing the format of the electronic document to extract the original text content; performing encoding and standardization processing on the structured text data, dialogue text and original text content to generate a standardized text data set containing semantic tags; Step S2, performing context semantic encoding on the standardized text data set, performing node alignment and semantic disambiguation processing on professional terms through a domain knowledge graph, and generating a text feature vector containing entity association relationships; Step S3, performing dynamic semantic matching on the text feature vector and the job requirement vector, optimizing the vector representation of polysemous words in the job context, analyzing the emotional tendency characteristics of the dialogue text, and generating personality trait evaluation parameters; Step S4, set four orthogonal detection points in the semantic matching results, construct a quadrilateral based on the semantic coverage, context dependence, term accuracy, and sentiment consistency indicators corresponding to the detection points, generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, and perform dimensional calibration on the text feature vector to obtain a calibrated text feature vector; Step S5, based on the calibrated text feature vector and personality trait evaluation parameters, generate an interactive evaluation report including a dynamic skill topology map, a fitness heat distribution, and a career development prediction curve in combination with the job requirement vector.
[0005] Further, convert the handwritten resume image into structured text data through OCR image recognition; convert the interview recording into dialogue text with time sequence tags through speech transcription; perform format parsing on the electronic document to extract the original text content; perform encoding standardization processing on the structured text data, dialogue text, and original text content to generate a standardized text data set containing semantic tags, including: When performing OCR image recognition on the handwritten resume image, use an adaptive threshold segmentation algorithm to perform illumination equalization processing on the image, segment the candidate text area, and parse the text area. The recognition result is classified and mapped to structured field data according to the resume fields. Among them, the education background field is parsed into the institution name, major, and degree level, the work experience field is parsed into the company name, position, and responsibility description, and the skill certificate field is parsed into the certificate name and the certification institution; when performing speech transcription on the interview recording, different speakers are identified based on voiceprint features. During the recognition process, the mixed audio is separated into independent audio tracks, and each audio track is segmented and transcribed in combination with voice endpoint detection to generate dialogue text with time sequence tags, and the voice emotion fundamental frequency features are marked in the dialogue text. The emotion fundamental frequency features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate, and the emotion fundamental frequency features are extracted through acoustic parameters and bound to the text paragraph; when performing format parsing on the electronic document, use a preset template engine according to the document type to identify the table, paragraph, and title hierarchical structure in the document, extract the unformatted original text content, and perform row-column association mapping on the text content. The association relationship of the merged cells across columns is realized through context semantic reconstruction; Input the structured field data, dialogue text with time sequence tags, and original text content into the encoding standardization module, add semantic tags to the skill names and project experience keywords through a preset industry term dictionary, unify the time format into the standard format, and normalize the location names into the full name of the administrative division to generate a standardized text data set containing semantic tags. The data set is stored in a structured document format according to the field type, and each field is associated with a semantic tag and a normalized attribute value.
[0006] Further, perform context semantic encoding on the standardized text dataset, perform node alignment and semantic disambiguation processing on professional terms through the domain knowledge graph, and generate a text feature vector containing entity association relationships, including: Based on the semantic tags and structured document format in the standardized text dataset, use context semantic encoding technology to perform semantic analysis on the text and generate an initial semantic vector; Input the initial semantic vector into the domain structured semantic network for term alignment. By traversing the professional term nodes in the semantic network, calculate the semantic similarity between the terms in the text and the nodes, map the terms with similarity higher than the preset threshold to the corresponding nodes, and expand the context semantics of the terms based on the association paths between the nodes. The expanded context semantics of the terms include upstream and downstream skill associations, industry scenario constraint conditions, and job ability dependency relationships; Construct a disambiguation weight using the co-occurrence relationship and hierarchical classification information of the nodes in the semantic network, and calculate the probability distribution of the ambiguous terms based on the context window of the terms in the text, and dynamically adjust the vector representation of the terms to eliminate semantic ambiguity and obtain the disambiguation result; Fuse the aligned term node features, expanded semantics, and disambiguation results, splice the attribute features of the nodes in the semantic network with the text semantic vector, and generate a text feature vector containing entity association relationships.
[0007] Further, perform dynamic semantic matching between the text feature vector and the job requirement vector, optimize the vector representation of the polysemous words in the job context, analyze the emotional tendency characteristics of the dialogue text, and generate personality trait evaluation parameters, including: Extract skill requirements, years of experience, and ability keywords according to the job requirement description to construct a job requirement vector. Among them, the skill requirement keywords are semantically enhanced based on the node attributes in the domain structured semantic network. The semantic enhancement extracts the upstream and downstream skill dependency relationships and industry scenario constraint conditions by traversing the association paths of the nodes in the semantic network, and generates a job requirement vector with the same dimension as the text feature vector; When expanding the co-occurrence word distribution of polysemous words through the association paths in the semantic network, dynamically load the co-occurrence word set related to the job requirements based on the upstream and downstream skill dependency relationships and industry scenario constraint conditions between the nodes; Count the occurrence frequency and distribution density of the co-occurrence words in the job requirement description to generate a coverage rate index; Adjust the vector weight of the polysemous words according to the coverage rate index to generate an optimized semantic matching score; Input the optimized semantic matching score into the multi-modal fusion layer, and segment the dialogue text by the speaker based on the time series markers, and extract the acoustic feature fundamental frequency parameters of each segment of the text. The acoustic features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate; Match the sentiment keywords in each text segment through a sentiment dictionary, and calculate the sentiment tendency intensity value by combining the fundamental frequency parameter of the acoustic features; Input the sentiment tendency intensity value and the text content into the sentiment calculation module, and fuse the acoustic features and text semantics through the time series alignment technology to output a smoothed sequence of sentiment polarity scores; Calculate the variance of the emotion stability index according to the fluctuation amplitude and frequency of the sentiment polarity scores in the time series; In the multi-modal fusion layer, perform feature-level fusion on the semantic matching score, the sentiment polarity score, and the emotion stability index to generate a joint vector, and analyze the data distribution pattern of the joint vector to generate personality trait evaluation parameters, where the parameters include dimensions of decision-making tendency, teamwork degree, and stress resistance ability.
[0008] Furthermore, input the optimized semantic matching score into the multi-modal fusion layer, segment the dialogue text by speaker based on the time series markers, and extract the fundamental frequency parameters of the acoustic features of each text segment. The acoustic features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate, including: Based on the time series markers of the dialogue text, cut the mixed audio into independent speech segments according to the speaker's track segmentation identification, perform fundamental frequency extraction on each speech segment to generate fundamental frequency change data; According to the first-order difference value of the fundamental frequency change data, count the number of mutations in the fundamental frequency change direction and amplitude within the speech segment, and calculate the intonation fluctuation frequency per unit time in combination with the segment duration; Perform voice endpoint detection on the speech segment, identify the starting position of the semantic boundary point, divide the speech segment into semantic unit segments based on the boundary point, and calculate the average silence interval between adjacent semantic unit segments as the semantic pause interval; Extract the total duration of the speech segment and the number of characters in the corresponding text segment, calculate the initial speech rate according to the ratio of the number of characters to the speech duration, and perform linear normalization on the initial speech rate based on the preset industry standard speech rate range to generate the speech rate change rate; Bind the intonation fluctuation frequency, semantic pause interval, and speech rate change rate with the time series markers of the corresponding text segment to generate structured acoustic feature fundamental frequency parameters.
[0009] Furthermore, input the sentiment tendency intensity value and the text content into the sentiment calculation module, and fuse the acoustic features and text semantics through the time series alignment technology to output a smoothed sequence of sentiment polarity scores, including: Based on the time series markers of the dialogue text and the timestamps of the acoustic feature fundamental frequency parameters, establish a time series mapping relationship between the text paragraphs and the speech segments to generate a time-aligned multi-modal data index; Dynamically clause-segment the text paragraph according to the timing mapping relationship, extract the voice segmentation time window corresponding to each sentence, and match the intensity weight of the sentiment keyword in the corresponding sentence based on the sentiment dictionary; Perform multimodal data splicing on the fundamental frequency parameter of the acoustic feature within the same time window and the intensity weight of the sentiment keyword, and extract the cross-modal context correlation feature vector by scanning the fused features on the time series with a convolutional kernel; Calculate the sentiment tendency offset based on the distribution density of the cross-modal context correlation feature vector, and generate an initial sentiment polarity score according to the projection direction and modulus length of the offset in the preset sentiment dimension coordinate system; Perform sliding window mean filtering on the initial sentiment polarity scores of consecutive time windows, and output the smoothed sentiment polarity score sequence.
[0010] Furthermore, set four orthogonal detection points in the semantic matching result, construct a quadrilateral according to the semantic coverage, context dependence, term accuracy, and sentiment consistency indicators corresponding to the detection points, generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, and perform dimension calibration on the text feature vector to obtain the calibrated text feature vector, including: Set four orthogonal detection points in the semantic matching result, corresponding to the semantic coverage, context dependence, term accuracy, and sentiment consistency indicators respectively; Construct the vertex coordinates of the quadrilateral according to the index values of the four orthogonal detection points, and generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral; Use the dynamic correction coefficient to perform dimension calibration on the text feature vector to generate the calibrated text feature vector.
[0011] Furthermore, construct the vertex coordinates of the quadrilateral according to the index values of the four detection points, and generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, including: Map the semantic coverage, context dependence, term accuracy, and sentiment consistency indicators to the four vertices in the two-dimensional coordinate system respectively. Among them, the semantic coverage and term accuracy are used as the positive and negative direction components of the horizontal axis, and the context dependence and sentiment consistency are used as the positive and negative direction components of the vertical axis to generate the initial quadrilateral vertex coordinate set; Calculate the geometric centroid coordinates of the quadrilateral based on the vertex coordinate set, and perform relative coordinate transformation on the four vertex coordinates with the centroid coordinates as the origin to generate the normalized vertex offset based on the centroid; Calculate the standard deviation based on the distribution dispersion of the normalized vertex offsets to generate a vertex distribution eigenvalue characterizing the vertex space dispersion; based on the set of initial quadrilateral vertex coordinates, calculate the Euclidean distance ratio between adjacent vertices as the side length ratio coefficient, extract the cosine values of the included angles of each side vector, and perform arithmetic averaging on the cosine values to generate a side length correlation coefficient. Linearly superimpose the vertex distribution eigenvalue and the side length correlation coefficient according to a preset weight coefficient to generate an initial correction factor, and perform interval normalization processing on the initial correction factor to output a dynamic correction coefficient.
[0012] In a second aspect, a talent evaluation and management system based on AI intelligence includes: A data acquisition module for extracting multi-source resume information through OCR recognition, speech transcription, and format parsing and standardizing it into a text data set with semantic tags. A semantic encoding module for performing context semantic encoding and domain knowledge graph disambiguation on the standardized text data set to generate a text feature vector with entity association relationships. A dynamic semantic matching module for dynamically matching the text feature vector with the job requirement vector and analyzing the dialogue sentiment tendency to generate personality trait evaluation parameters. A dynamic correction module for constructing an evaluation quadrilateral by setting orthogonal detection points, calculating eigenvalues and correlation coefficients to generate a correction coefficient, and generating a calibrated feature vector. An evaluation report module for generating an interactive evaluation report based on the calibrated feature vector, personality parameters, and job requirement vector.
[0013] In a third aspect, a computer-readable storage medium stores a program, and when the program is executed by a processor, the method described above is implemented.
[0014] The above solution of the present invention has at least the following beneficial effects: With the help of technologies such as OCR image recognition and speech transcription, different forms of data such as handwritten resumes, interview recordings, and electronic documents are uniformly converted into standardized text, breaking data barriers and avoiding information omission; at the same time, data structuring is achieved through semantic tag annotation. Using context semantic encoding and domain knowledge graphs, deeply mine text semantic information, accurately align professional terms and eliminate ambiguities, and construct a text feature vector containing entity association relationships, which can accurately capture the candidate's skill connotations and association relationships and avoid misjudgments caused by surface keyword matching. Realize the dynamic semantic matching of text features and job requirements, adapt to context for polysemous words, and at the same time generate personality trait parameters in combination with dialogue text sentiment analysis, comprehensively evaluate talents from multiple dimensions such as professional skills, semantic understanding, and emotional traits, and comprehensively depict the candidate portrait.
[0015] Construct a quadrilateral through four orthogonal detection points, and examine the balance of the matching results from dimensions such as semantic coverage and context dependence; generate a dynamic correction coefficient by calculating the geometric features of the quadrilateral, calibrate the dimensions of the text feature vector to ensure that the evaluation results are more in line with the actual job requirements, and improve the accuracy and reliability of the evaluation. Present the calibrated evaluation results in visual forms such as a dynamic skill topology map, a fitness heat distribution, and a career development prediction curve, and support interactive operations to help decision-makers quickly locate the advantages and disadvantages of candidates. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic flowchart of a talent evaluation management method based on AI intelligence provided by an embodiment of the present invention.
[0017] Figure 2 is a schematic diagram of a talent evaluation management system based on AI intelligence provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0019] As Figure 1 shown, an embodiment of the present invention proposes a talent evaluation management method based on AI intelligence, and the method includes the following steps: Step S1, convert the handwritten resume image into structured text data through OCR image recognition; convert the interview recording into dialogue text with time sequence marks through speech transcription; perform format parsing on the electronic document to extract the original text content; perform encoding standardization processing on the structured text data, dialogue text, and original text content to generate a standardized text data set containing semantic tags; Step S2, perform context semantic encoding on the standardized text data set, perform node alignment and semantic disambiguation processing on professional terms through a domain knowledge graph, and generate a text feature vector containing entity association relationships; Step S3, perform dynamic semantic matching between the text feature vector and the job requirement vector, optimize the vector representation of polysemous words in the job context, analyze the emotional tendency characteristics of the dialogue text, and generate personality trait evaluation parameters; Step S4, set four orthogonal detection points in the semantic matching results, construct a quadrilateral according to the semantic coverage, context dependence, term accuracy, and sentiment consistency indexes corresponding to the detection points, generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, and perform dimensional calibration on the text feature vector to obtain a calibrated text feature vector; Step S5, based on the calibrated text feature vector and the personality trait evaluation parameters, generate an interactive evaluation report including a dynamic skill topology map, a fitness heat distribution, and a career development prediction curve in combination with the job requirement vector.
[0020] In the embodiment of the present invention, by means of context semantic encoding and a domain knowledge graph, the implicit semantics behind the text are deeply mined, professional terms are accurately aligned and ambiguity is eliminated, and a text feature vector including skill association relationships is constructed, which can accurately capture the connotation and extension of the candidate's skills, avoid misjudgment caused by relying only on surface keywords, and make the talent ability evaluation more in-depth and accurate. Realize the dynamic semantic matching between the text feature vector and the job requirement vector, perform context adaptation for polysemous words, and ensure the accuracy of skill matching; at the same time, by analyzing the sentiment tendency and acoustic features of the dialogue text, the personality trait parameters of the candidate are quantified, filling the gap that soft skills are difficult to objectively evaluate in traditional talent evaluation, and comprehensively depicting the candidate portrait from multiple dimensions such as professional skills, semantic understanding, and emotional traits.
[0021] By setting four orthogonal detection points of semantic coverage, context dependence, term accuracy, and sentiment consistency, construct a quadrilateral, and test the balance of the semantic matching results from multiple angles, effectively avoiding evaluation biases caused by too high or too low single-sided indicators; generate a dynamic correction coefficient based on the geometric features of the quadrilateral, perform dimensional calibration on the text feature vector, automatically adjust the weights of each dimension, make the evaluation result more in line with the actual job requirements, and enhance the rationality and reliability of the evaluation. Decision-makers can quickly locate the advantages and disadvantages of candidates. At the same time, the career development prediction curve combines industry trends and candidate traits to provide a forward-looking reference for enterprises to formulate talent plans and for candidates to provide career development directions.
[0022] In a preferred embodiment of the present invention, in the above step S1, the handwritten resume image is converted into structured text data through OCR image recognition; the interview recording is converted into a dialogue text with time series marks through speech transcription; the original text content is extracted by parsing the format of the electronic document; the structured text data, the dialogue text, and the original text content are subjected to encoding standardization processing to generate a standardized text data set including semantic tags, which may include: In step S100, when performing OCR image recognition on a handwritten resume image, an adaptive threshold segmentation algorithm is used to perform illumination equalization processing on the image, segment candidate text regions, and parse the text regions. The recognition results are classified and mapped into structured field data according to resume fields. Among them, the education background field is parsed into the name of the institution, major, and degree level; the work experience field is parsed into the name of the employer, position, and responsibility description; the skill certificate field is parsed into the name of the certificate and the certifying agency. When performing speech transcription on the interview recording, different speaking subjects are identified based on voiceprint features. During the recognition process, the mixed audio is separated into independent tracks, and each track is segmented and transcribed in combination with voice endpoint detection to generate a dialogue text with time sequence markers, and the fundamental frequency features of speech emotion are marked in the dialogue text. The emotion fundamental frequency features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate, and the emotion fundamental frequency features are extracted through acoustic parameters and bound to the text paragraphs. When performing format parsing on the electronic document, a preset template engine is used according to the document type to identify the table, paragraph, and title hierarchical structure in the document, extract the original text content without formatting, and perform row-column association mapping on the text content. Among them, the association relationship of merged cells across columns is realized through context semantic reconstruction. In step S101, the structured field data, the dialogue text with time sequence markers, and the original text content are input into the encoding and standardization module. Semantic tags are added to the skill names and project experience keywords through a preset industry term dictionary, the time format is unified into the standard format, and the location names are normalized to the full name of the administrative division to generate a standardized text data set containing semantic tags. The data set is stored in a structured document format according to the field type, and each field is associated with a semantic tag and a normalized attribute value.
[0023] In the embodiment of the present invention, the specific logic of the adaptive threshold segmentation algorithm is as follows: The image is divided into multiple local sub-regions (such as an 8x8 pixel grid), the mean and variance of the pixel gray values are calculated for each sub-region, and the threshold is dynamically adjusted according to the light and dark characteristics of the sub-region (for example: the threshold is lowered in the dark region to highlight the text, and the threshold is raised in the bright region to suppress background noise). By processing each sub-region one by one, the influence of global illumination unevenness (such as local shadows or reflections in the resume scan) is eliminated, and the contrast between the text and the background is significantly improved.
[0024] The equalized image identifies the text contour through edge detection (such as the Canny operator), and combines morphological operations (dilation, erosion) to merge adjacent edges to form continuous rectangular candidate regions. Regions with too small area (such as less than 10x10 pixels) or abnormal aspect ratio (such as the height is much larger than the width, which may be a table line) are filtered out, and only regions that conform to the text arrangement characteristics (such as moderate height and width that can accommodate multiple characters) are retained.
[0025] Field parsing rules: Educational Background: When keywords such as "university", "college", "undergraduate", "master" are recognized, the educational background parsing logic is triggered. For example, when the text appears "A University|Computer Science and Technology|Master of Engineering", the university name (A University), major (Computer Science and Technology), and degree level (Master of Engineering) are extracted through the vertical bar separator or fixed sentence patterns (such as "Graduated from XX University, XX major").
[0026] Work Experience: Based on timeline keywords (such as "2020.01 - 2023.12", "up to now") or responsibility guiding words (such as "responsible for", "main work includes"), the work experience paragraphs are located. For example, "XX Company|Software Engineer|Responsible for backend system development, optimizing the database query efficiency by 30%" will be parsed as the employing company (XX Company), position (Software Engineer), and responsibility description (Responsible for backend system development, optimizing the database query efficiency by 30%).
[0027] Pre - collect the voiceprint samples of the interviewer and the candidate (such as the self - introduction segment at the beginning of the interview), and extract the voiceprint feature vectors (such as MFCC Mel - Frequency Cepstral Coefficients) through the Gaussian Mixture Model (GMM). During the transcribing process, the cosine similarity between the voiceprint features of the audio stream and the preset samples is calculated in real - time, and the segments with a similarity higher than 90% are assigned to the corresponding main audio tracks (such as the candidate audio track, the interviewer audio track) to achieve the separation of multi - person conversations.
[0028] The specific logic of Voice Activity Detection (VAD): Detect the start and end of speech through the double - threshold method: When the audio energy exceeds the "high threshold", the speech segment starts, and when the energy is continuously detected to be lower than the "low threshold" and lasts for more than 0.5 seconds, the speech segment ends. For example, when the candidate is answering a question, if the pause exceeds 0.5 seconds in the middle, it is regarded as the end of a semantic unit, and an independent text paragraph is generated and the start timestamp is marked (such as "00:02:15 - 00:03:40").
[0029] For each speech segment, extract the fundamental frequency curve (reflecting pitch changes), calculate the number of rising / falling inflection points in the curve, and divide it by the segment duration (in seconds) to obtain the number of intonation fluctuations per minute (for example, in a 20 - second speech segment, there are 5 inflection points, and the intonation fluctuation frequency is 15 times per minute). At the boundary of the semantic units obtained by voice activity detection, measure the silent duration between two adjacent units (for example, unit A ends at 00:02:30, unit B starts at 00:02:32, then the pause interval is 2 seconds), and take the average value of all pause intervals of this audio track as the feature value.
[0030] Calculate the original speech rate (number of text characters / speech duration, unit: characters per second), and then compare it with the industry standard speech rate range (such as 120 - 150 characters per minute). Map the original speech rate to the 0 - 1 interval through linear transformation (for example, if the original speech rate is 180 characters per minute, which exceeds the standard upper limit by 30%, then the normalized value is 1.3).
[0031] 1. Document type adaptation of the template engine: Word document: By parsing XML tags (such as <w:p>Paragraph <w:tbl>Table) Identify the title hierarchy (e.g., <w:outlinelvl>Marked heading levels) and table structures, and automatically ignore irrelevant areas such as headers and footers when extracting paragraph text.
[0032] PDF documents: Use a PDF parsing library to identify text streams and coordinate positions, and determine the paragraph hierarchy by the Y-axis value of the text coordinates (text with the same or similar Y-axis values is considered the same paragraph). For scanned PDF documents, call the OCR engine for secondary recognition.
[0033] Excel tables: Read cell coordinates and merge attributes. For cells merged across columns (such as A1:C1 merged), infer the actual value of the merged cell (such as "Sales Department" displayed across columns in A1:C1) by traversing the context logic of business data in the same industry (such as the content in the "Department" column of adjacent rows).
[0034] When detecting cells merged across columns, the system scans the same-column data in adjacent rows up / down to find keywords that appear repeatedly (such as "Project Name", "Responsible Person"), and infers the attributes of the merged cells through pattern matching. For example, in a certain table, A1:C1 is merged, and in the row below, A2 is "Project 1", B2 is "Zhang San", and C2 is "2023", then it is inferred that A1:C1 is the title row of "Project Information" to avoid content fragmentation caused by formatting issues.
[0035] The dictionary predefines the mapping relationships of standard terms in different industries (such as "JAVA" unified as "Java programming language", and "PM" mapped as "Project Manager" or "Product Manager" according to the context). Scan the text through the forward maximum matching algorithm, and add semantic tags to skill names (such as "AI development" matched to "Artificial intelligence development") and project keywords (such as "big data platform" matched to "Big data analysis platform") to ensure term consistency.
[0036] Normalization of time and location: Unification of time format: Identify multiple time expression forms (such as "May 2023", "2023.05", "2023-05"), and uniformly convert them to the "YYYY-MM" standard format; for fuzzy time (such as "nearly 3 years"), combine the current time (such as 2025) and convert it to the "2022-2025" interval.
[0037] Through technologies such as adaptive threshold segmentation, voiceprint recognition, and template engine, data errors caused by factors such as image blurring, multi-person speech confusion, and complex document formats are reduced; semantic label addition and format normalization processing make the data have a unified specification. Operations such as the annotation of voice emotional fundamental frequency features and the semantic reconstruction of cross-column cells add dimensional information such as emotional expression and content logic to the data, completely retaining the key content in the resume and interview and avoiding information omission. The standardized data set can be adapted to a variety of data analysis models and algorithms, reducing the time and cost of data preprocessing; the structured storage method also facilitates the rapid retrieval and invocation of data.
[0038] In a preferred embodiment of the present invention, in the above step S2, context semantic encoding is performed on the standardized text data set, and node alignment and semantic disambiguation processing are performed on professional terms through a domain knowledge graph to generate a text feature vector containing entity association relationships, which may include: Step S200, based on the semantic labels and structured document formats in the standardized text data set, perform semantic analysis on the text using context semantic encoding technology to generate an initial semantic vector; Step S201, input the initial semantic vector into a domain structured semantic network for term alignment. By traversing the professional term nodes in the semantic network, calculate the semantic similarity between the terms in the text and the nodes, map the terms with a similarity higher than the preset threshold to the corresponding nodes, and expand the context semantics of the terms based on the association paths between the nodes. The expansion of the context semantics of the terms includes upstream and downstream skill associations, industry scenario constraint conditions, and job ability dependencies; Step S202, construct a disambiguation weight using the co-occurrence relationship and hierarchical classification information of the nodes in the semantic network, and calculate the probability distribution of ambiguous terms based on the context window of the terms in the text, and dynamically adjust the vector representation of the terms to eliminate semantic ambiguity and obtain a disambiguation result; Step S203, fuse the aligned term node features, extended semantics, and disambiguation results, splice the attribute features of the nodes in the semantic network with the text semantic vector, and generate a text feature vector containing entity association relationships.
[0039] In an embodiment of the present invention, the standardized text data is split into independent text blocks by fields (such as "Educational Background", "Work Experience"), and each text block is tokenized (for example, splitting "Responsible for the back-end development of the e-commerce platform, using Java and Spring frameworks" into "Responsible for", "e-commerce platform", "back-end development", "using", "Java", "Spring framework"). A pre-trained language model (such as a BERT-like model) is used to analyze the context dependency relationship of each word. For example, in "Develop the back-end system using Java", "Java" has a strong association with "back-end development", while in "Introduction to Java Programming Language", it focuses more on basic concepts. The model captures this semantic difference through a multi-layer neural network and generates an embedding vector containing context information for each word (such as a 300-dimensional vector). The word vectors of the same text block are aggregated through average pooling or an attention mechanism to generate the initial semantic vector of the text block. For example, the vector of the "Work Experience" field needs to synthesize the semantic information of all keywords in the responsibility description.
[0040] Step S201, the semantic network is constructed based on the industry knowledge graph. The nodes are professional terms (such as "Python", "Cloud Computing", "Agile Development"), the edges are the association relationships between terms (such as "Python → Programming Language", "Cloud Computing → AWS / Azure", "Agile Development → Scrum Framework"), and the node attributes are marked (such as "Programming Language", "Tool Platform"). For example, when the term "Spring Boot" appears in the text, the system traverses the nodes of the semantic network and calculates its semantic similarity with the "Spring Framework" node (through cosine similarity or edit distance). If the similarity is higher than the threshold (such as 0.8), then "Spring Boot" is mapped to the "Spring Framework" node. After mapping, the semantics are extended along the association path of the semantic network nodes: Upstream and downstream skill associations: The "Spring Framework" node is associated with sub-nodes such as "Dependency Injection", "Aspect-Oriented Programming", and the extended semantics include these technical details; Industry scenario constraints: If the position is "Back-end development in the financial industry", the "Spring Framework" node will be associated with industry-specific constraints such as "Financial-level transaction processing"; Position ability dependencies: If the position requirements include "Microservices architecture", the "Spring Framework" node is further associated with ability items such as "Spring Cloud", "Service registration and discovery".
[0041] Step S202: Assign disambiguation weights to each term node by leveraging the co-occurrence relationships (e.g., "algorithm" often co-occurs with "data structure" and "machine learning") and hierarchical classifications (e.g., "algorithm" belongs to the "technical ability" hierarchy) in the semantic network. For example, in "machine learning algorithm optimization", the co-occurrence frequency of "algorithm" with the "machine learning" node is high, so the disambiguation weight inclines towards "machine learning-related algorithms".
[0042] Calculation of the probability of ambiguous terms: Centering around the term, take the 5 words before and after as the context window (e.g., "algorithm" in "optimize algorithm performance"), and statistically calculate the association strength between the words in the window and each node in the semantic network. For example, "performance" has a strong association with the "algorithm optimization" node and a weak association with the "data structure" node. Therefore, in this context, "algorithm" is more likely to refer to "optimization algorithm" rather than "basic algorithm".
[0043] Vector dynamic adjustment: Adjust the initial vector of the term according to the probability distribution of the context window. For example, if there is a 70% probability that "algorithm" belongs to the "optimization algorithm" node and a 30% probability that it belongs to the "basic algorithm" node, then its vector is the weighted sum of the vectors of the two nodes (weights 0.7:0.3), which can more accurately reflect the semantics after eliminating ambiguity.
[0044] Step S203: Concatenate the following features into the final text feature vector by dimension: Aligned node features: The attribute vectors of the term mapping nodes (e.g., the "programming language" and "enterprise-level development" attributes of the "Spring framework" node); Extended semantic features: Feature vectors such as upstream and downstream skills and industry scenarios extracted from the association path (e.g., "microservice architecture", "financial compliance"); Vector after disambiguation: The term vector after dynamic adjustment, which has excluded the interference of ambiguity; Original text semantic vector: The initial vector generated in Step S200, which retains the unstructured context information.
[0045] Example of vector dimension: Suppose the initial vector is 300-dimensional, the node attribute vector is 100-dimensional, the extended semantic vector is 200-dimensional, and the disambiguation vector is 300-dimensional. Then the total dimension of the concatenated text feature vector is 300 + 100 + 200 + 300 = 900 dimensions, which comprehensively covers the term semantics, association relationships, and context information.
[0046] By context semantic encoding, capture the subtle differences of terms in different scenarios (such as the different meanings of "product manager" in the Internet and traditional industries), and avoid the understanding deviation caused by "polysemy". The term alignment mechanism of the domain knowledge graph ensures the consistency of professional terms with industry standards (such as accurately mapping "AI development" to the node of "artificial intelligence development"), and reduces the interference of non-standard expressions. The extension of the association path of the semantic network (such as "Java → back-end development → microservices architecture") enables the text feature vector to contain deep skill dependency relationships, rather than isolated term stacks. After disambiguation processing, the vector can accurately reflect the specific context of the term (such as distinguishing between "deep learning algorithm" and "data structure algorithm"), and improve the refinement degree of semantic representation. The feature vector containing entity association relationships can be directly used for the dynamic matching of job requirements and candidate capabilities (such as detecting the skill chain matching degree between "Java development" and "SpringCloud microservices"). The structured feature vector adapts to advanced analysis models such as graph neural network (GNN), providing a rich semantic basis for subsequent personality trait assessment, career development prediction, etc. The disambiguation mechanism effectively processes the ambiguous expressions in the resume (such as "responsible for system optimization"), and infers the specific technical direction through the context (such as "database optimization" or "algorithm optimization"), reducing the impact of data noise.
[0047] In a preferred embodiment of the present invention, in the above step S3, dynamically semantically match the text feature vector with the job requirement vector, optimize the vector representation of the polysemous word in the job context, analyze the emotional tendency characteristics of the dialogue text, and generate personality trait assessment parameters, which may include: Step S300, extract skill requirements, years of experience, and ability keywords according to the job requirement description, and construct a job requirement vector. Among them, the skill requirement keywords are semantically enhanced based on the node attributes in the domain structured semantic network. The semantic enhancement extracts the upstream and downstream skill dependency relationships and industry scenario constraint conditions by traversing the association path of the nodes in the semantic network, and generates a job requirement vector with the same dimension as the text feature vector; Step S301, when extending the co-occurrence word distribution of the polysemous word through the association path in the semantic network, dynamically load the co-occurrence word set related to the job requirement based on the upstream and downstream skill dependency relationships and industry scenario constraint conditions between nodes; Step S302, count the occurrence frequency and distribution density of the co-occurrence word in the job requirement description, and generate a coverage rate index; Step S303, adjust the vector weight of the polysemous word according to the coverage rate index, and generate an optimized semantic matching score; Step S304, input the optimized semantic matching score into the multi-modal fusion layer, segment the dialogue text by speaker based on the timing tags, and extract the fundamental frequency parameters of the acoustic features of each text segment. The acoustic features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate, which specifically include: Step S3040, based on the timing tags of the dialogue text, cut the mixed audio into independent speech segments according to the speaker's track segment identification, perform fundamental frequency extraction on each speech segment, and generate fundamental frequency change data; Step S3041, according to the first-order difference value of the fundamental frequency change data, count the number of mutations in the fundamental frequency change direction and amplitude within the speech segment, and calculate the intonation fluctuation frequency per unit time in combination with the segment duration; Step S3042, perform voice endpoint detection on the speech segment, identify the starting position of the semantic boundary point, divide the speech segment into semantic unit segments based on the boundary point, and calculate the average silence interval between adjacent semantic unit segments as the semantic pause interval; Step S3043, extract the total duration of the speech segment and the number of characters in the corresponding text segment, calculate the initial speech rate according to the ratio of the number of characters to the speech duration, and perform linear normalization on the initial speech rate based on the preset industry standard speech rate range to generate the speech rate change rate; Step S3044, bind the intonation fluctuation frequency, semantic pause interval, and speech rate change rate with the timing tags of the corresponding text segment to generate structured acoustic feature fundamental frequency parameters; Step S305, match the emotional keywords in each text segment through an emotion dictionary, and calculate the emotional tendency intensity value in combination with the acoustic feature fundamental frequency parameters; Step S306, input the emotional tendency intensity value and the text content into the emotion calculation module, fuse the acoustic features and text semantics through the timing alignment technology, and output the smoothed emotional polarity score sequence, which specifically includes: Step S3060, based on the timing tags of the dialogue text and the timestamps of the acoustic feature fundamental frequency parameters, establish a timing mapping relationship between the text paragraphs and the speech segments, and generate a time-aligned multi-modal data index; Step S3061, perform dynamic sentence splitting on the text paragraphs according to the timing mapping relationship, extract the speech segment time window corresponding to each sentence, and match the intensity weight of the emotional keywords in the corresponding sentence based on the emotion dictionary; Step S3062, splice the acoustic feature fundamental frequency parameters and the emotional keyword intensity weights within the same time window for multi-modal data, and extract cross-modal context correlation feature vectors by scanning the fused features on the time series through a convolutional kernel; Step S3063: Calculate the sentiment tendency offset based on the distribution density of the cross-modal context correlation feature vector, and generate an initial sentiment polarity score according to the projection direction and magnitude of the offset in the preset sentiment dimension coordinate system. Step S3063: Perform a sliding window mean filtering process on the initial sentiment polarity scores of consecutive time windows, and output a smoothed sequence of sentiment polarity scores. Step S307: Calculate the variance of the emotional stability index according to the fluctuation amplitude and frequency of the sentiment polarity scores in the time series. Step S308: In the multi-modal fusion layer, perform feature-level fusion on the semantic matching score, the sentiment polarity score, and the emotional stability index to generate a joint vector, and analyze the data distribution pattern of the joint vector to generate personality trait evaluation parameters, including dimensions such as decision-making tendency, teamwork degree, and stress resistance ability.
[0048] In the embodiment of the present invention, the job requirement description is hierarchically analyzed as follows: Keyword extraction: Match the regular expression with the domain dictionary to extract the core elements from the text. For example, from "Recruit AI algorithm engineer, with more than 5 years of experience, need to master deep learning frameworks (such as TensorFlow), natural language processing technology, and have the ability of distributed training optimization", extract the skill keywords "deep learning framework", "natural language processing", the experience years "more than 5 years", and the ability keyword "distributed training optimization".
[0049] Semantic enhancement: Based on the domain knowledge graph (such as the IT industry graph), expand the context of the skill keywords. Taking "deep learning framework" as an example, the upstream and downstream nodes associated with this node in the graph include "TensorFlow / PyTorch ecosystem", "model training process", "hardware acceleration adaptation", etc. Traverse these paths and incorporate scenario constraints such as "experience in optimizing the TensorFlow framework for distributed training" into the keyword semantics.
[0050] Vector construction: Map the enhanced keywords to the same dimensional space as the text feature vector (such as 500 dimensions), and assign initial weights according to the keyword importance (core skill weight 0.6 - 0.8, auxiliary skill 0.2 - 0.4) to generate a job requirement vector.
[0051] Step S301: Take the polysemous word "model" as an example (such as the job requirement "need to optimize the performance of the recommendation model"): Semantic network traversal: Locate the "recommendation model" node in the knowledge graph. Its associated paths include upstream and downstream skills such as "collaborative filtering algorithm", "user feature engineering", "online learning mechanism", etc., as well as industry constraints such as "real-time requirements in e-commerce scenarios" and "cold start strategy for content platforms". According to the path weights (direct associated node weight > indirect association), dynamically load the top 5 - 10 highly relevant co-occurring words, such as "collaborative filtering", "cold start", "real-time recommendation", and exclude words irrelevant to the job scenario (such as "physical model").
[0052] Step S302, assume the co-occurring word set contains 8 words, among which "collaborative filtering", "cold start", and "real-time recommendation" appear in the job description, and the coverage rate is = 37.5%. At the same time, analyze the distribution density: If all 3 words appear in the "key skills" paragraph, the density index is 100%; if they are scattered in different paragraphs, the density index is weighted and calculated according to the paragraph weights (such as the weight of the "key skills" paragraph is 0.7, and the weight of the "other requirements" paragraph is 0.3) as 0.7× + 0.3× ≈ 0.63.
[0053] Step S303, according to the coverage rate (37.5%) and density (63%), adjust the vector weight of the polysemous word "model" according to the preset rules: If the coverage rate < 50% and the density < 70%, the weight drops by 0.1 - 0.2 (such as from the initial 0.5 to 0.4); combine the cosine similarity (range [-1, 1]) between the job requirement vector and the candidate text feature vector to generate a semantic matching score (such as 0.72).
[0054] Step S304, extraction of acoustic feature fundamental frequency parameters: Speech segmentation and fundamental frequency extraction: Cut the interview recording into short segments of 1 - 5 seconds according to the speaking subject, and use tools such as Praat to extract the fundamental frequency curve. For example, the fundamental frequency range of the candidate's speech segment is 120 - 180 Hz.
[0055] Intonation fluctuation frequency: Calculate the number of zero-crossing points of the first-order difference of the fundamental frequency curve. For example, if there are 4 fluctuations within 5 seconds, the frequency is 0.8 times per second.
[0056] Semantic pause interval: Detect the silent segments (such as > 200 ms) through the VAD algorithm, and calculate the average interval of adjacent semantic units. For example, for 3 pauses of 300 ms, 500 ms, and 400 ms respectively, the average value is 400 ms.
[0057] Speech rate change rate: The 10 - second speech corresponds to 60 - character text, with an initial speech rate of 6 characters per second. Normalize it according to the industry standard range (4 - 8 characters per second) to = 0.5.
[0058] Step S305: For the text "After the recommendation system led by me was launched, the click-through rate increased by 20%", the sentiment dictionary matches "led" (positive +0.6), "increased" (positive +0.8), and combined with acoustic features (speech rate 0.5, intonation fluctuation 0.8 times / second are both higher than the average, each with a weight of 0.1), the intensity value is ×(1 + 0.1 + 0.1) = 0.84.
[0059] Step S306: Align the text clauses and speech segments according to the timestamps (such as 00:01:00 - 00:01:10) to establish a multimodal index. Concatenate the acoustic parameters (intonation 0.8, pause 0.4 seconds, speech rate 0.5) and the sentiment keyword intensity (0.84) into a feature vector, and extract the temporal features through 1D convolution (such as the positive scores increasing in 3 consecutive windows). Apply a 3-window mean filter to the initial score sequence [0.7, 0.8, 0.9, 0.85] to obtain [0.77, 0.85, 0.85].
[0060] Step S307: For the sentiment polarity sequence [0.6, 0.9, 0.7, 0.8], calculate the variance as = 0.0125, reflecting relatively small emotional fluctuations.
[0061] Step S308: Fuse the semantic matching score (0.72), the average sentiment polarity (0.8), and the variance of emotional stability (0.0125) according to the weights (0.6:0.3:0.1) into a joint vector, and generate through a preset mapping rule: Decision tendency: High match + positive emotion → Tend to make decisive decisions (0.8 points); Stress resistance: High match + low variance → Strong stress resistance (0.9 points).
[0062] Enhancing the job requirement vector through the domain semantic network and optimizing the representation of polysemous words can deeply understand the connotation of job skill requirements, avoid misjudgment caused by surface keyword matching, and improve the accuracy of the matching between talents and jobs. Conducting sentiment analysis by combining the acoustic features and semantic information of the dialogue text can not only understand the professional abilities of candidates, but also evaluate their personality traits, comprehensively examine candidates from multiple dimensions, and provide a more abundant reference basis for recruitment decisions. Mechanisms such as dynamically loading the co-occurrence word set and adjusting weights according to the coverage rate can adapt to the demand changes of different jobs and industries, enhancing the versatility and flexibility of the system. Transforming abstract factors such as sentiment tendency and personality traits into specific quantitative indicators makes the evaluation results more objective and comparable, reducing the interference of human factors.
[0063] In a preferred embodiment of the present invention, in step S4, four orthogonal detection points are set in the semantic matching result, a quadrilateral is constructed according to the semantic coverage, context dependence, term accuracy, and sentiment consistency indexes corresponding to the detection points, and a dynamic correction coefficient is generated by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, and the text feature vector is dimensionally calibrated to obtain a calibrated text feature vector, which may include: Step S400, set four orthogonal detection points in the semantic matching result, corresponding to the semantic coverage, context dependence, term accuracy, and sentiment consistency indexes respectively; Step S401, construct the vertex coordinates of the quadrilateral according to the index values of the four orthogonal detection points, and generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, specifically including: Step S4010, map the semantic coverage, context dependence, term accuracy, and sentiment consistency indexes to the four vertices in the two-dimensional coordinate system respectively, where the semantic coverage and term accuracy are used as the positive and negative direction components of the horizontal axis, and the context dependence and sentiment consistency are used as the positive and negative direction components of the vertical axis to generate an initial set of quadrilateral vertex coordinates; Step S4011, calculate the geometric centroid coordinates of the quadrilateral based on the set of vertex coordinates, and perform relative coordinate transformation on the four vertex coordinates with the centroid coordinates as the origin to generate a normalized vertex offset based on the centroid; Step S4012, calculate the standard deviation according to the distribution dispersion of the normalized vertex offset to generate a vertex distribution eigenvalue representing the spatial dispersion degree of the vertices; based on the initial set of quadrilateral vertex coordinates, calculate the Euclidean distance ratio between adjacent vertices as the side length ratio coefficient, and extract the cosine value of the included angle of each side vector, and perform arithmetic averaging on the cosine values to generate a side length correlation coefficient; Step S4013, linearly superimpose the vertex distribution eigenvalue and the side length correlation coefficient according to a preset weight coefficient to generate an initial correction factor, and perform interval normalization processing on the initial correction factor to output a dynamic correction coefficient; Step S402, use the dynamic correction coefficient to perform dimensional calibration on the text feature vector to generate a calibrated text feature vector.
[0064] In the embodiment of the present invention, four orthogonal detection points are defined in the semantic matching result, and each detection point corresponds to an evaluation index: Semantic coverage: Measure the semantic matching range between the candidate's skills and the job requirements (for example, if the job requirements include 5 skills and the candidate covers 4, it is 80%); Context dependence: Evaluate the logical association strength between the terms in the text and the context (for example, the dependence of "microservice" in "SpringCloud microservice architecture" is higher than that when it appears alone); Term accuracy: To judge the standardization of the use of professional terms (e.g., whether "AI" accurately corresponds to "artificial intelligence" rather than other meanings); Emotional consistency: To detect the degree of fit between the emotional tendency of the dialogue text and the job requirement scenario (e.g., in a technical job interview, excessive emotional expression may have a lower consistency).
[0065] For example, the semantic coverage of a certain candidate is 75%, the context dependence is 60%, the term accuracy is 85%, and the emotional consistency is 70%.
[0066] Step S4010: Map the four index values to the four vertices of a two-dimensional coordinate system. The positive and negative directions of the horizontal axis are semantic coverage (positive) and term accuracy (negative) respectively, and the positive and negative directions of the vertical axis are context dependence (positive) and emotional consistency (negative) respectively. Assuming the index value range is [0, 1], then: Semantic coverage (75%) corresponds to vertex A: (0.75, 0); Context dependence (60%) corresponds to vertex B: (0, 0.6); Term accuracy (85%) corresponds to vertex C: (-0.85, 0); Emotional consistency (70%) corresponds to vertex D: (0, -0.7); Generate the initial quadrilateral vertex coordinate set: {A(0.75, 0), B(0, 0.6), C(-0.85, 0), D(0, -0.7)}.
[0067] Step S4011: Calculate the geometric centroid: The centroid coordinates G(x̄, ȳ) are the mean values of the vertex coordinates: x̄ = = -0.025, ȳ = = -0.025 Subtract the centroid coordinates from each vertex to obtain the normalized offset: A’: (0.75 + 0.025, 0 + 0.025) = (0.775, 0.025); B’: (0 + 0.025, 0.6 + 0.025) = (0.025, 0.625); C’: (-0.85 + 0.025, 0 + 0.025) = (-0.825, 0.025); D’: (0 + 0.025, -0.7 + 0.025) = (0.025, -0.675).
[0068] Step S4012: Vertex distribution eigenvalue (standard deviation): Calculate the standard deviation of the normalized offset on the x and y axes to reflect the degree of vertex dispersion: X-axis offset: 0.775, 0.025, -0.825, 0.025 → Standard deviation ≈ 0.65; Y-axis offset: 0.025, 0.625, 0.025, -0.675 → Standard deviation ≈ 0.58; Take the average value as the vertex distribution eigenvalue: = 0.615 (the larger the value, the more dispersed the distribution).
[0069] Side length correlation coefficient: Calculate the Euclidean distance between adjacent vertices (such as the distance from A' to B' is approximately equal to 0.774, the distance from B' to C' is approximately equal to 1.45, the distance from C' to D' is approximately equal to 0.827, and the distance from D' to A' is approximately equal to 1.45), and obtain the side length ratio coefficient ( ≈ ≈ 1.87); Calculate the cosine value of the included angle of each side vector (such as the cosine of the included angle between A'B' and B'C' is approximately equal to -0.98, close to a right angle). After arithmetic averaging, the side length correlation coefficient is approximately equal to 0.92 (the closer the value is to 1, the more regular the quadrilateral).
[0070] Step S4013, preset the weight of the vertex distribution eigenvalue as 0.6, the weight of the side length correlation coefficient as 0.4, and the initial correction factor = 0.6×0.615 + 0.4×0.92 = 0.731; map the correction factor to the interval [0, 1] (such as through linear transformation) to obtain the dynamic correction coefficient approximately equal to 0.78.
[0071] Step S402, use the dynamic correction coefficient (0.78) to perform weighted adjustment on each dimension of the text feature vector: If a certain dimension corresponds to a semantic coverage-related feature, its weight is multiplied by the correction coefficient (such as the original weight 0.3 → 0.3×0.78 = 0.234); for the dimension related to term accuracy, the weight adjustment direction is opposite (such as the original weight 0.2 → 0.2×(2 - 0.78) = 0.244), and finally a calibrated vector is generated.
[0072] The abstract indicators are transformed into geometric features through a quadrilateral model, which intuitively reflects the balance of semantic matching (for example, the more regular the quadrilateral, the more coordinated the four indicators), and avoids evaluation biases caused by too high or too low of a single indicator. A correction coefficient is generated based on the real-time matching results, which can automatically adjust the vector dimension weights. For example, if the semantic coverage is high but the term accuracy is low (the horizontal axis of the quadrilateral shifts greatly), the coverage weight is reduced and the accuracy weight is increased to make the vector more in line with the job requirements. By calculating the standard deviation of the vertex distribution and the side length ratio, outliers can be effectively identified (such as when an indicator suddenly deviates from other indicators), the data fluctuations are smoothed, and the stability of the feature vector is improved. The quadrilateral model provides interpretability for the evaluation process. Recruiters can quickly locate the weak links in the matching by observing the vertex distribution (such as the imbalance of the vertical axis of the quadrilateral caused by low emotional consistency), which helps with manual review.
[0073] In a preferred embodiment of the present invention, in the above step S5, based on the calibrated text feature vector and the personality trait evaluation parameters, an interactive evaluation report including a dynamic skill topology map, a fitness heat distribution, and a career development prediction curve is generated in combination with the job requirement vector, which may include: In the embodiment of the present invention, candidate skill nodes (such as "Java", "microservice architecture", "Python") are extracted from the calibrated text feature vector, and the skill mastery level is determined according to the vector weight (the higher the weight, the more proficient the mastery); target skill nodes (such as "SpringCloud", "distributed system") are extracted from the job requirement vector, and the upstream and downstream skills in the domain knowledge graph are associated (such as "SpringCloud → service registration and discovery → Eureka").
[0074] Graph structure construction: Candidate skill nodes are represented by circles, and the size is proportional to the weight (for example, the diameter of the "Java" node with a weight of 0.8 is 20px, and the diameter of the "Python" node with a weight of 0.5 is 12px); Job requirement nodes are represented by squares, with a blue color. The edges of the requirement nodes that the candidate has mastered are outlined in gold (such as if the "SpringCloud" node is covered, the edge is outlined in gold).
[0075] Edge rendering: The associated edges between skills (such as "Java → Spring framework") are represented by dashed lines, and the transparency is adjusted according to the co-occurrence frequency (the transparency of high-frequency associated edges is 0.8, and that of low-frequency ones is 0.3); the matching edges between the candidate skills and the job requirements are represented by solid red lines (such as the connection between "Java" and "SpringCloud", and the line width is proportional to the matching score).
[0076] When hovering over a node with the mouse, detailed information is displayed (such as the duration of skill mastery and project application cases); when clicking on a job requirement node, a sub-skill tree can be expanded (such as "Distributed System → Load Balancing → Nginx"), and the mastered sub-nodes of the candidate are highlighted. The calibrated text feature vectors and job requirement vectors are projected onto a two-dimensional space (such as through dimensionality reduction by principal component analysis PCA), with the horizontal axis representing "Technical Competence Matching Degree" and the vertical axis representing "Soft Skill Matching Degree"; each data point represents a skill or trait dimension (such as "Algorithm Ability" and "Communication Ability"), and the coordinates are determined by the corresponding dimension values in the vector (for example, if the algorithm ability matching value is 0.7 and the communication ability is 0.6, then the coordinates are (0.7, 0.6)). A red-yellow-blue gradient color is used, where the red area indicates high fitness (matching value > 0.8), yellow indicates medium fitness (0.5 - 0.8), and blue indicates low fitness (< 0.5); personality trait parameters (such as decision-making tendency and teamwork degree) are used as the third dimension and are represented by the size of the bubbles (for example, a teamwork degree of 0.9 corresponds to a bubble radius of 15px, and 0.5 corresponds to 8px).
[0077] Four quadrants are divided by dashed lines: The first quadrant (high technology + high soft skills): Core advantage area; the fourth quadrant (high technology + low soft skills): Area with both potential and risks.
[0078] Dragging the slider can switch the displayed dimension (such as switching from "Technical Competence" to "Industry Experience"); clicking on the heat area can view the list of specific skills or traits included in that area and the matching details.
[0079] Career development prediction curve generation: Candidate data: Calibrated skill vectors, years of work experience, project experience complexity (extracted from descriptions such as "responsible for XX million-user systems" in the text); Job trend data: Obtain the skill demand change trend of the target job through industry reports (such as the annual growth rate of the "Large Model Training" skill demand in the "AI Algorithm Engineer" job is 30%); Personality trait parameters: Stress resistance, learning ability, etc. are used to adjust the prediction slope (for example, for candidates with strong stress resistance, the skill improvement speed can be accelerated by 10% - 20%).
[0080] Skill growth curve: Basic growth rate: Set according to the industry average level (such as the "Distributed System" skill grows naturally by 5% per year); Acceleration factor: For every 1 unit increase in the candidate's learning ability (such as from 0.6 to 0.7), the growth rate increases by 2%; Generate the skill mastery curve for the next 1 - 3 years (such as the current mastery of "Large Model Training" is 0.4, predicted to be 0.6 after 1 year and 0.8 after 2 years).
[0081] Adaptability fluctuation curve: By combining the job demand trend (for example, when the demand for a certain skill decreases, the fitness curve drops) and the candidate's skill growth, a dynamic fitness curve is generated (for example, the current total fitness is 0.7, and it is predicted that it will drop to 0.65 in one year due to changes in job requirements, but will rise back to 0.78 after the candidate's skills improve).
[0082] The orange curve represents the skill growth trend, and the blue curve represents the change in fitness. The shaded area represents the prediction confidence interval (e.g., at 95% confidence, the skill mastery fluctuates by ±5%). Key time points are marked (e.g., "XX new skill needs to be mastered in 6 months") and associated training resource recommendations are given.
[0083] like Figure 2 As shown, the embodiment of the present invention also provides a talent evaluation management system based on AI intelligence, including: Data collection module, used to extract multi-source resume information through OCR recognition, speech transcription and format analysis and standardize it into a text dataset with semantic labels; The semantic encoding module is used to perform contextual semantic encoding and domain knowledge graph disambiguation on the standardized text dataset to generate text feature vectors containing entity association relationships; Dynamic semantic matching module, which is used to dynamically match text feature vectors with job requirement vectors and analyze conversation sentiment tendencies to generate personality trait assessment parameters; A dynamic correction module is used to construct an evaluation quadrilateral by setting orthogonal detection points, calculate eigenvalues and correlation coefficients to generate correction coefficients, and generate calibrated eigenvectors; The evaluation report module is used to generate an interactive evaluation report based on the calibrated feature vectors, personality parameters and job requirement vectors.
[0084] It should be noted that the system is a system corresponding to the above method, and all implementation methods in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.
[0085] The embodiment of the present invention also provides a computer-readable storage medium storing instructions, which, when executed on a computer, enable the computer to execute the method described above. All implementations in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.
[0086] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.< / w:outlinelvl> < / w:tbl> < / w:p>
Claims
1. A talent evaluation and management method based on AI intelligence, characterized in that, The method includes the following: Step S1: Convert the handwritten resume image into structured text data through OCR image recognition; convert the interview recording into conversation text with time sequence tags through speech transcription; perform format parsing on the electronic document to extract the original text content; perform encoding standardization processing on the structured text data, conversation text, and original text content to generate a standardized text data set containing semantic tags; Step S2: Perform context semantic encoding on the standardized text data set, perform node alignment and semantic disambiguation processing on professional terms through a domain knowledge graph, and generate a text feature vector containing entity association relationships; Step S3: Perform dynamic semantic matching between the text feature vector and the job requirement vector, optimize the vector representation of polysemous words in the job context, analyze the emotional tendency characteristics of the conversation text, and generate personality trait evaluation parameters; Step S4: Set four orthogonal detection points in the semantic matching result, construct a quadrilateral according to the semantic coverage, context dependence, term accuracy, and emotional consistency indicators corresponding to the detection points, generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, and perform dimension calibration on the text feature vector to obtain a calibrated text feature vector; Step S5: Based on the calibrated text feature vector and personality trait evaluation parameters, generate an interactive evaluation report containing a dynamic skill topology map, fitness heat distribution, and career development prediction curve in combination with the job requirement vector.
2. The talent evaluation and management method based on AI intelligence according to claim 1, characterized in that, Convert the handwritten resume image into structured text data through OCR image recognition; convert the interview recording into conversation text with time sequence tags through speech transcription; perform format parsing on the electronic document to extract the original text content; Perform encoding standardization processing on the structured text data, conversation text, and original text content to generate a standardized text data set containing semantic tags, including: When performing OCR image recognition on the handwritten resume image, use an adaptive threshold segmentation algorithm to perform illumination equalization processing on the image, segment the candidate text area, and parse the text area. The recognition result is classified and mapped to structured field data according to the resume fields. Among them, the education background field is parsed into the school name, major, and degree level, the work experience field is parsed into the company name, position, and responsibility description, and the skill certificate field is parsed into the certificate name and certification agency; when performing speech transcription on the interview recording, identify different speakers based on voiceprint features. During the recognition process, separate the mixed audio into independent tracks, perform segmented transcription on each track in combination with voice endpoint detection to generate conversation text with time sequence tags, and annotate the voice emotional fundamental frequency features in the conversation text. The emotional fundamental frequency features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate, and the emotional fundamental frequency features are extracted through acoustic parameters and bound to the text paragraph; when performing format parsing on the electronic document, use a preset template engine according to the document type to identify the table, paragraph, and title hierarchical structure in the document, extract the unformatted original text content, and perform row-column association mapping on the text content. The association relationship of merged cells across columns is realized through context semantic reconstruction; Input the structured field data, the dialogue text with time stamps, and the original text content into the encoding and standardization module. Add semantic tags to the skill names and project experience keywords through a preset industry term dictionary, unify the time format into the standard format, and normalize the location names into the full names of administrative regions to generate a standardized text data set containing semantic tags. The data set is stored in a structured document format according to the field type, and each field is associated with a semantic tag and a normalized attribute value.
3. The talent evaluation and management method based on AI intelligence according to claim 2, characterized in that, Perform context semantic encoding on the standardized text data set, and perform node alignment and semantic disambiguation processing on professional terms through a domain knowledge graph to generate a text feature vector containing entity association relationships, including: Based on the semantic tags and structured document format in the standardized text data set, use context semantic encoding technology to perform semantic analysis on the text to generate an initial semantic vector; Input the initial semantic vector into the domain structured semantic network for term alignment. By traversing the professional term nodes in the semantic network, calculate the semantic similarity between the terms in the text and the nodes, map the terms with similarity higher than the preset threshold to the corresponding nodes, and expand the context semantics of the terms based on the association paths between the nodes. The expanded context semantics of the terms include upstream and downstream skill associations, industry scenario constraint conditions, and job ability dependencies; Use the co-occurrence relationship and hierarchical classification information of the nodes in the semantic network to construct a disambiguation weight, and calculate the probability distribution of ambiguous terms based on the context window of the terms in the text to dynamically adjust the vector representation of the terms to eliminate semantic ambiguity and obtain a disambiguation result; Fuse the feature of the aligned term nodes, the extended semantics, and the disambiguation result, and splice the attribute features of the nodes in the semantic network with the text semantic vector to generate a text feature vector containing entity association relationships.
4. The AI intelligence-based talent evaluation and management method according to claim 3, wherein Perform dynamic semantic matching between the text feature vector and the job requirement vector, optimize the vector representation of polysemous words in the job context, and analyze the emotional tendency features of the dialogue text to generate personality trait evaluation parameters, including: Extract skill requirements, years of experience, and ability keywords according to the job requirement description to construct a job requirement vector. Among them, the skill requirement keywords are semantically enhanced based on the node attributes in the domain structured semantic network. The semantic enhancement extracts the upstream and downstream skill dependencies and industry scenario constraint conditions by traversing the association paths of the nodes in the semantic network to generate a job requirement vector with the same dimension as the text feature vector; When expanding the co-occurrence word distribution of polysemous words through the association path in the semantic network, dynamically load the co-occurrence word set related to the job requirements based on the upstream and downstream skill dependencies and industry scenario constraint conditions between the nodes; Count the occurrence frequency and distribution density of co-occurrence words in the job requirement description to generate a coverage rate index; Adjust the vector weight of polysemous words according to the coverage rate index to generate an optimized semantic matching score; Input the optimized semantic matching score into the multi-modal fusion layer, and segment the dialogue text by speaker based on the time stamps, and extract the fundamental frequency parameters of the acoustic features of each segment of text. The acoustic features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate; Match the sentiment keywords in each text segment through a sentiment dictionary, and calculate the sentiment tendency intensity value in combination with the fundamental frequency parameter of the acoustic features; Input the sentiment tendency intensity value and the text content into the sentiment calculation module, fuse the acoustic features and text semantics through the time series alignment technology, and output the smoothed sentiment polarity score sequence; Calculate the variance of the emotion stability index according to the fluctuation amplitude and frequency of the sentiment polarity score in the time series; In the multi-modal fusion layer, perform feature-level fusion on the semantic matching score, the sentiment polarity score, and the emotion stability index to generate a joint vector, and analyze the data distribution pattern of the joint vector to generate personality trait evaluation parameters, including decision-making tendency, teamwork degree, and stress resistance dimension; 5. The talent evaluation and management method based on AI intelligence according to claim 4, characterized in that, Input the optimized semantic matching score into the multi-modal fusion layer, segment the dialogue text by speaker based on the time series mark, and extract the fundamental frequency parameters of the acoustic features of each text segment. The acoustic features include intonation fluctuation frequency, semantic pause interval, and speech rate change rate, including: Based on the time series mark of the dialogue text, cut the mixed audio into independent speech segments according to the speaker's track segmentation identification, perform fundamental frequency extraction on each speech segment, and generate fundamental frequency change data; According to the first-order difference value of the fundamental frequency change data, count the number of mutations in the fundamental frequency change direction and amplitude within the speech segment, and calculate the intonation fluctuation frequency per unit time in combination with the segment duration; Perform voice endpoint detection on the speech segment, identify the starting position of the semantic boundary point, divide the speech segment into semantic unit segments based on the boundary point, and calculate the average silence interval between adjacent semantic unit segments as the semantic pause interval; Extract the total duration of the speech segment and the number of characters in the corresponding text segment, calculate the initial speech rate according to the ratio of the number of characters to the speech duration, and perform linear normalization on the initial speech rate based on the preset industry standard speech rate range to generate the speech rate change rate; Bind the intonation fluctuation frequency, semantic pause interval, and speech rate change rate to the time series mark of the corresponding text segment to generate the structured acoustic feature fundamental frequency parameter; 6. The talent evaluation management method based on AI intelligence according to claim 4, characterized in that Input the sentiment tendency intensity value and the text content into the sentiment calculation module, fuse the acoustic features and text semantics through the time series alignment technology, and output the smoothed sentiment polarity score sequence, including: Based on the time series mark of the dialogue text and the timestamp of the acoustic feature fundamental frequency parameter, establish the time series mapping relationship between the text paragraph and the speech segment, and generate the time-aligned multi-modal data index; Perform dynamic sentence splitting on the text paragraph according to the time series mapping relationship, extract the speech segment time window corresponding to each sentence, and match the intensity weight of the sentiment keyword in the corresponding sentence based on the sentiment dictionary; Perform multi-modal data splicing on the acoustic feature fundamental frequency parameter and the sentiment keyword intensity weight within the same time window, scan the fused features in the time series through the convolution kernel, and extract the cross-modal context correlation feature vector; Calculate the sentiment tendency offset according to the distribution density of the cross-modal context correlation feature vector, and generate the initial sentiment polarity score according to the projection direction and modulus length of the offset in the preset sentiment dimension coordinate system; Perform a sliding window mean filtering process on the initial sentiment polarity scores of consecutive time windows, and output a smoothed sequence of sentiment polarity scores.
7. The talent evaluation management method based on AI intelligence according to claim 6, characterized in that Set four orthogonal detection points in the semantic matching result. Construct a quadrilateral based on the semantic coverage, context dependence, term accuracy, and sentiment consistency indicators corresponding to the detection points. Generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, and perform dimensional calibration on the text feature vector to obtain a calibrated text feature vector, including: Set four orthogonal detection points in the semantic matching result, corresponding to the semantic coverage, context dependence, term accuracy, and sentiment consistency indicators respectively; Construct the vertex coordinates of the quadrilateral according to the index values of the four orthogonal detection points, and generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral; Use the dynamic correction coefficient to perform dimensional calibration on the text feature vector and generate a calibrated text feature vector.
8. The AI intelligence-based talent evaluation management method according to claim 7, wherein Construct the vertex coordinates of the quadrilateral according to the index values of the four detection points, and generate a dynamic correction coefficient by calculating the vertex distribution eigenvalue and side length correlation coefficient of the quadrilateral, including: Map the semantic coverage, context dependence, term accuracy, and sentiment consistency indicators to four vertices in the two-dimensional coordinate system respectively. Among them, the semantic coverage and term accuracy are used as the positive and negative direction components of the horizontal axis, and the context dependence and sentiment consistency are used as the positive and negative direction components of the vertical axis to generate an initial set of quadrilateral vertex coordinates; Calculate the geometric centroid coordinates of the quadrilateral based on the set of vertex coordinates, and perform relative coordinate transformation on the four vertex coordinates with the centroid coordinates as the origin to generate a normalized vertex offset based on the centroid; Calculate the standard deviation according to the distribution dispersion of the normalized vertex offset to generate a vertex distribution eigenvalue representing the spatial dispersion degree of the vertices; Based on the initial set of quadrilateral vertex coordinates, calculate the Euclidean distance ratio between adjacent vertices as the side length ratio coefficient, and extract the cosine values of the included angles of each side vector, and perform arithmetic averaging on the cosine values to generate a side length correlation coefficient; Linearly superimpose the vertex distribution eigenvalue and the side length correlation coefficient according to a preset weight coefficient to generate an initial correction factor, and perform interval normalization processing on the initial correction factor to output a dynamic correction coefficient.
9. A talent evaluation and management system based on AI intelligence, which implements the method described in any one of claims 1 to 8, characterized in that, Including: A data acquisition module for extracting multi-source resume information through OCR recognition, speech transcription, and format parsing and standardizing it into a text data set containing semantic tags; A semantic encoding module for performing context semantic encoding and domain knowledge graph disambiguation on the standardized text data set to generate a text feature vector containing entity association relationships; A dynamic semantic matching module for dynamically matching the text feature vector with the job requirement vector and analyzing the dialogue sentiment tendency to generate personality trait evaluation parameters; A dynamic correction module for constructing an evaluation quadrilateral by setting orthogonal detection points, calculating eigenvalues and correlation coefficients to generate a correction coefficient, and generating a calibrated feature vector; An evaluation report module for generating an interactive evaluation report based on the calibrated feature vector, personality parameters, and job requirement vector.
10. A computer-readable storage medium, characterized in that, A program is stored in the computer-readable storage medium, and when the program is executed by a processor, the method described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Management method and system based on man-post intelligent matching algorithm model
CN119204834A
Talent cultivation recommendation method based on federal learning and natural language processing
CN119988735A
Intelligent talent tag portrait analysis system based on big data
CN120087927A
System
JP2025051255A
Cited By
Voice call real-time transcription system and method
CN120526774A
A voice call real-time transcription system and method
CN120526774B
Talent background investigation method based on multi-source data evaluation
CN120598436A
Data processing method and device for integrated virtual employee care Saas platform
CN120708916A
Knowledge base generation method and system based on AI intelligent agent
CN120910279A