Talent portrait construction method and device based on multi-modal artificial intelligence, equipment and medium

Through multi-modal artificial intelligence technology, multi-dimensional analysis of user resumes has been solved, the problem of low quality of talent portraits in the existing technology has been achieved, more accurate and reliable talent portrait construction has been achieved, and the quality of corporate recruitment decisions has been improved.

CN120236175AInactive Publication Date: 2025-07-01ZHUHAI MEDIA SUNAC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510721086.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The quality of talent portraits in the existing technology is low and cannot provide a reliable basis for corporate recruitment decisions, mainly due to the limitations of resume content and possible false information.

Method used

Using a multi-modal artificial intelligence method, multi-dimensional data collection and analysis of the user's target resume, including speech analysis, video analysis and keyframe recognition, combined with speech generation and text recognition technology, users' reliability and skill scores under target problems are determined, thereby building a high-quality talent portrait.

Benefits of technology

Through multimodal data fusion, users' actual abilities can be reflected more comprehensively and accurately, the quality of talent portraits can be improved, the impact of false information can be reduced, and the reliability of recruitment decisions can be enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236175A_ABST
    Figure CN120236175A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a talent portrait construction method and device based on multi-modal artificial intelligence, equipment and a medium. The method comprises the following steps: acquiring a first audio of a user under a basic problem, and acquiring a second audio and a related video of the user under a target problem; performing voice analysis on the first audio to obtain initial voice features, and performing voice recognition on the second audio to obtain initial text information; performing voice generation according to the initial voice feature and the initial text information to obtain a third audio; determining a first reliability according to the second audio and the third audio in combination with the related video; performing key frame identification on the related video to obtain a target action, and determining a second reliability according to the target action; determining an initial skill score according to the target question and the initial text information; determining a target skill score according to the first reliability and the second reliability in combination with the initial skill score; and determining a target talent portrait corresponding to the user according to the target skill score in combination with the target question.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, equipment and medium for constructing a talent portrait based on multimodal artificial intelligence. Background Art

[0002] An enterprise forms a talent portrait corresponding to a user through the analysis of the user's resume. This portrait can help the enterprise quickly screen out candidates who meet the basic job requirements. However, there are many problems in establishing a talent portrait solely based on the user's resume. First of all, the content of the resume itself has limitations. In order to highlight their own advantages within a limited space, job seekers often focus on describing key experiences and skills, while some relatively minor but possibly important information for the position may be omitted. More seriously, there may also be false information in the resume. Some job seekers exaggerate their work experiences, performance results or skill levels in order to obtain more interview opportunities. Therefore, there is a problem in the prior art that the quality of establishing a talent portrait is low and cannot provide a reliable basis for the enterprise's recruitment decision-making. Summary of the Invention

[0003] The main purpose of the embodiments of the present invention is to provide a method, device, equipment and medium for constructing a talent portrait based on multimodal artificial intelligence, aiming to solve the problem in the related art that the quality of establishing a talent portrait is low and cannot provide a reliable basis for the enterprise's recruitment decision-making.

[0004] In a first aspect, an embodiment of the present invention provides a method for constructing a talent portrait based on multimodal artificial intelligence, including:

[0005] Analyze the target resume of the user to obtain basic questions and target questions, collect the first audio of the user under the basic questions, and collect the second audio and relevant videos of the user under the target questions;

[0006] Perform speech analysis on the first audio to obtain initial speech features, and perform speech recognition on the second audio to obtain initial text information;

[0007] Perform speech generation based on the initial speech features and the initial text information to obtain the third audio of the target question;

[0008] Determine the first reliability of the user under the target question according to the second audio, the third audio and the relevant videos;

[0009] Perform key frame recognition on the relevant videos to obtain target actions, and determine the second reliability of the user under the target question according to the target actions;

[0010] Determine the initial skill score of the user under the target problem according to the target problem and the initial text information;

[0011] Determine the target skill score according to the initial skill score by combining the first reliability and the second reliability;

[0012] Determine the target talent portrait corresponding to the user according to the target skill score in combination with the target problem.

[0013] In a second aspect, an embodiment of the present invention provides a talent portrait construction device based on multimodal artificial intelligence, including:

[0014] A data acquisition module, configured to analyze the target resume of the user to obtain basic questions and target questions, collect the first audio of the user under the basic questions, and collect the second audio and related videos of the user under the target questions;

[0015] A speech analysis module, configured to perform speech analysis on the first audio to obtain initial speech features, and perform speech recognition on the second audio to obtain initial text information;

[0016] A speech generation module, configured to perform speech generation according to the initial speech features and the initial text information to obtain a third audio of the target problem;

[0017] A first analysis module, configured to determine the first reliability of the user under the target problem according to the second audio and the third audio in combination with the related videos;

[0018] A second analysis module, configured to perform key frame recognition on the related videos to obtain target actions, and determine the second reliability of the user under the target problem according to the target actions;

[0019] A third analysis module, configured to determine the initial skill score of the user under the target problem according to the target problem and the initial text information;

[0020] A data fusion module, configured to determine the target skill score according to the initial skill score by combining the first reliability and the second reliability;

[0021] A portrait determination module, configured to determine the target talent portrait corresponding to the user according to the target skill score in combination with the target problem.

[0022] Thirdly, an embodiment of the present invention further provides a terminal device, which includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the computer program is executed by the processor, the steps of any one of the talent portrait construction methods based on multimodal artificial intelligence provided in the specification of the present invention are realized.

[0023] Fourthly, an embodiment of the present invention further provides a storage medium for computer-readable storage, which is characterized in that the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to realize the steps of any one of the talent portrait construction methods based on multimodal artificial intelligence provided in the specification of the present invention.

[0024] An embodiment of the present invention provides a method, device, equipment, and medium for constructing a talent portrait based on multimodal artificial intelligence. The method includes: analyzing the target resume of a user to obtain basic questions and target questions, collecting the first audio of the user under the basic questions, and collecting the second audio and relevant videos of the user under the target questions, thereby providing support for establishing a user portrait from multiple dimensions such as language expression, voice characteristics, and body movements. Then, performing speech analysis on the first audio to obtain initial speech characteristics, and performing speech recognition on the second audio to obtain initial text information; generating a third audio of the target question according to the initial speech characteristics and the initial text information; determining the first reliability of the user under the target question according to the second audio, the third audio, and the relevant videos, that is, determining the corresponding tone or intonation information of the user when they are more confident through the basic questions, and applying it to the audio generation of the target question. By comparing the differences between the second audio and the third audio, and combining the expression information in the relevant videos, the first reliability of the user under the target question can be determined. Then, performing key frame recognition on the relevant videos to obtain target actions, and analyzing the user's behavior according to the target actions to determine the second reliability of the user under the target question; determining the initial skill score of the user under the target question according to the target question and the initial text information; finally, determining the target skill score corresponding to the user under the target question according to the first reliability, the second reliability, and the initial skill score. The target skill score not only considers the user's knowledge and skill level, but also fully considers their reliability when answering questions, and can more comprehensively and accurately reflect the user's actual ability, thereby determining the corresponding target talent portrait of the user according to the target skill score and the target question. Through this multi-dimensional data collection, analysis, and processing method, the method can accurately adjust the initial skill score of the user under the target question, thereby obtaining an accurate target skill score. This not only provides strong support for constructing a high-quality target talent portrait, but also effectively solves the problem of low quality of talent portraits in related technologies, thereby effectively improving the quality of enterprise talent recruitment and reducing the waste of human resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 It is a schematic flowchart of a method for constructing a talent portrait based on multimodal artificial intelligence provided by an embodiment of the present invention;

[0027] Figure 2Schematic diagram of the module structure of a talent portrait construction device based on multimodal artificial intelligence provided by an embodiment of the present invention;

[0028] Figure 3 Schematic block diagram of the structure of a terminal device provided by an embodiment of the present invention. Detailed implementation manners

[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all contents and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may be changed according to the actual situation.

[0031] It should be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0032] An embodiment of the present invention provides a method, device, equipment, and medium for constructing a talent portrait based on multimodal artificial intelligence. Among them, the method for constructing a talent portrait based on multimodal artificial intelligence can be applied to a terminal device, and the terminal device can be an electronic device such as a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device. The terminal device can be a server or a server cluster.

[0033] Next, some embodiments of the present invention will be described in detail in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0034] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for constructing a talent portrait based on multimodal artificial intelligence provided by an embodiment of the present invention.

[0035] As Figure 1 shown, the method for constructing a talent portrait based on multimodal artificial intelligence includes steps S101 to S108.

[0036] Step S101: Analyze the user's target resume to obtain basic questions and target questions, collect the first audio of the user under the basic questions, and collect the second audio and relevant videos of the user under the target questions.

[0037] Exemplarily, the user is a candidate for corporate recruitment or a person who needs to be interviewed when submitting a target resume to a company. Then, based on natural language understanding, analyze the user's target resume to obtain the user's basic information and the corresponding professional skill information of the user under professional skills. Among them, the basic information is used to represent information such as the user's corresponding age, native place, school, etc., and the professional skill information is used to represent the professional skills possessed by the user, such as tools that can be used or functions that can be achieved, etc.

[0038] Exemplarily, write corresponding basic questions according to the user's basic information. For example, questions such as "Can you briefly introduce your basic situation, including but not limited to your name, age, and current city of residence?" can be written, and use models such as chatgpt to generate the user's corresponding target questions based on the user's professional skill information. That is, the basic questions are questions generated based on the user's basic situation, and the target questions are questions generated based on the user's professional skills.

[0039] Exemplarily, at the beginning of the interview, the interviewer clearly states the basic questions to be asked and turns on the audio collection device to obtain the first audio corresponding to the user when answering the basic questions. When entering the target question session, turn on the audio and video collection devices at the same time, and the audio and video devices should run synchronously to ensure that the collected audio and video are of the same time, so as to collect the second audio corresponding to the user when answering the target questions and the relevant videos synchronized with the second audio when answering the target questions.

[0040] Step S102: Perform speech analysis on the first audio to obtain initial speech features, and perform speech recognition on the second audio to obtain initial text information.

[0041] Exemplarily, with the help of professional speech analysis software such as Praat, Adobe Audition, etc., extract the acoustic features corresponding to the user from the first audio, and count the total number of words and the total duration in the first audio, and calculate the speech rate. The speed of the speech rate can reflect the thinking agility and emotional state of the speaker, and analyze the ups and downs of the pitch in the first audio, etc., so as to determine the initial speech features corresponding to the user under the basic questions based on the acoustic features, speech rate information, and pitch ups and downs, etc.

[0042] Exemplarily, use a speech recognition model based on DeepSpeech or Transformer-ASR to perform speech recognition on the second audio, so as to obtain the initial text information corresponding to the user under the target questions.

[0043] Step S103: Generate speech based on the initial speech features and the initial text information to obtain a third audio of the target question.

[0044] Exemplarily, determine a speech generation model such as Tacotron, WaveNet, FastSpeech, etc., and then preprocess the initial text information, including word segmentation, part-of-speech tagging, adding prosody marks, etc. According to the requirements of speech generation, adjust the prosody of the initial text information, such as determining pauses, high and low changes in intonation, etc., to generate a more natural and fluent speech. Then, perform necessary conversion and normalization processing on the initial speech features to make them meet the input requirements of the selected speech generation model. For example, adjust the dimension and range of the features to ensure that the model can correctly process these features, and extract key style features from the initial speech features, such as pitch, timbre, speech rate, etc., in order to simulate the user's style during the speech generation process.

[0045] Exemplarily, input the processed initial text information into the text encoder of the speech generation model. The model will analyze and encode the text, extract the semantic and prosody information of the text, and then combine the initial speech features. The model predicts the corresponding acoustic features, such as Mel spectrogram, fundamental frequency, etc. These acoustic features describe the acoustic characteristics of the speech and are the basis for generating the speech waveform. Therefore, use the acoustic feature decoder to convert the predicted acoustic features into a speech waveform. In addition, during the synthesis process, use the previously extracted speaker style features to adjust the generated speech waveform to make it have a style similar to the first audio, so as to obtain a third audio in which the user has the same initial speech features when answering the target question as when answering the basic question.

[0046] Step S104: Determine the first reliability of the user under the target question according to the second audio, the third audio, and the relevant video.

[0047] Exemplarily, obtain the first waveform corresponding to the second audio and the second waveform corresponding to the third audio, calculate the distance information between the first waveform and the second waveform, and determine the difference value between the user's answers to the basic question and the target question according to the distance information.

[0048] Exemplarily, perform emotion recognition on the third audio according to the emotion recognition model to obtain the emotion type corresponding to the user when answering the target question. When the emotion type is a positive type such as pleasure or excitement, determine the first sub-reliability calculated by the user under the third audio as the basic reliability plus the difference value between the basic question and the target question. When the emotion type is a negative type such as tension or sadness, determine the first sub-reliability calculated by the user under the third audio as the basic reliability minus the difference value between the basic question and the target question.

[0049] Exemplarily, facial recognition is performed on the relevant video to obtain the facial information corresponding to the user, and then the expression type recognition is performed on the facial information of the user to obtain the expression type sequence corresponding to the user. And the facial expression change characterization value corresponding to the user when answering the target question is calculated according to the entropy value of the expression type sequence. Furthermore, the first quantity corresponding to the positive type when the user answers the target question and the second quantity corresponding to the negative type when the user answers the target question are obtained according to the expression type sequence. Thus, the second sub-reliability corresponding to the user when answering the target question is determined by combining the first quantity and the second quantity with the facial expression change characterization value. Furthermore, the first sub-reliability and the second sub-reliability are fused to determine the first reliability of the user under the target question.

[0050] In some embodiments, the determining the first reliability of the user under the target question according to the second audio, the third audio, and the relevant video includes: performing discrete processing on the second audio to obtain first discrete data, and performing discrete processing on the third audio to obtain second discrete data; calculating the distance between the first discrete data and the second discrete data to obtain the target distance information between the first discrete data and the second discrete data; constructing a target matrix for the second audio and the third audio according to the target distance information; aligning the data of the second audio and the third audio according to the target matrix to obtain an audio alignment result; determining the user audio difference between the second audio and the third audio according to the audio alignment result, and determining the first emotion information corresponding to the user and the first weight information corresponding to the first emotion information according to the user audio difference; obtaining the target expression information corresponding to the user from the relevant video, and performing expression analysis on the target expression information to obtain the second emotion information corresponding to the user and the second weight information corresponding to the second emotion information; fusing the first emotion information and the second emotion information according to the first weight information and the second weight information to obtain the target emotion information of the user under the target question; determining the first reliability of the user under the target question according to the target emotion information.

[0051] Exemplarily, the second audio is sampled at a certain sampling interval to convert the audio signal into a series of discrete values to obtain the first discrete data. The selection of the sampling interval should be determined according to the characteristics of the audio. Generally speaking, a higher sampling frequency can retain more audio details. The third audio is processed in the same sampling manner as the second audio to obtain the second discrete data, ensuring that the two discrete data are comparable.

[0052] Exemplarily, distance calculation is performed on any two data between the first discrete data and the second discrete data according to distance calculation methods such as Euclidean distance and Manhattan distance to obtain target distance information.

[0053] Exemplarily, the first data quantity corresponding to the first discrete data and the second data quantity corresponding to the second discrete data are obtained, and then an initial matrix is constructed according to the first data quantity and the second data quantity. Each matrix value in the initial matrix is 0, and thus the target distance information is filled into the initial matrix according to the first position of the corresponding first discrete data and the second position of the second discrete data to obtain a target matrix.

[0054] Exemplarily, using the dynamic time warping algorithm, data alignment of the second audio and the third audio is performed according to the target matrix, and then the best matching path between the second audio and the third audio is found, so that the overall difference between the two audios on this path is the smallest, thereby obtaining an audio alignment result.

[0055] Exemplarily, by analyzing the audio alignment result, comparing the differences between the second audio and the third audio in terms of pitch, volume, timbre, speech rate, etc., determining the user audio difference between the second audio and the third audio, establishing a mapping relationship between audio features and emotions, judging the first emotion information corresponding to the user according to the user audio difference. For example, an accelerated speech rate and increased volume may indicate that the user is in an excited emotional state, and determining the first weight information corresponding to the first emotion information according to factors such as the significance of the audio difference and the importance of emotion judgment.

[0056] For example, when the user audio difference is that aspects such as the pitch, volume, timbre, and speech rate of the second audio are significantly higher than those of the third audio, it indicates that the user is more confident or has more confidence when answering the target question than when answering the basic question, that is, the first emotion information is significantly confident. When the user audio difference is that aspects such as the pitch, volume, timbre, and speech rate of the second audio are significantly lower than those of the third audio, it indicates that the user is less confident or has no information when answering the target question than when answering the basic question, that is, the first emotion information is lack of confidence.

[0057] Exemplarily, the initial text information is used to generate a third audio based on the initial speech features under the user's basic question, and thus the first emotion information corresponding to the second audio is determined with the third audio as a reference, so as to accurately measure the emotion information of the user when answering the target question.

[0058] Exemplarily, a sequence of facial images of the user is extracted from the relevant video, these images are processed to identify the facial expression features of the user, such as the degree of eyebrow raising, the curvature of the mouth corner, etc., to obtain the target expression information, and then the expression recognition model is used to analyze the target expression information to judge the second emotion information corresponding to the user, such as happy, sad, angry, etc. The second weight information corresponding to the second emotion information is determined according to factors such as the accuracy of expression recognition and the obviousness of expression features.

[0059] Exemplarily, according to the first weight information and the second weight information, the first emotion information and the second emotion information are weighted and fused. For example, the two emotion information are multiplied by their respective weights and then added together to obtain the target emotion information of the user under the target question, so that the information from both audio and expression aspects can be comprehensively considered to more accurately judge the emotion of the user.

[0060] Exemplarily, according to factors such as the clarity of the target emotion information and the consistency of the two information sources (audio and expression), the first reliability of the user under the target question is determined. If the emotions reflected by the audio and expression are highly consistent and the emotion features are significantly positive features, the first reliability is higher; otherwise, the first reliability is lower. Among them, the first reliability is used to represent the characterization value of the user's emotion when answering the target question.

[0061] Specifically, by combining the information from both audio and expression aspects to judge the emotion of the user, compared with a single information source, it can more comprehensively and accurately understand the true emotion state of the user and reduce the misjudgment caused by inaccurate single information.

[0062] In some embodiments, calculating the target distance information between the first discrete data and the second discrete data includes: obtaining a first position corresponding to first data from the first discrete data and a second position corresponding to second data from the second discrete data; determining an adjacent position corresponding to the first data and the second data according to the first position and the second position; obtaining third data and fourth data corresponding to the adjacent position from the first discrete data and the second discrete data; calculating a first distance between the third data and the fourth data and obtaining a minimum distance corresponding to the first distance; calculating a second distance between the first data and the second data, and obtaining adjacent data corresponding to the first data from the third data and the fourth data; determining an adjustment parameter corresponding to the second distance according to the first data and the adjacent data in combination with an adjustment factor; determining the target distance information between the first discrete data at the first position and the second discrete data at the second position according to the adjustment parameter in combination with the minimum distance and the second distance; wherein, the target distance information is obtained according to the following formula:

[0063] ;

[0064] wherein, represents the target distance information between the first discrete data at the i-th first position and the second discrete data at the j-th second position, represents taking the minimum value, represents the first distance between the third data at the (i - 1)-th adjacent position corresponding to the i-th first position and the fourth data at the (j - 1)-th adjacent position corresponding to the j-th second position, represents the first distance between the third data at the (i - 1)-th adjacent position corresponding to the i-th first position and the second data at the j-th second position, represents the first distance between the first data at the i-th first position and the fourth data at the (j - 1)-th adjacent position corresponding to the j-th second position, and represents the adjustment factor, represents the second distance between the first data at the i-th first position and the second data at the j-th second position, represents the first discrete data at the (i - 1)-th first position, represents the first discrete data at the i-th first position, and abs represents taking the absolute value.

[0065] Exemplarily, for the first discrete data, each data element is associated with its position. This position can be the index of the data in the sequence, numbered sequentially from the start of the sequence. For example, if the first discrete data is an array, the position number of the first element of the array is 0, the second is 1, and so on. By traversing the first discrete data and recording the position corresponding to each data, the first position corresponding to the first data is obtained. The second discrete data is processed in the same way as the first discrete data. Each data element in the second discrete data is associated with its position in the sequence to obtain the second position corresponding to the second data.

[0066] Exemplarily, the first position and the second position are analyzed. According to the preset adjacent rule, the corresponding adjacent positions between the first data and the second data are determined. For example, if the adjacent rule is defined as the difference in position numbers being less than or equal to a certain threshold (such as 1), then the position numbers of the first position and the second position are compared to find the position pairs that satisfy this rule, and these position pairs are the adjacent positions.

[0067] Exemplarily, according to the determined adjacent positions, the data at the corresponding positions are extracted from the first discrete data and the second discrete data. The third data corresponding to the adjacent position is found in the first discrete data, and the fourth data corresponding to the adjacent position is found in the second discrete data.

[0068] Exemplarily, according to a distance metric method, such as Euclidean distance, Manhattan distance, etc., the distance between the third data and the fourth data is calculated to obtain the first distance. The first distances calculated for all adjacent positions are compared to find the minimum value, i.e., the minimum distance.

[0069] Exemplarily, using the same selected distance metric method, the distance between the first data and the second data is calculated to obtain the second distance. Among the third data and the fourth data, the adjacent data corresponding to the first data in the adjacent position relationship is found. A regulation factor is introduced, and this regulation factor can be set according to specific application scenarios and requirements to adjust the influence degree of the second distance.

[0070] Exemplarily, the first data and the adjacent data are comprehensively considered, combined with the regulation factor, and the regulation parameter corresponding to the second distance is determined through the following formula:

[0071]

[0072] where c represents the regulation parameter, and represents the said regulation factor, denotes the first discrete data corresponding to the (i - 1)-th first position and the first discrete data corresponding to the i-th first position, and abs represents taking the absolute value.

[0073] Exemplarily, the adjustment parameter, the minimum distance, and the second distance are comprehensively calculated according to the following formula to finally determine the target distance information between the first discrete data at the first position and the second discrete data at the second position. Among them, the target distance information is obtained according to the following formula:

[0074] ;

[0075] Among them, denotes the target distance information between the first discrete data at the i-th first position and the second discrete data at the j-th second position, represents taking the minimum value, denotes the first distance between the third data at the (i - 1)-th adjacent position corresponding to the i-th first position and the fourth data at the (j - 1)-th adjacent position corresponding to the j-th second position, denotes the first distance between the third data at the (i - 1)-th adjacent position corresponding to the i-th first position and the second data at the j-th second position, denotes the first distance between the first data at the i-th first position and the fourth data at the (j - 1)-th adjacent position corresponding to the j-th second position, and denotes the adjustment factor, denotes the second distance between the first data at the i-th first position and the second data at the j-th second position, denotes the first discrete data at the (i - 1)-th first position, denotes the first discrete data at the i-th first position, and abs represents taking the absolute value.

[0076] Exemplarily, by considering the data at adjacent positions and introducing the adjustment parameter, the relationship between two discrete data can be analyzed more comprehensively, avoiding the one-sidedness that may be brought by calculating the distance based on only a single data, so that the target distance information can more accurately reflect the true difference between the two discrete data. In addition, the introduction of the adjustment factor enables the distance calculation to be flexibly adjusted according to different application scenarios. In some scenarios, the weight of the adjustment factor can be increased to pay more attention to the influence of adjacent data; in other scenarios, the weight of the adjustment factor can be decreased so that the calculation result focuses more on the difference between the first data and the second data itself, improving the adaptability and flexibility of the method.

[0077] Step S105: Identify key frames from the relevant video to obtain the target action, and determine the second reliability of the user under the target question based on the target action.

[0078] Exemplarily, extract features from each frame image of the relevant video, such as color histograms, texture features, etc. Calculate the feature differences between adjacent frames. When the difference exceeds a pre-set threshold, mark this frame as a key frame. Further process the extracted key frames to extract features related to the action. For example, the position information of the key points of the human skeleton, and the coordinates of each joint point of the human body can be obtained through the human pose estimation algorithm. Motion trajectory features can also be extracted to record the movement path of the target object or person in the key frame.

[0079] Exemplarily, establish an action feature library containing feature templates of various common actions. Match the action features extracted from the key frames with the templates in the feature library, and use a suitable matching algorithm, such as the nearest neighbor algorithm, support vector machine, etc. to determine the target action. For example, determine whether the target action is any one of actions such as crossing the arms, touching the nose, etc.

[0080] Exemplarily, after identifying the target action, count the number of frames for each target action to obtain the number of frames corresponding to the target action. After counting the number of frames of all target actions, calculate the probability of occurrence of each target action. The calculation method of the probability is to divide the number of frames corresponding to each target action by the total number of frames of the relevant video. Assume that the video has a total of N frames, and the number of frames corresponding to a certain target action A is nA. Then the probability of action A occurring, PA = nA / N. Furthermore, obtain the target entropy corresponding to the target action in the relevant video based on the probability corresponding to the target action. When the target entropy is larger, it indicates that the user frequently changes the target action in the video. Because frequently changing actions often implies that the user is in a state of guilt or lack of confidence, which means that the user's performance is unstable and the reliability of their answer to the target question is relatively low. Thus, set the mapping relationship between the reliability and the target entropy. A corresponding table or functional relationship between the target entropy and the second reliability can be established based on experience or experimental data. For example, when the target entropy is in a certain higher interval, the corresponding second reliability is smaller; when the target entropy is in a lower interval, the corresponding second reliability is larger. Then, based on the calculated target entropy, determine the second reliability of the user under the target question through the established mapping relationship above.

[0081] In some embodiments, the key frame recognition of the relevant video to obtain the target action includes: obtaining the video frame sequence corresponding to the relevant video, and determining the coordinate information of the skeleton points corresponding to the user in each first sub-video frame in the video frame sequence; determining the initial clustering center according to the skeleton point coordinate information, and performing data clustering according to the initial clustering center combined with the skeleton point coordinate information to obtain an initial clustering result; extracting features of each second sub-video frame in each subclass cluster in the initial clustering result by using the skeleton information to obtain the skeleton point features corresponding to the second sub-video frame; constructing a skeleton feature matrix corresponding to the subclass cluster according to the skeleton point features, and obtaining a difference hash matrix corresponding to the subclass cluster by comparing the skeleton feature matrix; determining the target distance information corresponding to the second sub-video frames in the subclass cluster according to the difference hash matrix; adjusting the initial clustering result according to the target distance information to obtain a target clustering result corresponding to the first sub-video frame; obtaining a video label sequence corresponding to the video frame sequence by classifying the video frame sequence corresponding to the relevant video according to the target clustering result; determining the target video frame corresponding to the relevant video according to the video label sequence; and performing action recognition on the target video frame by using an action recognition model to obtain the target action corresponding to the target video frame.

[0082] Exemplarily, the relevant video is sampled at a fixed frame rate, such as 25 frames per second or 30 frames per second, and split into a series of consecutive video frames in chronological order to form a video frame sequence. For each first sub-video frame in the video frame sequence, a human pose estimation algorithm is used to identify the skeleton points corresponding to each key part of the user's body (such as the head, shoulders, elbows, wrists, hips, knees, ankles, etc.), and the coordinate information of these skeleton points in the video frame image is recorded, so as to obtain the coordinate information of the skeleton points corresponding to the user in each first sub-video frame.

[0083] Exemplarily, the initial clustering center is determined by randomly selecting or selecting based on the data distribution characteristics from the skeleton point coordinate information of all first sub-video frames. Then, based on the initial clustering center, a clustering algorithm such as K-means clustering is used to classify all first sub-video frames according to the similarity between their skeleton point coordinate information and the initial clustering center, and an initial clustering result is obtained. Each subclass cluster contains several similar first sub-video frames.

[0084] Exemplarily, for each second sub-video frame in each sub-class cluster of the initial clustering result, its skeletal point information is used for feature extraction. Features such as the relative distances and angles between skeletal points can be calculated, for example, calculating the distances between adjacent skeletal points, the bending angles of joints, etc., so as to obtain the skeletal point features corresponding to each second sub-video frame. Furthermore, the skeletal point features of all second sub-video frames in each sub-class cluster are combined to form the skeletal feature matrix corresponding to this sub-class cluster. Each row of the matrix can represent the skeletal point feature vector of a second sub-video frame. The skeletal feature matrix is processed, and by comparing the feature differences between different rows in the matrix, the difference information is converted into binary hash values to construct the difference hash matrix corresponding to the sub-class cluster.

[0085] Exemplarily, according to the difference hash matrix, the distances between the second sub-video frames in the sub-class cluster are calculated. For example, by calculating the Hamming distance between the hash values, etc., the corresponding target distance information between the second sub-video frames is obtained. Based on the target distance information, the initial clustering result is adjusted. If the target distances between some sub-video frames are too large and exceed a certain threshold, they may need to be reclassified into other sub-class clusters or form new sub-class clusters separately, and finally the target clustering result corresponding to the first sub-video frame is obtained.

[0086] Exemplarily, for the video frame sequence corresponding to the relevant video, according to the target clustering result, a class cluster label is assigned to each video frame to form the video label sequence corresponding to the video frame sequence. In this way, each video frame has a unique class cluster identifier.

[0087] Exemplarily, starting from the second element of the video label sequence, each element (i.e., the label of each video frame) is checked in turn. The label of the current video frame is compared with the previously recorded label: if the current label is different from the previous label, it means that the sequence number has changed. At this time, the current video frame is determined as the target video frame. If the current label is the same as the previous label, it means that the sequence number has not changed, and the label of the next video frame is continued to be checked, and the operation is repeated until the entire video label sequence is traversed.

[0088] Exemplarily, the action recognition model can be based on a neural network or deep learning. Furthermore, the trained action recognition model is used to recognize the actions of the target video frames, that is, the target video frames are input into the action recognition model. The action recognition model outputs the target actions corresponding to the target video frames according to the features and patterns it has learned, such as crossing the arms, touching the nose, etc.

[0089] Exemplarily, when obtaining the target action, relevant video frames with the same label as the target video frame can also be obtained according to the video label sequence, and then the action recognition model is used to classify the actions of the target video frame and the relevant video frames respectively to obtain the initial actions, and then data statistics are performed based on all the initial actions to obtain the target action corresponding to the target video frame.

[0090] Specifically, adjusting the initial clustering result based on the target distance information can further optimize the clustering effect, making the video frames within each subclass cluster more similar and the differences between different subclass clusters more obvious, which helps the action recognition model to more clearly distinguish different action categories.

[0091] In some embodiments, the determining the second reliability of the user under the target question according to the target action includes: determining the action type corresponding to the target action and the action score corresponding to the action type, and determining the first action quantity corresponding to the action type according to the video label sequence; determining the second action quantity corresponding to the target action according to the video label sequence; determining the target entropy information corresponding to the video label sequence according to the second action quantity, and calculating the autocorrelation coefficient according to the video label sequence to obtain the target correlation coefficient; determining the first degree of confusion of the user under the target question according to the target entropy information and the target correlation coefficient; determining the second degree of confusion of the user under the target question according to the action score and the first action quantity; and fusing the first degree of confusion and the second degree of confusion to determine the second reliability of the user under the target question.

[0092] Exemplarily, a large number of psychological literature, research reports and relevant professional materials are consulted to understand the typical action characteristics that the human body will show in different psychological states. For example, when a person is not confident, they may have actions such as lowering their head, folding their arms, and avoiding eye contact; while when confident, they will have actions such as raising their head and chest, having a firm look in their eyes, and natural and generous gestures. These actions corresponding to different psychological states are classified and sorted to form a comprehensive and clear action type library. The action type library can be divided according to psychological states, such as "unconfident action type", "confident action type", "nervous action type", "relaxed action type", etc., and then the relevant actions under different action types are obtained, so as to search for the target action in the action type library to obtain the action type corresponding to the target action.

[0093] Exemplarily, an appropriate action score is assigned to each action type according to expert experience or historical experience. For example, a relatively high action score is set for the "unconfident action type" and the "relaxed action type", and a relatively low action score is set for the "unconfident action type" and the "nervous action type".

[0094] Exemplarily, each label in the video label sequence corresponds to a video frame. Combining with the action type determined previously, count the number of occurrences of each action type in the video label sequence, and this number is the first action quantity corresponding to the action type. Similarly, based on the video label sequence, count the total number of occurrences of the target action, that is, the second action quantity. When traversing the video label sequence, as long as the target action is recognized, perform cumulative counting.

[0095] Exemplarily, divide the quantity of each target action in the second action quantity by the total sum of the second action quantity to obtain the probability of occurrence of each target action, and thus obtain the target entropy information corresponding to the video label sequence according to the calculation method of information entropy combined with the probability. The target entropy information reflects the degree of chaos of the target action distribution. The larger the entropy value, the more dispersed the target actions, indicating that the user's actions change more frequently.

[0096] Exemplarily, the autocorrelation coefficient is used to measure the correlation of the video label sequence itself at different time intervals. Obtain the target correlation coefficient by calculating the correlation of the video label sequence at different delays. For example, the correlation between adjacent video frame labels and the correlation between video frame labels separated by a certain number of frames can be calculated, and then the target correlation coefficient is obtained by synthesizing these correlation results. The larger the target correlation coefficient, the stronger the regularity of the video label sequence; conversely, the weaker the regularity. When the target correlation coefficient is smaller, it indicates that the user's actions change more frequently, and when the target correlation coefficient is larger, it indicates that the user's actions are relatively stable.

[0097] Exemplarily, comprehensively consider the target entropy information and the target correlation coefficient to determine the first degree of chaos. Different weights can be assigned to the target entropy information and the target correlation coefficient according to experience or experimental data, and then the two are weighted and summed to obtain the first degree of chaos. The larger the target entropy and the smaller the target correlation coefficient, the higher the first degree of chaos.

[0098] Exemplarily, determine the total score corresponding to the relevant video according to the action score and the first action quantity, and then determine the score standard deviation and score variance corresponding to the relevant video according to the total score and the action score, and then determine the second degree of chaos according to the score standard deviation and score variance.

[0099] Exemplarily, assign corresponding weights to the first degree of chaos and the second degree of chaos, and perform weighted summation on the two to obtain a comprehensive degree of chaos value. Then, according to the inverse relationship between the degree of chaos and the reliability, convert the comprehensive degree of chaos value into the second reliability. The higher the degree of chaos, the lower the second reliability.

[0100] Specifically, the first degree of chaos and the second degree of chaos respectively reflect the state of the user from aspects such as the distribution law of actions and the degree of abnormality of actions. When the user is in a state of guilt or lack of confidence, these chaos degree indicators will increase correspondingly, thereby reducing the second reliability and helping to detect the abnormal performance of the user in a timely manner.

[0101] Step S106: Determine the initial skill score of the user under the target problem according to the target problem and the initial text information.

[0102] Exemplarily, different skill levels and corresponding skill scores for the target problem are set according to expert experience or the experience of the R & D leader. Then, multiple first target answers corresponding to the target problem at different skill levels are generated by using ChatGPT. Next, the text similarity values between the initial text information and different first target answers are calculated. Then, the maximum value of the text similarity values is obtained, and this maximum value is compared with a preset value. When the maximum value is greater than the preset value, the second target answer corresponding to this maximum value is obtained. Then, the skill level corresponding to this second target answer is determined as the target skill level corresponding to the user, and the skill score corresponding to this target skill level is determined as the initial skill score of the user under the target problem.

[0103] In some embodiments, the determining the initial skill score of the user under the target problem according to the target problem and the initial text information includes: generating multiple initial relevant texts corresponding to the target problem according to the target problem in combination with a large model; obtaining a first text vector by encoding the initial text information using the text encoding layer of the quality classification model; respectively encoding the initial relevant texts using the text encoding layer of the quality classification model to obtain multiple second text vectors; calculating the difference between the first text vector and the second text vectors to obtain a first difference vector between the first text vector and the second text vectors according to the global difference layer of the quality classification model; calculating the difference between the first text vector and the second text vectors to obtain a second difference vector between the first text vector and the second text vectors according to the content difference layer of the quality classification model; performing attention fusion on the first difference vector and multiple second difference vectors by the information fusion layer of the quality classification model to obtain a target difference vector; classifying the content according to the first text vector and the target difference vector by the information classification layer of the quality classification model to obtain the content quality type corresponding to the initial text information and the target probability corresponding to the content quality type; and determining the initial skill score of the user under the target problem according to the target probability and the content quality type.

[0104] Exemplarily, the target problem is input into a large model such as Chatgpt or deepseek. Utilizing the powerful language generation ability of the large model, multiple initial relevant texts related to the target problem are generated. Then, the initial text information is input into the text encoding layer of the quality classification model. The text encoding layer will digitally process the initial text information and convert it into a first text vector that can be understood and processed by a computer, and this vector can represent the semantic features of the initial text information. In the same way, each initial relevant text is sequentially input into the text encoding layer of the quality classification model, and multiple second text vectors are obtained respectively, and these vectors respectively represent the semantic features of the respective initial relevant texts.

[0105] Exemplarily, using the global difference layer of the quality classification model, the first text vector and each second text vector are compared and calculated. The global difference layer will analyze the differences between the two vectors from an overall perspective and obtain a first difference vector between the first text vector and each second text vector, and this vector reflects the global difference situation between the initial text information and the initial relevant texts. Using the content difference layer of the quality classification model, the first text vector and each second text vector are compared and calculated again. The content difference layer will focus on analyzing the differences in content details between the two vectors and obtain a second difference vector between the first text vector and each second text vector, and this vector reflects the differences in content details between the initial text information and the initial relevant texts.

[0106] Exemplarily, the first difference vector and multiple second difference vectors are input into the information fusion layer of the quality classification model. The information fusion layer will adopt an attention mechanism to perform weighted fusion according to the importance of each difference vector. The attention mechanism can highlight important difference information and suppress unimportant information, and finally obtain a comprehensive target difference vector. The first text vector and the target difference vector are input into the information classification layer of the quality classification model. The information classification layer will classify the content of the initial text information according to the information contained in these two vectors. Through a classification algorithm, the content quality type corresponding to the initial text information is determined, and at the same time, the target probability corresponding to this content quality type is calculated, and this target probability represents the likelihood that the initial text information belongs to this content quality type. Among them, the content quality types include but are not limited to excellent, good, and poor.

[0107] Exemplarily, from all the calculated target probabilities, the one with the largest value is found. This largest target probability means that the text content is most likely to belong to the content quality type it corresponds to.

[0108] Exemplarily, the score mapping table is carefully set in advance by experts in the relevant field. The experts assign corresponding scores to different content quality types. These scores represent the scores that users can obtain under the corresponding content quality type. For example, for the "high-quality" content quality type, the experts may set its score to 100 points; the "good" type corresponds to 80 points; the "average" type corresponds to 60 points; the "poor" type corresponds to 20 points. Thus, according to the content quality type determined previously, a query is made in the score mapping table to obtain the score corresponding to this content quality type. Taking the above example, when the content quality type is determined to be "average", the corresponding score of 60 points is found from the score mapping table. Then, the score obtained by querying the content quality type corresponding to this maximum target probability in the score mapping table is determined as the initial skill score of the user under the target question.

[0109] Exemplarily, in the score mapping table, the experts may also assign corresponding score ranges to different content quality types. Then, the content quality type corresponding to this maximum target probability is queried in the score mapping table to obtain the target score range. Next, the difference between the maximum value and the minimum value in the target score range is calculated. Then, the maximum target probability is multiplied by this difference and then added to the minimum value in the target score range to obtain the initial skill score of the user under the target question.

[0110] Step S107: Determine the target skill score according to the first reliability, the second reliability, and the initial skill score.

[0111] Exemplarily, the minimum value of the first reliability and the second reliability is obtained, that is, the minimum reliability is obtained. Then, the minimum reliability is normalized and converted to the range of 0 to 1 to obtain the normalized reliability.

[0112] Exemplarily, according to the initial text information, the content quality type corresponding to the user under the target question is determined. Then, based on the clear understanding of the R & D leader about the basic levels corresponding to different content quality types, the minimum score corresponding to each content quality type is set in advance. This minimum score represents the basic threshold of the user's performance under this quality type. For example, for the "excellent" quality type, the experts may think its minimum score is 80 points; the minimum score of the "good" quality type is 60 points; the minimum score of the "medium" quality type is 40 points, and so on.

[0113] Exemplarily, in order to more precisely measure the user's performance beyond the minimum level in this quality type, subtract the previously determined minimum score from the initial skill score, and the resulting difference is the additional score. The additional score reflects the extra score obtained by the user in this content quality type, reflecting the degree to which the user's performance exceeds the basic requirements. Then, multiply the additional score by the normalized reliability. This process is to adjust the additional score according to the reliability. If the normalized reliability is high, indicating that the evaluation result is relatively reliable, then the additional score will be retained to a large extent; conversely, if the normalized reliability is low, indicating that there is a certain degree of uncertainty in the evaluation result, the additional score will be correspondingly reduced. Then, add the adjusted additional score to the minimum score, and the resulting value is the target skill score corresponding to the user for the target question. This target skill score comprehensively considers the user's basic performance (minimum score) and the performance beyond the basic level (adjusted additional score), and also takes into account the reliability of the evaluation result, and can more accurately reflect the user's true skill level for the target question.

[0114] Step S108: Determine the target talent profile corresponding to the user according to the target skill score in combination with the target question.

[0115] Exemplarily, determine the target skills involved in different target questions, cluster the target questions according to the target skills to obtain all the questions involved in different target skills, and then obtain the target skill scores corresponding to all the questions in different target skills respectively. Then, take the average of the target skill scores to obtain the average skill score corresponding to this target skill. Next, compare the average skill score with the score ranges of different portrait labels of this target skill to obtain the target label corresponding to the user for this target skill. Finally, determine the target talent profile corresponding to the user according to all the target labels.

[0116] In some embodiments, the determining the target talent profile corresponding to the user according to the target skill score in combination with the target question includes: determining the initial difficulty level and question type corresponding to the target question; classifying the target question according to the initial difficulty level and the question type to obtain the question cluster corresponding to the question type at the initial difficulty level and the target difficulty level corresponding to the question cluster; determining the average skill score corresponding to the question cluster according to the target skill score; performing score fusion according to the average skill score and the target difficulty level to obtain the target score corresponding to the user for the question type; determining the target label corresponding to the user for the question type according to the target score, and determining the target talent profile corresponding to the user according to the target label.

[0117] Exemplarily, experts or experienced practitioners in the field of organization evaluate and classify the difficulty of the target problem based on various factors such as the depth of knowledge involved in the target problem, the complexity of skills required to solve the problem, and the requirements for relevant background knowledge, thereby determining the initial difficulty level corresponding to the target problem, and determining the problem type corresponding to the target problem according to the experts or experienced practitioners. The problem type is used to characterize the skill type to which the target problem belongs.

[0118] Exemplarily, the target problem is classified according to the initial difficulty level and the problem type to obtain the problem cluster corresponding to the same problem type at the same initial difficulty level and the target difficulty level corresponding to the problem cluster. A problem cluster refers to a group of problems with the same skill type and the same difficulty level.

[0119] Exemplarily, the relevant skill scores corresponding to each sub-problem in the problem cluster are obtained from the target skill scores, and then the average skill score corresponding to the problem cluster is calculated by taking the mean of all the relevant skill scores in the problem cluster. This average skill score reflects the overall performance of the user's skill level under this problem cluster.

[0120] Exemplarily, all the problem clusters under the same problem type are obtained, and then the average skill scores corresponding to different problem clusters under the same problem type are obtained. Next, the target difficulty levels corresponding to different problem clusters under the same problem type are obtained. Then, considering the average skill score and the target difficulty level comprehensively, different weights are assigned to the average skill score according to the level of the target difficulty level. For example, for a problem cluster with a high difficulty level, a higher weight can be given to highlight the challenge of this problem cluster to the user's skill level. Then, through methods such as weighted calculation, the average skill score and the target difficulty level are integrated to obtain the target score corresponding to the user under this problem type. This target score can more comprehensively reflect the user's actual ability and performance in this problem type.

[0121] Exemplarily, according to the pre-set score intervals and corresponding labels, the user's target score is compared with these intervals to determine the target label corresponding to the user under this problem type. For example, if the problem type is java development, it is set that a score above 80 is an "excellent java developer", 60 - 80 is a "good java developer", and below 60 is a "java developer to be improved". Thus, when the target score is greater than 80, the target label corresponding to the user under the problem type is "excellent java developer", and when the target score is less than 80 and greater than 60, the target label corresponding to the user under the problem type is "good java developer", and so on.

[0122] Exemplarily, by combining the target tags of the user under multiple question types, all the target tags are jointly determined as the target talent profile corresponding to the user. This target talent profile can intuitively display the user's ability characteristics and development potential, providing a reference basis for further cultivation, selection, etc.

[0123] Please refer to Figure 2 , Figure 2 FIG. 200 is a talent profile construction device based on multi-modal artificial intelligence provided by an embodiment of the present application. The talent profile construction device 200 based on multi-modal artificial intelligence includes a data acquisition module 201, a speech analysis module 202, a speech generation module 203, a first analysis module 204, a second analysis module 205, a third analysis module 206, a data fusion module 207, and a profile determination module 208. Among them, the data acquisition module 201 is configured to analyze the target resume of the user to obtain basic questions and target questions, collect the first audio of the user under the basic questions, and collect the second audio and related videos of the user under the target questions; the speech analysis module 202 is configured to perform speech analysis on the first audio to obtain initial speech features, and perform speech recognition on the second audio to obtain initial text information; the speech generation module 203 is configured to perform speech generation based on the initial speech features and the initial text information to obtain the third audio of the target question; the first analysis module 204 is configured to determine the first reliability of the user under the target question according to the second audio, the third audio, and the related videos; the second analysis module 205 is configured to perform key frame recognition on the related videos to obtain target actions, and determine the second reliability of the user under the target question according to the target actions; the third analysis module 206 is configured to determine the initial skill score of the user under the target question according to the target question and the initial text information; the data fusion module 207 is configured to determine the target skill score according to the first reliability, the second reliability, and the initial skill score; the profile determination module 208 is configured to determine the target talent profile corresponding to the user according to the target skill score and the target question.

[0124] In some embodiments, the talent profile construction device 200 based on multi-modal artificial intelligence can be applied to a terminal device.

[0125] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described talent profile construction device 200 based on multi-modal artificial intelligence can refer to the corresponding process in the foregoing embodiment of the talent profile construction method based on multi-modal artificial intelligence, and will not be described herein again.

[0126] Please refer to Figure 3 ,Figure 3 A schematic block diagram of a terminal device provided by an embodiment of the present invention.

[0127] As Figure 3 shown, the terminal device 300 includes a processor 301 and a memory 302. The processor 301 and the memory 302 are connected through a bus 303, and this bus is, for example, an I2C (Inter - integrated Circuit) bus.

[0128] Specifically, the processor 301 is used to provide computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a Central Processing Unit (CPU), and this processor 301 can also be other general - purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field - Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general - purpose processor can be a microprocessor or this processor can also be any conventional processor, etc.

[0129] Specifically, the memory 302 can be a Flash chip, a Read - Only Memory (ROM), a magnetic disk, an optical disc, a USB flash drive, or a mobile hard disk, etc.

[0130] Those skilled in the art can understand that Figure 3 the structure shown in

[0131] is only a block diagram of a part of the structure related to the solution of the embodiment of the present invention, and does not constitute a limitation on the terminal device to which the solution of the embodiment of the present invention is applied. A specific server may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0132] In an embodiment, the processor is used to run a computer program stored in the memory and, when executing the computer program, implement the following steps:

[0133] Analyze the user's target resume to obtain basic questions and target questions, collect the first audio of the user under the basic questions, and collect the second audio and related videos of the user under the target questions;

[0134] Perform speech analysis on the first audio to obtain initial speech features, and perform speech recognition on the second audio to obtain initial text information;

[0135] Perform speech generation based on the initial speech features and the initial text information to obtain the third audio of the target question;

[0136] Determine the first reliability of the user under the target question according to the second audio, the third audio and the related videos;

[0137] Perform key frame recognition on the related videos to obtain target actions, and determine the second reliability of the user under the target question according to the target actions;

[0138] Determine the initial skill score of the user under the target question according to the target question and the initial text information;

[0139] Determine the target skill score according to the first reliability, the second reliability and the initial skill score;

[0140] Determine the target talent portrait corresponding to the user according to the target skill score and the target question.

[0141] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described terminal device can refer to the corresponding process in the embodiment of the method for constructing a talent portrait based on multimodal artificial intelligence described above, and will not be repeated here.

[0142] The embodiment of the present invention also provides a storage medium for computer-readable storage. The storage medium stores one or more programs, and one or more programs can be executed by one or more processors to implement the steps of any one of the methods for constructing a talent portrait based on multimodal artificial intelligence provided in the specification of the embodiment of the present invention.

[0143] Among them, the storage medium can be an internal storage unit of the terminal device in the foregoing embodiment, such as the hard disk or memory of the terminal device. The storage medium can also be an external storage device of the terminal device, such as a plug-in hard disk equipped on the terminal device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0144] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be executed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. The computer storage medium includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0145] It should be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this text, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article, or system. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article, or system comprising that element.

[0146] The serial numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments. The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for constructing a talent profile based on multimodal artificial intelligence, characterized in that The method includes: Analyzing the user's target resume to obtain basic questions and target questions, collecting the first audio of the user under the basic questions, and collecting the second audio and related video of the user under the target questions; Performing speech analysis on the first audio to obtain initial speech features, and performing speech recognition on the second audio to obtain initial text information; Performing speech generation based on the initial speech features and the initial text information to obtain the third audio of the target question; Determining the first reliability of the user under the target question according to the second audio, the third audio and the related video; Performing key frame recognition on the related video to obtain target actions, and determining the second reliability of the user under the target question according to the target actions; Determining the initial skill score of the user under the target question according to the target question and the initial text information; Determining the target skill score according to the first reliability, the second reliability and the initial skill score; Determining the target talent portrait corresponding to the user according to the target skill score and the target question.

2. The method according to claim 1, wherein, The determining the first reliability of the user under the target question according to the second audio, the third audio and the related video includes: Performing discrete processing on the second audio to obtain first discrete data, and performing discrete processing on the third audio to obtain second discrete data; Calculating the distance between the first discrete data and the second discrete data to obtain the target distance information between the first discrete data and the second discrete data; Constructing a distance matrix for the second audio and the third audio according to the target distance information to obtain a target matrix; Aligning the data of the second audio and the third audio according to the target matrix to obtain an audio alignment result; Determining the user audio difference between the second audio and the third audio according to the audio alignment result, and determining the first emotion information corresponding to the user and the first weight information corresponding to the first emotion information according to the user audio difference; Obtaining the target expression information corresponding to the user from the related video, and performing expression analysis on the target expression information to obtain the second emotion information corresponding to the user and the second weight information corresponding to the second emotion information; Fusing the first emotion information and the second emotion information according to the first weight information and the second weight information to obtain the target emotion information of the user under the target question; Determining the first reliability of the user under the target question according to the target emotion information.

3. The method according to claim 2, wherein The calculating the distance between the first discrete data and the second discrete data to obtain the target distance information between the first discrete data and the second discrete data includes: Obtaining the first position corresponding to the first data from the first discrete data and obtaining the second position corresponding to the second data from the second discrete data; Determining the adjacent position corresponding to the first data and the second data according to the first position and the second position; Obtain the corresponding third data and fourth data at the adjacent positions from the first discrete data and the second discrete data; Calculate the first distance between the third data and the fourth data, and obtain the minimum distance corresponding to the first distance; Calculate the second distance between the first data and the second data, and obtain the adjacent data corresponding to the first data from the third data and the fourth data; Determine the adjustment parameter corresponding to the second distance according to the first data, the adjacent data and the adjustment factor; Determine the target distance information between the first discrete data at the first position and the second discrete data at the second position according to the adjustment parameter, the minimum distance and the second distance; Wherein, the target distance information is obtained according to the following formula: ; Among them, represents the target distance information between the first discrete data at the i-th first position and the second discrete data at the j-th second position, represents taking the minimum value, represents the first distance between the third data at the (i - 1)-th adjacent position corresponding to the i-th first position and the fourth data at the (j - 1)-th adjacent position corresponding to the j-th second position, represents the first distance between the third data at the (i - 1)-th adjacent position corresponding to the i-th first position and the second data at the j-th second position, represents the first distance between the first data at the i-th first position and the fourth data at the (j - 1)-th adjacent position corresponding to the j-th second position, and represents the adjustment factor, represents the second distance between the first data at the i-th first position and the second data at the j-th second position, represents the first discrete data at the (i - 1)-th first position, represents the first discrete data at the i-th first position, and abs represents taking the absolute value.

4. The method according to claim 1, characterized in that, The key frame recognition of the relevant video to obtain the target action includes: Obtain the video frame sequence corresponding to the relevant video, and determine the coordinate information of the skeleton points corresponding to the user in each first sub-video frame in the video frame sequence; Determine the initial clustering center according to the skeleton point coordinate information, and perform data clustering according to the initial clustering center and the skeleton point coordinate information to obtain the initial clustering result; Extract features of each second sub-video frame in each sub-cluster of the initial clustering result by using the skeleton point information to obtain the skeleton point features corresponding to the second sub-video frame; Construct the skeleton feature matrix corresponding to the sub-cluster according to the skeleton point features, and perform comparison according to the skeleton feature matrix to obtain the difference hash matrix corresponding to the sub-cluster; Determine the target distance information corresponding to the second sub-video frames in the sub-cluster according to the difference hash matrix; Adjust the initial clustering result according to the target distance information to obtain the target clustering result corresponding to the first sub-video frame; Perform cluster labeling on the video frame sequence corresponding to the relevant video according to the target clustering result to obtain the video labeling sequence corresponding to the video frame sequence; Determine the target video frame corresponding to the relevant video according to the video labeling sequence; Perform action recognition on the target video frame by using the action recognition model to obtain the target action corresponding to the target video frame.

5. The method according to claim 4, wherein The determination of the second reliability of the user under the target question according to the target action includes: Determine the action type corresponding to the target action and the action score corresponding to the action type, and determine the first action quantity corresponding to the action type according to the video labeling sequence; Determine the second action quantity corresponding to the target action according to the video labeling sequence; Determine the target entropy information corresponding to the video labeling sequence according to the second action quantity, and calculate the autocorrelation coefficient according to the video labeling sequence to obtain the target correlation coefficient; Determine the first degree of confusion corresponding to the user under the target question according to the target entropy information and the target correlation coefficient; Determine the second degree of confusion corresponding to the user under the target question according to the action score and the first action quantity; Determine the second reliability of the user under the target question by integrating the first degree of confusion and the second degree of confusion.

6. The method according to claim 1, wherein The determining the initial skill score of the user under the target question according to the target question and the initial text information includes: Generate multiple initial relevant texts corresponding to the target question according to the target question in combination with a large model; Use the text encoding layer of the quality classification model to perform text encoding on the initial text information to obtain a first text vector; Use the text encoding layer of the quality classification model to perform text encoding on the initial relevant texts respectively to obtain multiple second text vectors; According to the global difference layer of the quality classification model, perform difference calculation on the first text vector and the second text vectors respectively to obtain a first difference vector between the first text vector and the second text vectors; According to the content difference layer of the quality classification model, perform difference calculation on the first text vector and the second text vectors respectively to obtain a second difference vector between the first text vector and the second text vectors; According to the information fusion layer of the quality classification model, perform attention fusion on the first difference vector and multiple second difference vectors to obtain a target difference vector; According to the information classification layer of the quality classification model, perform content classification on the first text vector and the target difference vector to obtain the content quality type corresponding to the initial text information and the target probability corresponding to the content quality type; Determine the initial skill score of the user under the target question according to the target probability and the content quality type.

7. The method according to claim 1, wherein The determining the target talent portrait corresponding to the user according to the target skill score in combination with the target question includes: Determine the initial difficulty level and question type corresponding to the target question; Perform question classification on the target question according to the initial difficulty level and the question type to obtain the question cluster corresponding to the question type at the initial difficulty level and the target difficulty level corresponding to the question cluster; Determine the average skill score corresponding to the question cluster according to the target skill score; Perform score fusion according to the average skill score and the target difficulty level to obtain the target score corresponding to the user under the question type; Determine the target label corresponding to the user under the question type according to the target score, and determine the target talent portrait corresponding to the user according to the target label.

8. A talent portrait construction device based on multimodal artificial intelligence, characterized in that, Includes: A data collection module, configured to analyze the target resume of the user to obtain basic questions and target questions, collect the first audio of the user under the basic questions, and collect the second audio and relevant videos of the user under the target questions; A speech analysis module, configured to perform speech analysis on the first audio to obtain initial speech features, and perform speech recognition on the second audio to obtain initial text information; A speech generation module, configured to perform speech generation according to the initial speech features and the initial text information to obtain a third audio of the target question; A first analysis module, configured to determine a first reliability of the user under the target question according to the second audio and the third audio in combination with the relevant video; A second analysis module, configured to perform key frame recognition on the relevant video to obtain a target action, and determine a second reliability of the user under the target question according to the target action; A third analysis module, configured to determine an initial skill score of the user under the target question according to the target question and the initial text information; A data fusion module, configured to determine a target skill score according to the first reliability and the second reliability in combination with the initial skill score; A portrait determination module, configured to determine a target talent portrait corresponding to the user according to the target skill score in combination with the target question; 9. A terminal device, characterized in that, The terminal device includes a processor and a memory; The memory is used to store a computer program; The processor is configured to execute the computer program and implement the method for constructing a talent portrait based on multimodal artificial intelligence according to any one of claims 1 to 7 when executing the computer program.

10. A computer storage medium for computer storage, characterized in that, The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the method for constructing a talent portrait based on multimodal artificial intelligence according to any one of claims 1 to 7.