COMPUTER SYSTEM AND METHOD FOR PERSONIFIED ANALYSIS FOR IN PARTICULAR CAREER PLANNING AND EMPLOYMENT PLACEMENT OF A PERSON
An AI-supported audio and video system addresses the limitations of predefined dialogue systems by analyzing multi-modal user inputs to provide personalized career advice and job recommendations, enhancing user interaction and reducing misinterpretations.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- JOBNINJA GMBH
- Filing Date
- 2024-11-15
- Publication Date
- 2026-05-21
AI Technical Summary
Current digital dialogue systems for career planning and job placement are limited by predefined interpretation routines, failing to effectively analyze and interpret user inputs beyond simple yes/no answers or keyword-based processing, thus lacking personalization and depth in career recommendations.
An AI-supported audio and video system that analyzes visual, auditory, and textual inputs to provide personalized career suggestions, using a digital avatar for interactive communication, capable of recognizing speech patterns, facial expressions, and document analysis to create comprehensive user profiles and tailored job recommendations.
Enables personalized career advice by analyzing user's strengths, weaknesses, and aspirations, providing detailed job recommendations and improving resume quality, while reducing misinterpretations through synchronized avatar responses.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for interactive communication between a natural person and a computer system.
[0002] While known digital dialogue systems are designed to interpret linguistic, acoustic, or textual inputs, interactive communication is limited in the current state of the art insofar as the interpretation analysis, as implemented by means of computer data analysis, follows predetermined, unchanging routines.
[0003] These dialogue systems, within the framework of predefined dialogue content, are tailored to predetermined interpretation routines. Predefined questions are posed to a user via computer infrastructure, formulated in such a way that only a limited and therefore small number of possible answers are possible. For example, in the simplest case, a question that can only be answered with a user's "yes" or "no".
[0004] The invention described herein focuses on the invented methodology using the example of a dialogue involving career planning and job placement. Currently, some job providers and job boards offer the option of video applications, where videos recorded by applicants can be forwarded directly to the employer. However, even here, the processing of the video content is limited to the aforementioned predefined interpretation routines. Furthermore, job boards increasingly offer so-called matching algorithms that find suitable job postings based on keywords in the uploaded resume. These methods are essentially based on keyword data processing.
[0005] Against this background, the object of the present invention is to provide an improved analysis of recorded personal data by means of intensified interaction between the person on the one hand and a computer-based system on the other, in particular with regard to a dialogue topic of career planning and job placement.
[0006] This problem is solved by a method having the features of claim 1 and by a system as defined in claim 6.
[0007] The invention discloses a method for interactive real-time communication between a natural person and an AI-supported computer application with an audio and video system, whereby, depending on the visual and verbal capture of personal data by means of appropriate sensors, an AI-supported analysis result is output by the AVS (audio and video computer system) following the interactive communication. Furthermore, the invention relates to an AI-supported audio and video communication computer system (AVS) that is configured for interactive communication via image and sound with a natural person.
[0008] The present invention essentially relates to a digital system for video-based analysis and evaluation of data, particularly for career planning and job placement of a natural person, using an AI-supported algorithm. The system enables a user to interact with the computer system in natural language via video call with a digital avatar.
[0009] The computer system analyzes speech, video, and optionally text information from the user to provide suitable career recommendations and job offers. The system according to the invention, particularly for career planning, goes beyond purely data-based processing. Rather, it aims to enable personalized suggestions by analyzing the user's audiovisual input, taking into account their strengths, weaknesses, and career aspirations. The invention creates a comprehensive analytical capability that goes beyond the mere processing of text data. It combines visual, auditory, and textual input to provide a better basis for decisions, especially regarding career suggestions.
[0010] In other words, the present invention provides a digital "headhunter" that offers personalized advice to the user through video-based interaction with an avatar. The system can advantageously analyze the user's speech with regard to pitch, fluctuations, and pauses in order to draw conclusions about the user's confidence and persuasiveness. Simultaneously, the user's facial expressions and gestures can be analyzed through visual recognition to obtain a comprehensive picture of the individual's personality and communication skills.
[0011] Advantageously, in addition to the above-mentioned auditory and visual recording, the system can also capture textual input from the user via a suitable interface, for example a keyboard.
[0012] Furthermore, documents can optionally be digitally imported or uploaded to support the analysis. The avatar, assisted by the computer-implemented system, prompts the user to upload data sets. This prompt is then followed by uploading documents, forms, or data sets. Advantageously, the method then includes an analysis of the uploaded documents to identify spelling errors, grammatical errors, and / or inconsistencies in the text. In an advantageous embodiment, the method can also include a comparison of the data from an uploaded data set with data acquired from the system's sensors.
[0013] Based on this data, suitable career suggestions, missing qualifications, and relevant job postings are identified. These suggestions may also include improving the resume, participating in application training, or attending further education courses or seminars. Ideally, a suitable job offer can be made.
[0014] Ultimately, the invention creates a digital career assistant that, through video-based interaction with an avatar, can provide users with valuable career recommendations for their future professional development. Based on these analyses, recommendations for job postings can also be generated directly.
[0015] The avatar's response is output via at least one output device in the form of words and images, whereby the avatar's response is achieved by synchronizing an animated representation of a specific person with a synthesized speech response using that person's voice. For example, the avatar can visually represent a person from a company's human resources department with whom a job interview might take place.
[0016] The avatar's movement, design, and / or environment are linked to the information output by the speech system. This defined interlinking ensures that the information is always conveyed clearly and supported visually. Misinterpretations of the speech system's response, or even a complete lack of understanding, are thus prevented. The avatar's movements, design, and / or environment are automatically synchronized with the response.Thus, the situation is always such that the avatar behaves correctly, presents itself correctly, and is displayed in the correct environment, thereby visually supporting the content of the information to be transmitted with precision, and thus virtually eliminating any misinterpretations regarding the perceptibility and interpretation of the information to be transmitted.
[0017] Preferably, the avatar's movements, design, and / or visually represented environment are controlled by the information output by the speech system. This means that the information forms the fundamental basis for how the avatar moves, how its environment is displayed, or how the avatar itself is designed. Through this control, the avatar's movements, design, and / or environment representation are precisely and precisely tailored to the intended message, based on the information being transmitted and its content. An emotional state conveyed through the information can then be communicated more effectively by the avatar.
[0018] In particular, the correct phonetic pronunciation of words is stored in the speech system, and the avatar's movement, design, and / or visually represented environment are linked to this correct pronunciation. This allows for an extremely realistic representation of an action or speech, significantly increasing the user's attention to the information presented by the speech system in the form of a response.
[0019] Preferably, the linking is carried out automatically depending on an input recognized by the language system.
[0020] It can also be provided that the linking is dependent on a setting made by the user. In this context, it can be provided that, for example, the user can set the specific national language in which they wish to receive the response from the language system. Thus, the output can be set to German, English, or similar languages. The phonetic pronunciations underlying the words of the specific national language are then automatically linked to the precise movement of the avatar or a body part of the avatar, or to its design, or to the visually represented environment of the avatar.
[0021] The digital "career assistant" analyzes the user's communication on several levels: linguistic characteristics such as voice fluctuations and intelligibility, visual features such as facial expressions and gestures, as well as textual information from an uploaded resume or a description of professional experience. Based on this analysis, the system can: i) Identify the user's strengths and weaknesses, ii) identify missing qualifications for certain professions, iii) suggest suitable job advertisements, and / or iv) optimize the career path through targeted recommendations.
[0022] The method according to the invention is started by activating an application (app) on a computer-based gadget, handheld, tower, laptop, mobile phone, smartphone, or the like. Advantageously, by installing the application on their own computer, the user can use their own computer to apply the method. This has the advantage that communication can take place in the user's familiar environment, allowing the user to make their inputs in an intuitive and familiar format.
[0023] The application software can be stored on the computer's internal memory or be cloud-based software that can be accessed via a website on the internet or that obtains data segments from a computer network to execute the procedure.
[0024] In an alternative embodiment, the method according to the invention is executed without application software being installed on the computer itself, in a kind of "thin client" or "cloud computing" model. The computer boots a minimal operating system or a special program (client) capable of establishing network connections. The computer connects to the internet or a network server where the required software (application) is provided. The client requests the necessary software via the network. This software is then streamed to the computer or executed remotely. The actual computing power and software execution take place on the server, while the local computer primarily serves for input and display. This model reduces the need for local computing power and storage, since the main processing is performed on an external server.
[0025] To be able to access all results of the interactive communication and personalized data analysis at any later time, it is advantageous for the user to register with the service provider who provides the application. During this registration process, some personalized data may be requested, which can also be useful for later analysis.
[0026] The process can proceed in several steps. Depending on the context of career planning and / or job searching, the following content segments can be implemented during the application process: 1. Avatar greeting: Upon launching the app or website, the user is greeted by an avatar. The avatar introduces itself and asks if the user is looking for a job and if they would like to upload a completed resume. 2. Recording of professional history: If the user does not provide a resume, they will be asked to describe their professional history to date. The user begins with their most recent education or studies and details their previous positions and responsibilities. 3. Speech and Video Analysis: The user's responses are recorded via the system's video and audio input. Voice analysis includes the identification of fluctuations, pauses, tremors in the voice, and other indicators of uncertainty or decisiveness. Visual analysis captures the user's facial expressions and gestures to create a coherent picture. This includes, for example, checking whether the facial expression matches what was said and whether the person speaks a lot or a little. Any texts uploaded by the person are analyzed for spelling errors, logic / comprehensibility of their statements, sentence structure, sentence length, hierarchical or chronological structure. 4. Creating a user profile: Based on the analyzed data, the system creates a detailed user profile that includes their strengths, weaknesses, and career history. Using this information, the system generates suggestions for suitable job offers and provides recommendations for necessary qualifications. 5. (Optional) Feedback loops: The user receives suggestions and has the opportunity to rate and adjust them. If an application is submitted, there is a feedback loop from the employer. This information is used to improve the algorithm and further optimize the suggestions.
[0027] A communication system according to the invention comprises a digital speech and video system designed for communication with a natural person in words and images. The digital communication system, which interacts with a natural person, is configured to output a response based on the person's voice input. The speech system is designed to process acoustic and visual input from the person and, optionally, to process at least one other piece of information distinct from it. The information presentation and processing capabilities of the communication system can be significantly enhanced by allowing documents to be read in and / or instructions to be entered into the system via keyboard. The system includes the following technical aspects, which enable in-depth analysis and interaction with the user: 1. Natural speech input and output via suitable auditory sensors: The system is capable of recognizing, transcribing, and analyzing the user's speech. This analysis takes place in real time during the video call with the avatar. 2. Visual analysis via suitable camera-supported video sensors: The system captures images of the user via the camera, with the video algorithm recognizing and analyzing visual features such as facial expressions. The algorithm can also analyze the congruence of facial expressions with spoken words. For example, inappropriate gestures and facial expressions can be interpreted as potential indicators of uncertainty. 3. Text analysis: In addition to voice and video input, the system also analyzes uploaded or entered textual information such as resumes or keyboard input. The data is checked for logical consistency and potential areas for improvement.
[0028] Furthermore, in an advantageous implementation, the system can utilize a dual feedback loop in which the user evaluates the system's recommendations and provides feedback to potential employers regarding applications. This feedback is used to optimize the system's algorithms.
[0029] The algorithm described below is suitable for a system and procedure for generating and controlling an interactive avatar that enables natural communication with a human person: The algorithm consists of several coordinated modules that enable the acquisition, analysis, and synthesis of input data, as well as the output of corresponding reactions by the avatar. The goal of the algorithm is to to realize the most natural and context-related interaction possible between a natural person and a virtual avatar. 1. Data acquisition unit (Module A)
[0030] The data acquisition unit comprises various sensors and input methods used to record communication signals from the individual. These recorded communication signals include, among others: • Voice input via microphones, • Image and video recordings by cameras for the recognition of facial expressions, gestures and emotional expressions, • Text-based input via a keyboard unit or other writing input devices.
[0031] The acquisition unit is equipped with a preprocessing component that checks the quality of the recorded data and applies noise reduction or filtering algorithms if necessary. 2. Analysis module (Module B)
[0032] The analysis module is responsible for processing the data collected by the data acquisition unit. This module consists of several sub-steps, which are described in detail below: • Speech recognition and transcription: Using a trained neural network for speech recognition (ASR - Automatic Speech Recognition), spoken input is converted into text. Speech characteristics such as pitch, volume, and speaking rate are recognized to infer the emotional state of the individual. • Emotion recognition: Based on speech analysis and facial recognition data (facial expressions), an emotional analysis is performed. Advantageously, a convolutional neural network (CNN) and recurrent neural networks (RNNs) are used to extract emotional features from image, video, and audio data. • Context analysis: To ensure context-related communication, the transcribed text undergoes semantic analysis. A transformer model based on the BERT approach (Bidirectional Encoder Representations from Transformers) is used for this purpose, determining the meaning of the captured input. The model considers historical interactions and contextual information to enable a consistent response from the avatar. 3. Reaction generation (Module C)
[0033] Based on the analysis results, the reaction generation module creates a suitable reaction for the avatar. This is done in the following steps: • Generation of the linguistic response: A language-generating neural network, for example a GPT model (Generative Pre-trained Transformer), creates an appropriate textual response based on semantic and emotional analysis. • Synthesis of facial expressions and gestures: Parallel to the verbal response, an animation of the avatar is generated, which is tailored to the analyzed emotional state of the natural person. Here, the avatar's gestures and facial expressions are calculated using a model for movement and emotion synthesis (e.g., GAN - Generative Adversarial Network). • Speech synthesis: The textual response is converted into a spoken output using a text-to-speech model (e.g., Wavenet). This model takes into account prosodic elements such as stress, intonation, and rhythm to enhance the emotional component of the response. 4. Output module (Module D)
[0034] The output module includes all components necessary for displaying and playing back the avatar, including a display for visual representation and speakers for audio output. The animation and speech synthesis generated by the reaction generation module are output synchronously to ensure consistent and natural communication.
[0035] Example algorithm flow: a) Capture of input data (Module A): • Capture of speech via a microphone. • Capturing facial expressions using a camera. b) Conducting an analysis (Module B): • Transcription of the language using the ASR model. • Recognition of emotions through analysis of facial expressions and voice. • Performing a semantic analysis of the transcribed text using a transformer model. c) Generating a reaction (Module C): • Generation of a text-based response using a GPT model. • Synthesis of a suitable animation of the avatar (facial expressions and gestures) based on the recognized emotion. • Conversion of the text-based response into spoken language using a text-to-speech model. d) Output of the reaction (module D): • Synchronized output of the voice response and the animated avatar reaction.
[0036] The algorithm is implemented through a combination of machine learning models and signal processing modules. The neural networks (e.g., CNN, RNN, Transformer) are trained using a comprehensive training dataset to ensure accurate recognition of speech, emotion, and context. Furthermore, real-time optimizations are performed to enable lag-free interaction between the human user and the avatar.
[0037] To adapt the described algorithm to the specific application of a career planning or job search dialogue, additional modules and functions are integrated, tailored to the particular requirements of this domain. The focus here is on the recognition, processing, and synthesis of career-specific information and individualized advice.
[0038] A key enhancement is the integration of a comprehensive knowledge database containing information on professions, career paths, further training options, qualifications, skills, market trends, salaries, and other relevant topics. This database is automatically updated regularly to ensure that the information provided reflects current labor market trends and requirements.
[0039] The system according to the invention further uses ontologies or taxonomic models to link professional skills, qualifications, and roles. These models help to understand the relationship between skills, required qualifications, and possible career steps for different professions. Specific models or vector space analyses are used to determine which skills are essential for which professions.
[0040] An extension of the analysis module includes the processing of input data relating to the professional history and qualifications of the individual. This can be achieved using a semantic analysis model that analyzes the resume and professional experience, identifying strengths, weaknesses, and potential development opportunities. The model utilizes NLP (Natural Language Processing) techniques to process work-related texts.
[0041] The system also recognizes explicitly stated professional interests, as well as subtle preferences in the natural person's speech input. For this purpose, special sentiment analysis and preference models are used, which correctly interpret statements such as "I am interested in the IT industry" or "I would like to take on more responsibility."
[0042] The system analyzes the individual's short- and long-term career goals, including desired positions, industries, salary expectations, and work-life balance preferences. It uses an extended semantic model based on contextual goal analysis.
[0043] Furthermore, the system generates detailed recommendations tailored to the individual's professional history and preferences. To do this, the model draws on the knowledge database and the professional profile. The model provides recommendations on: • Further training options and necessary qualifications. • Potential career steps and advancement opportunities. • Job search tips such as resume design, interview preparation and networking opportunities.
[0044] In addition to simply generating a response, the system can create a complete development plan that defines the steps for achieving the goal. These include: • Time estimates for developing specific skills. • Suggestions for professional networks and mentoring programs. • Recommendations for specific industries or companies based on the preferences and abilities of the individual.
[0045] The system can also ask targeted questions to obtain more information about the individual's career aspirations. These questions are based on a decision tree optimized for career counseling. For example, the avatar might ask: "Would you like to learn more about further training opportunities in the IT industry?" or "Are you interested in management positions?"
[0046] A specific module for simulating job interviews can be added. The avatar can act as a virtual interviewer to provide a realistic interview experience. Typical interview questions are asked, and the system provides feedback and suggestions for improvement based on the answers.
[0047] As part of the technical adjustments to ensure the timeliness of career-specific information, the system can regularly update publicly available job boards, industry reports, and salary surveys using web scraping techniques. This makes it possible to provide real-time information during consultations.
[0048] Optionally, the system continuously learns through interactions with the individual. Using a reinforcement learning approach, it dynamically adapts its recommendations, questions, and responses to the individual's preferences and progress.
[0049] The following is an example of an algorithmic process for career planning: 1. Recording of professional history and interests through targeted questions and analysis of the resume. 2. Analysis of skills, qualifications, and professional preferences. 3. Generating recommendations for possible career paths, including necessary further training and qualifications. 4. Creation of a development plan that includes specific measures and time estimates. 5. Simulation of job interviews to prepare for specific job opportunities.
[0050] The invention is explained below with reference to an exemplary embodiment shown in the accompanying drawing, in which In Fig. 1. A system diagram of the digital career assistant is shown.
[0051] Step 10 begins with a greeting. The avatar introduces itself. Step 12 then asks initial questions, such as whether the user is looking for a job and whether they have a resume (curriculum vitae). If a resume is available, the user is asked to upload it in step 14. If no resume is available, the user is asked to briefly describe their professional background in step 16. This description is captured audibly and visually via appropriate sensors, and the digital data is then transmitted to the communication system.
[0052] In step 18, the person is asked to provide information about their most recent education or studies up to their current job. This description is also captured audibly and visually via sensors, with the digital data being transmitted to the communication system. In step 20, the captured data is then analyzed by the communication system. This analysis of the input data uses a neural network trained to recognize speech patterns, facial expressions, and contextual information.
[0053] Voice analysis includes the detection of the user's confidence, insecurity, pauses, and voice fluctuations. Visual analysis considers facial expressions and the congruence between facial expressions and spoken words. During the communication phase between the person and the digital, AI-powered computer system, the avatar regularly responds based on the analyzed information.
[0054] In step 22, an interaction takes place between the avatar and the test subject, during which questions about abnormalities and possible inconsistencies are clarified.
[0055] Finally, in step 24, a reaction output is generated as a result of the conversation via the avatar by at least one output device in word and image, whereby the reaction of the avatar is achieved by synchronizing an animated representation of a specific person and a synthesized speech response with the voice of the specific person.
[0056] The analysis also results in the creation of a profile and the generation of suggestions for the user: Based on the analyzed data, the system creates a user profile. Depending on the user's career history, career goals, and the relevance of their existing qualifications, suitable job offers and qualification suggestions are generated.
[0057] Finally, in step 26, there is a double feedback loop where the user can evaluate and adjust the suggestions. If the user ultimately decides to apply, the employer provides feedback. This information is then used to optimize the algorithm.
Claims
A method for real-time communication between a natural person as user and a digital AI-supported computer system, in which the system can generate suggestions for job offers and career paths by collecting user data, wherein the system includes at least video and audio sensors for data collection, and performs an analysis of the collected data from the user, wherein the data is generated within the framework of interactive communication between the user and the computer system, and wherein the output of the analysis is provided using an avatar generated by the computer system, wherein the method comprises the following steps: - capturing at least voice data and image data from the user;- Analyzing the captured data to determine textual, semantic, and emotional information, wherein the analysis of the input data uses a neural network trained to recognize speech patterns, facial expressions, and contextual information; - Generating an avatar response based on the determined information; and - Outputting the avatar response in word and image through at least one output device, wherein the avatar response is achieved by synchronizing an animated representation of a specific person's face and a synthesized speech response with the specific person's voice as part of real-time communication with the user. Method according to claim 1, characterized in that, as further information for data collection, information entered or read in via an interface is recognized as input. Method according to one of the preceding claims, characterized in that the video analysis includes facial expressions, mimetic gestures and gestures of the user. Method according to one of the preceding claims, characterized in that the audio analysis includes voice characteristics such as voice fluctuations, pauses and intelligibility. Method according to one of the preceding claims, characterized in that the user is prompted to upload data records and this prompt is followed by uploading documents, forms, or data records. The method according to claim 5, characterized in that the method comprises an analysis of the uploaded documents in which spelling errors, expression errors, grammatical errors and / or inconsistencies in the text are identified. Method according to claim 5 or 6, characterized in that the method comprises a comparison of the data from an uploaded data set and data acquired from the system sensors. Method according to one of the preceding claims, characterized in that a double feedback loop is set up by the user and future employer. A communication system comprising a natural person as user and an AI-supported computer system designed to communicate with the person, generating real-time interaction between the person and the AI-supported computer system based on recorded auditory and visual data from the person and, if applicable, textual data about the person. This interaction is based on an analysis for content-related evaluations of the data, in particular for suggestions of job offers and career paths, according to a method of claims 1 to 8, wherein the system comprises a networkable computer and auditory and visual sensors by means of which the user's data is collected in real time.