Information processing system
Patent Information
- Application Number
- CN202610262502.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-05
- Publication Date
- 2026-09-22
AI Technical Summary
通过上述构成,本发明能够在标准化流程中自动执行对话解析、情感识别以及生成式人工智能分析,实现对考生人物特性和情感面的客观、系统、可重复的评估,有效解决现有技术中评价不全面、依赖人工经验及难以深入挖掘潜在特征的问题
服务器在获得分析用文本和结构化数据后,生成评价文档数据。服务器在一种实施方式中通过以下方式提高处理效率和精度:
Smart Images

Figure CN122796014A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot speech in response to the user's speech.
[0003] The problem this invention aims to solve is that current recruitment and selection examinations primarily rely on human interviews and written document review for candidate evaluation, which presents the following issues: First, relying solely on resumes and written answers makes it difficult to grasp candidates' personality traits, emotional states, and behavioral tendencies in a timely, objective, and systematic manner, resulting in an incomplete candidate profile. Second, the capture and analysis of multimodal information such as candidate speech content, facial expressions, and tone of voice in traditional interviews heavily depends on the interviewer's subjective experience, easily leading to evaluation bias and inconsistency. Third, there is a lack of repeatable and scalable methods for in-depth analysis of potential traits and abilities that are difficult to fully understand or quantify during the written screening stage. Fourth, existing generative artificial intelligence applications are mostly limited to simple question-and-answer or text generation, failing to effectively integrate with emotion recognition technology and lacking a mechanism for structured analysis of candidate characteristics and outputting standardized reports that can be used for recruitment decisions. Therefore, there is an urgent need for a system that can integrate dialogue content and emotion recognition results to automatically generate objective reports on candidate characteristics and emotional aspects, and further analyze aspects not clearly defined in the written screening process, thereby improving the efficiency and evaluation quality of the selection process. Summary of the Invention
[0004] To address the aforementioned issues, this invention provides an information processing system comprising a processor, an emotion recognition engine, and a generative artificial intelligence model. The processor is configured to conduct a dialogue between a candidate and the generative artificial intelligence model on a pre-defined topic, and to parse the dialogue content to obtain textual information reflecting the candidate's thinking patterns, communication style, and problem-solving process. The processor is further configured to utilize the emotion recognition engine to analyze multimodal information such as the candidate's speech, facial expressions, and tone of voice during the dialogue to estimate the candidate's emotional state, thereby obtaining an emotion recognition result. The processor is also configured to generate a report evaluating the candidate's characteristics and emotional aspects based on the dialogue parsing result and the emotion recognition result, outputting assessment information about the candidate's profile in a structured manner.
[0005] Furthermore, to delve deeper into key points not fully understood during the written screening stage, the processor is configured to generate and send prompt text requesting the generative AI model to perform additional analysis, enabling the model to supplement the candidate's characteristics and abilities from a specific perspective based on previous dialogue content and existing reports. Further, the processor is configured to evaluate the candidate's responses based on the prompt text, integrating the analysis results from the generative AI model into the corresponding characteristic and ability evaluation sections of the report, thereby obtaining a more refined, comparable, and easily usable comprehensive assessment result for recruitment personnel. Through these configurations, the present invention can automatically execute dialogue parsing, sentiment recognition, and generative AI analysis within a standardized process, achieving an objective, systematic, and repeatable assessment of the candidate's personality traits and emotional aspects, effectively solving the problems of incomplete evaluation, reliance on human experience, and difficulty in deeply uncovering potential characteristics in existing technologies.
[0006] "System" refers to a device or combination of devices consisting of hardware and / or software used to perform dialogue between a test taker and a generative artificial intelligence model, parse the dialogue content, perform emotion recognition, and generate an evaluation report.
[0007] A "processor" refers to a hardware component or equivalent device that can execute program instructions, perform calculations and logical control on input data, and realize functions such as dialogue management, data parsing, model calling, result integration and report generation, including but not limited to CPU, GPU, application-specific integrated circuit or combinations thereof.
[0008] An “emotion recognition engine” refers to a software or hardware module or combination thereof used to receive and analyze voice, image and / or text information related to test takers in order to identify or infer the emotional state, emotional tendencies and other emotional characteristics of test takers.
[0009] "Generative artificial intelligence models" refer to artificial intelligence models trained through machine learning or deep learning methods that are able to generate natural language responses, analysis results, or evaluation content based on input text, speech, or other information. These include, but are not limited to, large language models, multimodal generative models, or their variations.
[0010] "Pre-defined topics" refer to one or more topics, sets of questions, or scenario settings that are pre-configured by the system administrator or recruiter before the dialogue begins to guide candidates in communicating with the generative artificial intelligence model.
[0011] "Conversation content" refers to all interactive information generated during the interaction between the examinee and the generative artificial intelligence model around a pre-set topic, including voice, text, and facial expressions input or emitted by the examinee, as well as response text or voice output by the generative artificial intelligence model.
[0012] "Speaking" refers to the content expressed by the candidate in the form of voice or text during the dialogue, including answering questions, stating opinions, describing experiences, etc.
[0013] "Facial expressions" refer to the visual characteristics presented by the changes in facial muscles during a conversation, which are non-verbal information used to reflect the candidate's emotions or attitudes, including but not limited to smiling, frowning, and changes in eye contact.
[0014] "Pitch" refers to the acoustic characteristics of a candidate's speech, such as pitch, volume, speed, and intonation, which are used to reflect their emotional state, level of tension, or attitude.
[0015] "Emotion" refers to the subjective psychological or emotional state inferred by analyzing information such as the candidate's speech, facial expressions, and tone of voice, including but not limited to happiness, tension, anxiety, confidence, and neutrality.
[0016] "Emotion recognition results" refer to the structured or semi-structured output data about the candidate's emotional state or emotional changes generated by the emotion recognition engine based on the analysis of the candidate's speech, facial expressions and tone of voice.
[0017] "Analysis results" refers to the analytical data obtained by the processor after performing semantic understanding, structuring processing, and feature extraction on the dialogue content between the test taker and the generative artificial intelligence model. This includes the extraction results of information such as the theme, attitude, and behavioral patterns in the test taker's answers.
[0018] "Characteristics" refers to the analysis and description of relatively stable individual attributes such as personality traits, behavioral tendencies, communication styles, and value orientations of test takers, based on the results of dialogue analysis and emotion recognition.
[0019] "Emotional aspect" refers to the dimensions or aspects related to emotional state, emotional stability, emotional expression, and emotional regulation ability in the evaluation of candidates.
[0020] A “report” is a document or data structure generated by a processor based on the parsing results and sentiment recognition results, used to evaluate the characteristics and emotional aspects of a candidate. It can be presented in the form of an electronic document, a page, or structured data.
[0021] "Written screening" refers to the process of preliminary selection and judgment of candidates based on their resumes, CVs, application forms, papers, or other written materials submitted before an interview or conversation.
[0022] "Additional analysis" refers to the process of further evaluating existing dialogue content and related information based on new analytical objectives or focuses after the initial dialogue analysis and report generation.
[0023] "Prompt text" refers to the instructional text content that is constructed by the processor and sent to the generative artificial intelligence model to instruct the model to perform a specific analysis task or generate a specific type of output. It includes prompts, question lists, analysis requirements, etc.
[0024] "Response" refers to the answer or reaction of the candidate to a question posed by a pre-set topic or generative artificial intelligence model, including text response, voice response and its corresponding non-verbal behavior. Attached Figure Description
[0025] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0026] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0027] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0028] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0029] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0030] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0031] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0032] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0033] Figure 9 This represents an emotion map that maps multiple emotions.
[0034] Figure 10 This represents an emotion map that maps multiple emotions.
[0035] Figure 11This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0036] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0037] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0038] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0039] Hereinafter, an example of an implementation of the system to which the technology of this disclosure relates will be described with reference to the accompanying drawings.
[0040] First, let me explain the terminology used in the following instructions.
[0041] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0042] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0043] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0044] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0045] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0046] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0047] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0048] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0049] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0050] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0051] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0052] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0053] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0054] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0055] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0056] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0057] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0058] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0059] In existing computing-based job evaluation technologies, most systems merely use generative AI models as simple question-and-answer generation tools. The server side typically only acts as a relay for dialogue data, lacking fine-grained control over the conversation flow and instruction-level control over the generative AI model, leading to the following technical problems: (1) The server cannot perform structured management of the context of multi-turn dialogues. The generative artificial intelligence model relies on rough historical text splicing when generating questions each time, resulting in context loss, question redundancy and unstable dialogue path, thereby reducing the accuracy of capturing the features of the evaluated object.
[0060] (2) The server lacks a prompt generation mechanism for analysis tasks. The call to the generative artificial intelligence model usually only stays at the level of "generating the next reply". It cannot dynamically construct targeted prompts based on predefined evaluation perspectives or uncovered evaluation points and drive the model to generate high-value analysis text, resulting in the unutilization of the computational power of the generative artificial intelligence model.
[0061] (3) The text analysis on the server side is mostly single-dimensional (e.g., only sentiment analysis or keyword statistics are performed), and the structured parsing results output by the language processing program are not combined with the analysis text output by the generative artificial intelligence model. There is a lack of unified data structure and processing flow to support the generation of traceable and quantifiable character evaluation documents.
[0062] (4) In actual recruitment scenarios, traditional systems have difficulty automatically identifying “unevaluated items” or “insufficient information points” based on the information gaps exposed during the document review stage, and further supplementing the analysis content by driving the generative artificial intelligence model through additional prompts. As a result, the evaluation report still has structural gaps, requiring a lot of manual writing and correction. The system’s technical effect in reducing manpower burden and improving objectivity is limited.
[0063] (5) Existing multi-turn dialogue systems generally do not perform fine-grained real-time evaluation of “each response text”. The server cannot dynamically generate prompts corresponding to the predetermined evaluation indicators for each round of response during the dialogue process and call generative artificial intelligence models. Therefore, it is impossible to form a correspondence between the “response level” evaluation results and the dialogue session information, making it difficult to achieve subsequent reverse tracking and local re-analysis of specific answers.
[0064] Therefore, it is necessary to provide a new system and its implementation that enables the server to: (i) Structured management of multi-turn dialogues at the conversation level; (ii) At the analytical level, combine local language processing procedures with generative artificial intelligence models in synergy; (iii) Automatically generate targeted prompts that are closely related to the evaluation perspective at the prompt statement level; (iv) Automatically complete unevaluated items and generate structured evaluation documents at the report level; This will improve the process of collecting, processing, and evaluating natural language dialogue data at the computer technology level, and enhance the automation, consistency, and interpretability of character evaluations.
[0065] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0066] In this invention, the server includes: a data storage unit for pre-storing multiple topic information and associating user authentication information, conversation history information, and evaluation result information; an authentication control unit for receiving authentication information sent by the user through a terminal device and performing authentication processing, and prompting the user with a list of topic information upon successful authentication; a conversation generation unit for receiving identification information of selected topic information, generating conversation information corresponding to the user based on the identification information, and generating an initial prompt statement for input to a generative artificial intelligence model; a dialogue control unit for sending query data to an external generative artificial intelligence model based on the initial prompt statement and conversation history, obtaining question text for conversation from the generative artificial intelligence model, and sending it to the terminal device to advance multi-turn conversations with the user; and a dialogue control unit for sequentially obtaining user response text, accumulating the obtained response text as conversation information, and storing the conversation information... At least a portion is resent as contextual information to the generative AI model, thereby enabling the generative AI model to generate response management unit for the next round of conversational question text; a text parsing unit is used to perform text parsing processing, including word extraction, sentiment inference, and syntactic parsing, on the overall conversational information using a language processing program, and generate structured data from the parsing results; an analysis text acquisition unit is used to generate prompt statements based on the structured data and analysis data containing conversational information, send the analysis data with the prompt statements attached to it to the generative AI model, and obtain analysis text containing character evaluation information from the generative AI model; and a document generation unit is used to generate evaluation document data, including at least a portion of user personality traits, thinking patterns, communication styles, and role suitability information, based on the analysis text and the structured data, and convert the evaluation document data into an output format for storage or distribution. This allows for the formation of a complete technical chain on the server side, encompassing session management, prompt generation, text parsing, and evaluation document output. It not only fully leverages the reasoning and generation capabilities of generative AI models through automatic and iterative prompt generation, but also unifies the management of local structured parsing results with the model's output analysis text. This enables efficient computational processing of natural language dialogue data and high-precision character evaluation, significantly improving the functionality and performance of job applicant evaluation systems based on generative AI models at the computer technology level.
[0067] The “data retention unit” refers to a data processing module used to store information in a searchable manner in a storage device. This module records, updates, and manages the association of topic information, authentication information, session history information, evaluation result information, etc., enabling the server to access and manipulate the information based on a unified data structure.
[0068] "Authentication information" refers to identifying data used to identify and verify the identity of a user in a computer system, including but not limited to user IDs, passwords, tokens, certificates, etc., which are used by the server to perform login verification and access control processing.
[0069] "Conversation history information" refers to the accumulated record of conversation-related data exchanged between the terminal device and the server during multiple rounds of dialogue, including the question text proposed by the generative artificial intelligence model, the user's response text, and the time and sequence information of these texts.
[0070] "Evaluation result information" refers to summary or quantitative data generated based on user characteristics, behavior, or suitability, including rating results, grade results, and tag information, which are used to generate evaluation documents or provide reference to other business modules.
[0071] The "Authentication Control Unit" refers to the functional module in the server that performs processes such as receiving and verifying authentication information and generating session tokens. This module controls whether users can access the topic information overview and subsequent session functions by verifying the authentication information.
[0072] "Topic information" refers to preset topic data used to limit the scope of dialogue between generative artificial intelligence models and users, including topic-related names, descriptions, identifiers, and optional initial guidance content.
[0073] "Session information" refers to a session-level data structure used to identify and manage a complete multi-turn dialogue process, including session identifier, associated user identifier, topic identifier, dialogue status, start time, end time, and message record references associated with the session.
[0074] The “Session Generation Unit” refers to a functional module in the server used to generate new session information based on selected topic information and to construct initial prompt statements for input into the generative artificial intelligence model.
[0075] "Generative AI models" refer to AI computing models that are based on machine learning algorithms, especially deep learning and natural language processing technologies, and can automatically generate natural language text based on input text and prompts. These models are used to generate question text, analysis text, or other language output.
[0076] "Initial prompts" refer to the first set of instructional texts constructed by the server and sent to the generative artificial intelligence model at the beginning of a multi-turn dialogue. These texts are used to set the model's role, dialogue objectives, topic scope, and output style, serving as the starting point for subsequent dialogue generation.
[0077] "Query data" refers to the set of input data that a server constructs and sends in order to request the output of a generative artificial intelligence model. This data includes prompts, contextual information, and control parameters, and is used to drive the model to perform text generation or analysis processing.
[0078] "Question text" refers to user-oriented questioning natural language text generated by generative artificial intelligence models based on query data, used to guide users to respond in multi-turn dialogues.
[0079] The “dialogue control unit” refers to the functional module in the server used to manage multi-turn dialogue processes. This module is responsible for sending question texts to terminal devices, receiving response texts, maintaining the session state, and coordinating the interaction with generative artificial intelligence models.
[0080] "Response text" refers to the natural language text content that users input into the terminal device and send to the server in response to questions posed by generative artificial intelligence models, used to express the user's views, experiences, or attitudes.
[0081] The “Response Management Unit” refers to a functional module in the server used to receive and accumulate response text, and to resend at least a portion of the session information as context information to the generative artificial intelligence model to trigger the generation of the next round of question text.
[0082] "Contextual information" refers to the set of dialogue content used to provide historical environment and context for generative artificial intelligence models in multi-turn dialogues. It includes the question text and response text of the previous one or more rounds, and may also include relevant metadata when necessary.
[0083] A "language processing program" refers to a computer program that runs on a server and is used to automatically analyze text data. It can perform natural language processing operations such as word segmentation, part-of-speech tagging, sentiment analysis, and syntactic parsing.
[0084] “Text parsing processing” refers to the comprehensive analysis of conversational information using language processing programs. This process includes at least word extraction, sentiment inference, and syntactic parsing, and is used to extract structured information from natural language text.
[0085] "Word extraction processing" refers to the process of identifying and extracting keywords, key phrases, or words with specific semantic categories from the original text for subsequent statistical and feature analysis.
[0086] "Emotional presumption processing" refers to the analysis process of identifying the emotional or attitudinal tendencies (such as positive, neutral, negative, etc.) of text content in order to infer the user's emotional state in a conversation.
[0087] "Syntactic parsing" refers to the process of analyzing the sentence structure of a text, and obtaining higher-level semantic structural information by identifying syntactic components such as subject, predicate, and object and their relationships.
[0088] "Structured data" refers to data forms that are transformed from the results of text parsing and processing, and have clear fields and hierarchical relationships. Examples include information such as vocabulary statistics, sentiment tags, and syntactic relationships represented by key-value pairs, tables, or tree structures.
[0089] "Analysis data" refers to the set of data prepared for generative artificial intelligence models to perform character analysis, including structured data, conversational information, and auxiliary information related to evaluation, which are used as input for the model to generate analytical text.
[0090] "Prompt statements" refer to natural language instruction texts automatically generated by the server to instruct generative artificial intelligence models to perform specific tasks or follow specific output formats, including but not limited to question generation instructions, analysis instructions, and report generation instructions.
[0091] "Analysis text" refers to natural language text output by generative artificial intelligence models based on analysis data and prompts, containing information on character evaluations or related analytical conclusions, which is used to generate evaluation document data in the future.
[0092] "Character evaluation information" refers to the analysis results data representing users' personality traits, thinking style, communication style, role suitability, etc., which can be in the form of descriptive text, tags, or ratings.
[0093] "Evaluation document data" refers to structured text data organized in document form to display information on the evaluation of a person. It includes at least a portion of information on personality traits, thinking style, communication style, and role suitability, and can be further converted into an output file format.
[0094] The “document generation unit” refers to a functional module in the server used to generate evaluation document data based on the text and structured data used for analysis. This module is also responsible for converting the evaluation document data into an output format and performing storage or distribution processing.
[0095] "Document review information" refers to the judgment results data on the completeness and coverage of information obtained when reviewing existing materials or temporarily generated evaluation documents, which is used to identify unevaluated items or insufficient information.
[0096] "Unevaluated item information" refers to a set of items that have not yet been fully evaluated or for which corresponding analysis results have not yet been generated in a pre-defined evaluation indicator system, such as a personality dimension or ability dimension.
[0097] "Insufficient information" refers to a situation where the amount of information provided for a certain evaluation item or dimension in the existing dialogue data and evaluation documents is insufficient, lacks detail, or lacks supporting evidence.
[0098] "Additional condition information" refers to the constraints or key information determined based on document review information, used to indicate the content that needs to be supplemented for analysis, including the evaluation dimensions, problem angles, or information types that need to be supplemented.
[0099] "Additional prompts" refer to prompts automatically generated by the server to supplement unevaluated or insufficient information, containing additional conditional information, and are used to drive generative artificial intelligence models to output additional analysis text.
[0100] "Additional analytical text" refers to analytical natural language text generated by generative artificial intelligence models based on additional prompts, used to supplement unevaluated or insufficient information.
[0101] "Evaluation perspective information" refers to a predefined set of indicators used to evaluate user response texts from different dimensions or angles, such as the definitions and focus points of various evaluation dimensions like leadership, collaboration, and problem-solving ability.
[0102] "Evaluation result information" refers to the evaluation conclusions given by generative artificial intelligence models or other analysis programs based on evaluation perspective information for each response text or the entire conversation, including tags, ratings, or descriptive text.
[0103] This invention will focus on servers, terminals, and users, providing a detailed description of the system architecture, data structure, algorithm flow, generative artificial intelligence model, and its technical effects. The following embodiments are merely examples and do not limit the technical scope of this invention.
[0104] I. Overall System Composition In this invention, the server operates as the central processing node. The server includes: at least one multi-core central processing unit (CPU), an optional graphics processing unit (GPU), main memory (RAM), persistent storage (such as an SSD), a network interface, and an operating system. In one embodiment, the server uses a Linux-based server operating system (e.g., Ubuntu Server) and runs web server software (e.g., Nginx or Apache HTTP Server) and application server software (e.g., the Java-based Spring Boot framework or the Python-based Django framework).
[0105] The server installs a relational database management system (such as MySQL or PostgreSQL) on persistent storage, and optionally also installs a key-value data storage system (such as Redis) for session caching. The server communicates bidirectionally with the terminal via a network interface, using the HTTPS protocol to ensure the encryption and integrity of data transmission.
[0106] In this invention, the terminal can be a smartphone, tablet device, or personal computer. The terminal runs an operating system (e.g., Android, iOS, Windows, macOS) and a web browser (e.g., Chrome, Safari, Edge) or native applications. The terminal provides a user interface for displaying conversation content and report information sent by the server and for collecting text information input by the user.
[0107] Users access the application interface provided by the server using a terminal. In the interface, users enter authentication information, select a dialogue topic, and answer questions posed by the generative artificial intelligence model. Users can complete the evaluation process without needing to understand the server's internal processing.
[0108] II. Data Structure and Storage Method The server defines various data tables or equivalent data structures in the database, including but not limited to: 1. Subject Information Form The server assigns a unique topic identifier to each topic in the topic information table. This table contains: a topic identifier field, a topic name field, a topic description field, and an optional topic category field. The server uses this table to provide a list of topics to the terminal.
[0109] 2. User Information Form The server stores user identifiers, encrypted authentication information, and basic user-related data in the user information table. The server uses password hashing algorithms (such as bcrypt, PBKDF2, or Argon2) to hash user passwords and stores only the hash value, thereby reducing the risk of authentication information leakage.
[0110] 3. Session Information Table The server generates a session identifier for each multi-turn dialogue in the session information table and records information such as user identifier, topic identifier, session status (in progress, ended), start time, and end time for structured session management.
[0111] 4. Message Information Table The server stores each dialogue message in a message information table, including: message identifier, session identifier, sender type (generative AI model or user), message text, timestamp, and optional evaluation result reference. Using this table, the server can replay the complete dialogue in chronological order and attach analysis results to each message.
[0112] 5. Structured parsing results table The server stores structured data output by the language processing program in the structured parsing results table, including keyword lists, word frequency statistics, sentiment scores, syntactic relationships, and sentence category labels. The server saves this data as JSON fields or through multi-table joins for subsequent joint processing with the analysis output of generative artificial intelligence models.
[0113] 6. Evaluation Document Form For each generated evaluation report, the server records the document identifier, corresponding session identifier, document storage path, summary information, and generation time in the evaluation document table. The server can store the report content as a PDF file in the file system (such as a local disk or object storage service) and record the path in this table.
[0114] The above data structure design enables the server to access dialogue content and analysis results in an efficient and indexable manner internally, thereby reducing unnecessary data scanning and redundant calculations during multi-turn dialogues, analysis, report generation, and other processing, and improving data management efficiency.
[0115] III. Generative Artificial Intelligence Models and Their Structure In one embodiment of the present invention, the server calls an externally provided generative artificial intelligence model interface. This generative artificial intelligence model, in a specific form, is a multi-layer neural network based on the Transformer architecture. The model includes the following structural elements: 1. Word embedding layer When preparing input, the server segments the text into sub-word units or lexical tokens and assigns an integer ID to each lexical token. The generative AI model uses a trained embedding matrix to map the lexical token IDs to a high-dimensional vector space (e.g., 768 or higher), thus forming a continuous real-number representation.
[0116] 2. Position encoding or relative position representation Generative AI models incorporate positional encodings into the input sequence to indicate the position of each word within the sequence. These positional encodings can be fixed sine or cosine codes, or learnable parameter vectors, allowing the network to consider sequential information during attention computation.
[0117] 3. Multi-head self-attention layer Generative AI models utilize a multi-head self-attention mechanism to calculate weighted relevance between input vectors at each layer. Each attention head learns relevance weights in different subspaces, thereby extracting contextual dependencies. The server provides complete or truncated historical dialogue context, enabling the model to capture the semantic association between the user's preceding text and the current input under the attention mechanism.
[0118] 4. Feedforward network layer Generative AI models include a feedforward fully connected network in each Transformer layer to perform non-linear transformations on the representation at each position. This structure helps the model perform complex feature transformations and abstractions at the lexical level.
[0119] 5. Output layer and probability distribution Generative AI models map the hidden states of the last Transformer layer back to a vector the size of the vocabulary at the output, and calculate the probability of each word being the next output using a softmax function. The model generates the output sequence word by word in an autoregressive manner.
[0120] When invoking a model, the server passes in an input sequence containing prompts and context, and sets sampling parameters such as temperature, top-k, or top-p to control output diversity and determinism. By properly setting these parameters, the server can achieve a balance between stability and personalization.
[0121] IV. Language Processing Programs and Algorithms In this invention, in addition to calling the generative artificial intelligence model, the server also runs a language processing program locally. In one implementation, the server uses a natural language processing library (such as spaCy, NLTK, or a Transformers-based library) to achieve the following functions: 1. Word segmentation and part-of-speech tagging The server performs word segmentation on each user response text and labels each word with its part-of-speech tag. Based on the part-of-speech information, the server filters functional words and retains content words such as nouns, verbs, and adjectives for subsequent keyword extraction and statistics.
[0122] 2. Keyword extraction and statistics The server uses the TF-IDF algorithm, mutual information, or other statistical methods to extract high-importance keywords from the overall session information. For example, the server can calculate a keyword weight vector for each session and store the top few keywords in a structured parsing result table.
[0123] 3. Sentiment Analysis The server uses a sentiment analysis model (which can be a classifier based on traditional machine learning or a fine-tuned classifier head from a pre-trained language model) to calculate a sentiment score for each response text and label it as positive, neutral, or negative. The server records these sentiment scores in structured data for subsequent visualization of the evaluated documents.
[0124] 4. Syntactic Analysis The server uses a dependency parser to decompose sentences into a tree structure, identifying subject-verb-object relationships, modification relationships, and so on. Based on this, the server can determine whether the user's answer contains clear causal structure, goal description, specific behavioral description, and other features.
[0125] The server performs multi-dimensional analysis of the text using the aforementioned algorithm, and then combines the analysis results as structured features with the text used for analysis by the generative artificial intelligence model. Because the server supplements the model's output with structured features, the system can provide more interpretable metrics and charts in the evaluation report generation, thereby improving overall accuracy and credibility.
[0126] V. Examples of the Construction and Use of Prompt Statements In this invention, the server exerts fine-grained control over the generative artificial intelligence model by automatically constructing prompt statements. The server generates different types of prompt statements based on different stages.
[0127] 1. Example of prompts during the session generation phase At the start of the session, the server generates the following prompt statement to set the model role and question style: "You are a virtual interviewer responsible for recruitment assessment. The theme for this interview is 'Leadership.' Please use multiple rounds of questions to understand the candidate's actual performance in leadership behavior, decision-making style, and team communication. Ask only one specific question at a time and encourage the candidate to provide examples." The server passes the prompt along with the topic information to the generative AI model so that the model can consistently generate content around that topic in subsequent conversations.
[0128] When generating the first round of issues, the server can generate the following prompt: "Please talk about the most memorable experience you had leading a team to complete a task. Please explain the background, your role, the actions you took, and the result." 2. Examples of prompting statements during the follow-up questioning phase After receiving the user's answer, the server generates a prompt for follow-up questions, for example: "Based on the applicant's answers above, please ask a further question to focus on his leadership behavior in team conflict management or stressful situations. Please ask a specific question." Based on this, the generative AI model analyzes the previous answers using a self-attention mechanism and outputs targeted questions for the next round.
[0129] 3. Example of prompts during the report generation phase After the dialogue ends, the server generates the following prompt statement for the report: Below is a complete transcript of a conversation between a job applicant and a virtual interviewer on the topic of 'leadership'. Based on this conversation, please analyze and summarize the applicant from the following perspectives: 1. Personality traits (such as level of responsibility, initiative, stability, etc.); 2. Mindset (e.g., whether it is systematic, whether one is good at reflection, and whether one values data and facts); 3. Communication and collaboration style (e.g., listening, persuasion, and coordination skills); 4. Leadership characteristics and potential (such as goal setting, task allocation, motivation methods, and conflict resolution methods).
[0130] Please produce a report using structured headings and keeping the language objective and neutral. You may provide appropriate illustrative references, but do not disclose sensitive personal information. The server combines these prompts with session information and structured parsing results before sending them to the generative AI model. Because the prompts explicitly define the analysis dimensions and output structure, the model's output text is more stable in format and content distribution, which is beneficial for the server's subsequent automatic parsing and formatting.
[0131] VI. Evaluation of Document Generation and Technical Effectiveness After obtaining the text and structured data for analysis, the server generates evaluation document data. In one implementation, the server improves processing efficiency and accuracy by: 1. Templated report generation The server maintains a set of report templates, each defining the chapter structure, heading order, and insertion position. Based on the dimensions mentioned in the analysis text, the server populates the corresponding paragraphs into the appropriate chapters. The server embeds the ratings or distribution results from the structured data into the report in chart form. For example, the server can create radar charts or bar charts based on scores for personality dimensions.
[0132] 2. Local Updates and Incremental Generation After analyzing the document review information, the server generates additional prompts only for dimensions that were not evaluated or had insufficient information, retrieves the additional analysis text, and updates the corresponding sections, without regenerating the entire report. This incremental update mechanism reduces the number of calls to the generative AI model, thereby reducing communication load and computation time.
[0133] 3. The effects of data-driven technologies By storing dialogue data, parsing results, and analytical text in a unified structured data structure, the server allows subsequent retrieval and comparison of data from different sessions or candidates to focus solely on indexing, eliminating the need for extensive text parsing and model inference. This data management approach significantly reduces redundant computation and improves overall system throughput in large-scale deployments involving multiple users and sessions.
[0134] VII. Explanation of Model Training and Parameter Update In one extended implementation, the server can deploy its own generative artificial intelligence model, rather than simply calling external services. This model can be trained using the following methods: 1. Pre-training phase The model is pre-trained on a large-scale unlabeled corpus with language modeling as the objective. A server or external training environment uses the cross-entropy loss function to calculate the difference between the model's output sequence and the target sequence, and updates the network weights using backpropagation. The optimization algorithm can be Adam or a variant thereof.
[0135] 2. Fine-tuning stage The model is fine-tuned on dialogue corpora relevant to the evaluation scenario. At this stage, the server adds specific task labels, such as "good / poor answer quality" and "whether behavioral details are included." The model updates its parameters under supervised feedback to better adapt to interview-style dialogues and character analysis tasks.
[0136] 3. Data Augmentation When building the training set, the server can perform data augmentation through methods such as synonym replacement, sentence rearrangement, and noise injection, thereby improving the model's robustness to diverse expressions. Through these training and augmentation steps, the model's ability to generalize to user input during the inference phase is improved, resulting in higher quality generated questions and reports.
[0137] VIII. Explanation of Technological Improvements and Causal Relationships The server in this invention achieves substantial improvements in computer technology through the following key technical points: 1. Context-structured management The server manages each message and its relationships through session information tables and message information tables, avoiding the simple concatenation of all historical text each time the generative AI model is invoked. The server can selectively truncate historical content based on length limits and importance, retaining only key rounds and summaries, thereby reducing input length without significant semantic loss. This reduces model inference time and network traffic, improving system response speed.
[0138] 2. Local parsing and model collaboration The server first performs structured parsing locally using a language processing program, then inputs the parsed results along with the original text into a generative artificial intelligence model. Because the model can directly utilize the key information extracted from the parsing results, it becomes more focused when generating questions and analyzing text, reducing irrelevant output and thus improving analytical accuracy and effective information density. This combined processing approach differs from the traditional method of relying solely on the model's raw output, substantially improving the computational process.
[0139] 3. Automatic generation and dynamic control of prompt statements The server dynamically generates prompts corresponding to the evaluation perspective and automatically adjusts the angle and depth of the questions based on the user's actual responses during the dialogue. Because the prompts contain clear task descriptions and evaluation dimensions, the output of the generative AI model is constrained within a narrower semantic space, improving output stability and predictability. This mechanism, through a combination of server-side rules and model generation, achieves a non-traditional control flow, offering greater adaptability and programmability compared to purely human operation or fixed scripts.
[0140] 4. Hierarchical evaluation and traceability The server generates prompts corresponding to the evaluation perspective for each response text and obtains the evaluation results, then stores these results in a one-to-one correspondence with specific messages. This fine-grained data structure allows the system to trace which response a particular evaluation conclusion originated from during later analysis, facilitating correction and interpretation. Compared to the traditional black-box output method at the entire report level, the design of this invention improves system interpretability and debugging efficiency.
[0141] 5. Improve incremental reporting The server utilizes document review information and additional prompts to perform supplementary analysis and update reports only for unevaluated items or areas with insufficient information, avoiding repeated full generation of all content. This incremental strategy reduces model calls and network round trips, thereby significantly reducing the communication burden between the server and the generative AI model service in high-concurrency scenarios, and improving the overall system's scalability and performance.
[0142] Through the above-described embodiments, the server, terminal, and user together constitute a technical system that can be directly applied in the real world. The server not only executes simple automated business logic, but also improves the internal data processing flow and resource utilization methods of the computer through structured data management, model collaborative control, prompt statement design, and incremental update mechanisms. This results in verifiable technical effects in terms of processing speed, analysis accuracy, data management efficiency, and communication load.
[0143] use Figure 11The processing flow is explained.
[0144] Step 1: The user opens the system interface on the terminal and enters login information. The input consists of the user ID and password string entered by the user on the terminal interface. The terminal packages the user ID and password into a request message using the HTTPS protocol and sends it to the server. Specific actions performed by the terminal in this step include: displaying the login page, receiving keyboard or touch input, performing basic format validation on the input, and calling the operating system's network stack to send the HTTP request to the login interface specified by the server. The output is an encrypted network request containing the user ID and password.
[0145] Step 2: The server receives login requests from terminals and performs authentication. The input is an HTTP request message containing the user identifier and password. The server uses a web framework to parse the request body, extracting the user identifier and password strings. It then accesses the database to query the corresponding user record and reads the pre-stored password hash value. The server performs the same hash operation on the input password using a password hash library and compares the resulting hash with the hash value in the database. Based on the comparison result, the server generates a boolean value indicating whether the login was successful. If successful, it generates a session token and records the mapping between the user ID and the session token in session storage. The output is an HTTP response message containing the login result and the session token (if successful).
[0146] Step 3: The terminal receives the login result returned by the server and updates the local session state. The input is the server's HTTP response, which includes a login success flag and a session token or error message. The terminal parses the response content. When successful login is detected, the terminal writes the session token to the browser cookie or local storage and updates the interface state to redirect to the theme selection page. When login failure is detected, an error message is displayed on the interface, and the user's entered identifier is retained. The specific data processing performed by the terminal in this step includes: parsing the JSON response, handling conditional statements, and locally storing the session token. The output is the updated interface state and the locally saved session token.
[0147] Step 4: The user selects a conversation topic in the terminal's topic selection interface. The input consists of a list of topics retrieved and displayed by the terminal from the server, and the user's clicks on the interface. The terminal obtains the identifier of the selected topic based on the clicked interface element and encapsulates this topic identifier along with the session token into request data. The terminal performs only a simple index lookup on the topic list data to obtain the topic identifier, without modifying the content. The output is an HTTP request containing the session token and topic identifier, ready to be sent to the server's session initiation interface.
[0148] Step 5: The server receives a topic selection request and creates a session record and an initial prompt statement. The input is a request message containing the user's session token and the selected topic identifier. The server first retrieves and verifies the user's identity from the session storage based on the session token; then, it inserts a new session record into the database, writing the user identifier, topic identifier, session status, and start time, and generating a unique session ID. The server then reads the topic name and preset description from the topic information table, populates them into a pre-designed prompt statement template, and generates the initial prompt statement through string replacement and concatenation operations. The output is the newly created session ID and the generated initial prompt statement; this data will be used for subsequent calls to the generative artificial intelligence model.
[0149] Step 6: The server constructs input for the generative AI model based on the initial prompt statement and sends a query. The input includes the initial prompt statement, the topic name, and user and system role settings. The server organizes this text into the dialogue format required by the model, labels the content for different roles as system messages or user messages, concatenates the text into an ordered sequence, and then uses an HTTP client library to assemble it into an API call request, setting parameters such as model name, temperature, and maximum generation length. The server sends the packaged request data via HTTPS to the server endpoint where the generative AI model resides. The output is a model call request containing structured prompt statements and dialogue context information.
[0150] Step 7: The server receives the initial question text returned by the generative AI model and stores it as a dialogue message. The input is the response content from the generative AI model, the main body of which is natural language question text generated based on the prompts. The server parses the response JSON, extracts the question string, and inserts it as a new message into the message information table, recording the session ID, sender type (model), message content, and timestamp. The server also encapsulates this question text into an HTTP response body and sends it back to the terminal. The data processing includes JSON parsing, database insertion, and response message construction. The output is an HTTP response containing the initial question text and a persistently stored message record.
[0151] Step 8: The terminal displays the initial questions posed by the generative AI model and receives user responses. The input is the question text from the server's response. The terminal renders this question text onto the chat interface and updates UI components to display the message from "System / Interviewer." Subsequently, the user enters their response in the text input box and clicks the send button. The terminal listens for this event and encapsulates the user's input response string along with the session ID and session token into a message sending request. The processing performed by the terminal in this step includes: UI rendering, event listening, string reading, and HTTP request construction. The output is the user's human-computer interface display status and a user response request ready to be sent to the server.
[0152] Step 9: The server receives the user's response text and updates the session history. The input is an HTTP request containing the session ID, the user's response text, and a session token. The server first verifies the request's validity using the session token, then inserts a new message record with the user as the sender into the message information table, writing the session ID, message text, and timestamp. Immediately after insertion, the server can invoke a local language processing program to perform preprocessing on this response text, including tokenization and basic cleaning (removing redundant whitespace, etc.). These preprocessing results can be temporarily stored in memory or directly written to a structured parsing result table. The output consists of the updated database record and the preprocessed text content, which will serve as the context for the next call to the generative AI model.
[0153] Step 10: The server constructs new prompts based on the session history and latest responses, and then invokes a generative AI model to generate the next round of questions. The input consists of the most recent message records corresponding to the session ID and the current user's response text. The server extracts a specified number of historical messages from the message information table, sorted by time, and combines the content into a dialogue context. Then, based on a pre-configured evaluation perspective (e.g., leadership, collaboration), it generates follow-up prompts, such as, "Based on the applicant's answers, here's a further question, focusing on their leadership behavior in team conflict management or stressful situations." The server constructs the follow-up prompts and the session context into a model input sequence and sends it again to the generative AI model via an HTTP request. The output is a new model query request, which drives the model to generate the next round of questions.
[0154] Step 11: The server receives the next round of questions generated by the generative AI model and loops through the dialogue process. The input is the new question text from the model's response. The server parses the response and writes the question text into the message information table as a new message, then returns the question to the terminal as an HTTP response. The terminal displays the new question, and the user answers again. Through this loop, the server continuously receives user responses, updates the session history, constructs prompts, and invokes the generative AI model, enabling the multi-turn dialogue to continue. The output of this step is the updated message record and the new question text returned to the terminal, forming a multi-turn question-and-answer sequence.
[0155] Step 12: After the conversation ends, the server collects all message text from the entire session and performs text parsing. The input is the session ID and all associated message records. The server reads all question and response text from the message information table for that session, concatenating the user's responses in chronological order to form the analysis object. The server inputs this text into a language processing program, which sequentially performs word segmentation, part-of-speech tagging, keyword extraction, sentiment analysis, and syntactic parsing. Each step maps the raw text to a higher-level feature space: for example, keyword extraction calculates word weights using TF-IDF, sentiment analysis outputs a sentiment score for each response using a classifier, and syntactic parsing outputs a dependency tree structure. The server converts these results into structured data and records them in a structured parsing result table. The output is a set of structured feature data corresponding to the session, providing a foundation for subsequent user evaluation.
[0156] Step 13: The server generates analysis prompts based on the structured parsing results and dialogue text, and requests the generative AI model to output the analysis text. The input consists of structured feature data of the conversation and the original dialogue text. The server constructs analysis prompts that detail the dimensions to be analyzed by the generative AI model (such as personality traits, thinking style, communication style, role suitability), and instructs the model to use a report format with chapters and subheadings. The server combines these prompts with the entire conversation content and summarized structured features (such as a list of high-frequency keywords and sentiment trends) to form an analysis request. The server sends the analysis request to the generative AI model server via an HTTP interface. The output is an analysis request that drives the model to generate structured analysis text.
[0157] Step 14: The server receives and post-processes the analysis text returned by the generative AI model. The input is the analysis text, which contains multiple paragraphs of character evaluations as required. The server performs format checks and basic cleaning on the text, such as removing redundant punctuation, standardizing title formatting, and segmenting it into multiple sections based on preset tags (e.g., "I," "II,"). The server writes the analysis text as a record to the database and associates it with the corresponding session ID. The server can also apply local rules or simple classifiers to the analysis text, highlighting or annotating certain keywords if necessary. The output consists of a well-structured analysis text record and a database entry associated with the session.
[0158] Step 15: The server detects unevaluated items or insufficient information based on document review information and generates supplementary prompts. The input consists of initially generated analysis text and a list of predefined evaluation metrics. The server checks whether each evaluation dimension is covered in the report through string search or structured parsing. If some dimensions are missing or overly brief, document review information is generated in memory to mark these gaps. The server then constructs supplementary prompts based on the review information, such as "Please supplement the analysis of the applicant's performance in stress tolerance and provide specific evidence." The server then sends the supplementary prompts and a summary of the original dialogue back to the generative AI model. The output consists of several supplementary prompts and corresponding model call requests used to generate supplementary analysis text.
[0159] Step 16: The server receives supplementary analysis text and integrates it into the evaluation document data. The input consists of supplementary analysis text returned by the generative AI model and existing analysis text. The server inserts or appends the supplementary analysis text to the corresponding sections of the original report, updating the overall structure of the analysis text. The server then generates an evaluation document data object based on the complete analysis text and structured feature data, containing descriptions of each evaluation dimension and optional quantitative scores. The server uses document templates to map this content to section titles, paragraphs, and chart areas, forming a renderable internal data structure. The output is a complete evaluation document data record.
[0160] Step 17: The server converts the evaluation document data into an output file and saves or distributes it to the terminal. The input consists of the evaluation document data object and possible output parameters (such as PDF or HTML format). The server calls a document generation library to format text paragraphs, insert charts, generate PDF or HTML files, and writes the generated files to the file system or object storage service. The server registers the file path and metadata in the evaluation document table. When a terminal sends a report viewing request, the server locates the corresponding file path based on the report identifier in the request and sends the file content to the terminal via an HTTP response. The output consists of a report file stored in the file system and a file data stream sent to the terminal, which the terminal then presents to the user or recruitment manager.
[0161] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0162] While generative AI models have been used for dialogue generation and text analysis in existing technologies, the following technical challenges remain in real-world human-computer interaction scenarios, especially those requiring comprehensive evaluation and decision support based on real-time dialogue regarding the subject's characteristics, emotional state, and individualized needs: First, existing systems mostly use generative artificial intelligence models as simple question-answering engines, lacking technical solutions for integrated collection, analysis, and structured representation of "conversation content + multimodal signals (voice, facial expressions)," making it difficult to perform high-precision comprehensive modeling of the emotional state and personality characteristics of the evaluated object.
[0163] Second, existing character evaluation systems generally use fixed questionnaires or preset rules for analysis, lacking a mechanism for dynamically generating prompts based on the conversation process and automatically driving generative artificial intelligence models for in-depth analysis. They cannot flexibly adjust the analysis perspective, evaluation indicators, and evaluation benchmarks according to the real-time dialogue content, thus easily leading to rigid evaluation dimensions and insufficient information mining.
[0164] Third, existing systems typically separate "offline assessment" from "online assistance," generating reports only after the session ends. They fail to utilize generative AI models to infer the needs and preferences of the evaluated individual in real time during the conversation and provide low-latency feedback to the front-end display device, thus hindering effective on-site response. This deficiency makes it difficult for computer systems to proactively assist in decision-making during real-time interactions.
[0165] Fourth, the lack of a unified computer implementation framework in existing technologies makes it impossible to form an organic, pipelined processing process for dialogue content parsing, emotion recognition, prompt generation, generative artificial intelligence model invocation, report generation, and auxiliary information output within the same information processing device. This results in a complex system architecture, high data coupling between modules, and low processing efficiency, making it difficult to meet the requirements of real-time performance, scalability, and maintainability in practical applications.
[0166] Therefore, it is necessary to propose a new system and its computer implementation scheme. By integrating functions such as dialogue parsing, prompt statement generation, multimodal emotion recognition, and collaborative processing of generative artificial intelligence models on the server side, a unified technical framework can be constructed that can infer the needs / preferences of the evaluated object in real time during the conversation and output auxiliary information, while generating a high-quality person evaluation report after the conversation. This will improve the processing power, response efficiency, and intelligence level of the human-computer interaction system from the perspective of computer technology.
[0167] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0168] In this invention, the server includes: a processing unit for processing audio, image, and other data collected by a front-end terminal around a pre-defined conversation topic, executing a dialogue between an information processing device and a generative artificial intelligence model through an evaluated object, and converting the dialogue content into time-series text information for parsing; a processing unit for automatically generating prompt statements as input to the generative artificial intelligence model based on the text information, and sending the prompt statements and the text information to the generative artificial intelligence model to obtain analysis results on the evaluated object's personality traits and preferences; and a processing unit for analyzing the spoken content contained in the text information, the acoustic features extracted from the speech signals corresponding to the spoken content, and the facial expression detection... The system comprises: a processing unit that performs emotion recognition calculations on image information acquired by the testing device to infer the emotional state of the evaluated object; a processing unit that generates report data containing evaluation information on the characteristics, hobbies, and emotions of the evaluated object based on the analysis results and the inferred emotional state; a processing unit that, during the dialogue, infers the needs and preferences of the evaluated object in real time based on the analysis results obtained from the generative artificial intelligence model, and generates auxiliary information for output on the display device to support the dialogue participants' responses; and a storage control unit that stores the report data and the auxiliary information in a recordable and reusable data structure. This allows for the formation of an integrated processing pipeline on the same server, encompassing multimodal conversation data acquisition and parsing, automatic generation of prompts, invocation of the generative artificial intelligence model, emotion recognition, real-time auxiliary information output, and post-meeting report generation. This enables high-precision modeling and real-time inference of the characteristics, emotions, and needs of the evaluated object, improving the processing efficiency and intelligent assistance capabilities of the computer system in complex interactive scenarios, thereby improving human-computer interaction and evaluation technology based on generative artificial intelligence models.
[0169] "System" refers to a collection of devices consisting of one or more information processing devices, storage devices, and input / output devices, used to perform the processing steps described in this invention.
[0170] "Information processing device" refers to an electronic computing device that includes a processor, memory, and communication interface, used to execute programs to process input data and control data flow.
[0171] "Generative artificial intelligence models" refer to artificial intelligence models trained based on machine learning and deep learning algorithms that can generate new text, perform semantic analysis, or make inferences based on input text or other data.
[0172] "Prompt statements" refer to control text information generated by an information processing device and input into a generative artificial intelligence model, used to instruct the model on its analysis objectives, output format, and evaluation criteria.
[0173] "Dialogue" refers to a multi-round language interaction process between the evaluated object and the information processing device (through a terminal) around a preset conversation topic.
[0174] "Conversation topic" refers to a thematic concept or topic that is set before the conversation begins to limit the scope of the conversation and the direction of analysis.
[0175] "The evaluated subject" refers to the individual whose speech, behavior, emotional state, personality traits, and hobbies are analyzed and evaluated by this system during the dialogue process.
[0176] "Temporal text information" refers to the text of dialogue content obtained by speech recognition or other methods and its associated time information, arranged in chronological order of occurrence.
[0177] "Speech content" refers to the textual representation of the spoken content delivered by the person being evaluated during a dialogue.
[0178] "Speech signal" refers to analog or digital audio data acquired by a data acquisition device that reflects the voice produced by the object being evaluated.
[0179] "Acoustic feature information" refers to the set of parameters obtained from speech signals through signal processing and feature extraction algorithms, used to characterize attributes such as pitch, volume, speech rate, and timbre.
[0180] "Face detection device" refers to an image acquisition and processing device used to collect facial images of the evaluated object and detect and analyze its expressions.
[0181] "Image information" refers to image data acquired by an expression detection device that reflects the facial or body posture and other states of the person being evaluated.
[0182] "Emotion recognition processing" refers to the computational processing of speech signals, text information, and image information by performing pattern recognition and feature analysis to infer the emotional state of the evaluated object.
[0183] "Emotional state" refers to the emotional or psychological state of the evaluated object at a specific point in time, inferred through emotion recognition processing.
[0184] "Character traits" refers to the characteristics of the evaluated person in terms of personality tendencies, behavioral styles, and ways of thinking, inferred from the content of the dialogue and the results of the analysis.
[0185] "Preference" refers to the persistent preference and tendency to choose goods, services, topics, or behaviors that the evaluated object exhibits.
[0186] "Analysis results" refers to the structured or semi-structured information output by the generative artificial intelligence model after processing the characteristics, hobbies, and needs of the evaluated object based on prompts and input data.
[0187] "Report information" refers to document-based or data-based information generated based on analysis results and emotional state inferences, used to comprehensively evaluate the characteristics, hobbies, and emotions of the evaluated person.
[0188] "Assessment information" refers to evaluative data or descriptions included in the report information, used to quantitatively or qualitatively describe the characteristics, emotional state, and related indicators of the evaluated object.
[0189] "Demand" refers to the need for a certain product, service, information, or support expressed directly or indirectly by the person being evaluated during a dialogue.
[0190] "Preference" refers to the stable tendency or preferred direction of an evaluated object to choose certain options among multiple options.
[0191] "Supporting information" refers to output information generated based on analysis results and inferences, used to provide suggestions, prompts, or recommendations to participants during the dialogue.
[0192] "Dialogue participants" refers to individuals or system users who, other than the person being evaluated, participate in the dialogue and respond to it by referring to supplementary information.
[0193] "Display device" refers to an output device used to visually present auxiliary or report information to participants in a dialogue, including but not limited to head-mounted display devices, terminal displays, or other visual output devices.
[0194] The “storage control unit” refers to the functional module used to control the recording, reading and management of report information and auxiliary information in the storage medium.
[0195] "Reusable format" refers to a storage method that preserves information in a structured or semi-structured data format, enabling the stored data to be accessed and used again in subsequent processing, retrieval, or analysis.
[0196] In one embodiment of the present invention, a server, a terminal, and a user collaboratively constitute a system for performing multimodal dialogue analysis and real-time auxiliary output on the evaluated object. The server executes a program stored on a non-transitory computer-readable medium to acquire, process, compute, and store audio data, image data, and text data, and interacts with a generative artificial intelligence model to generate report information and auxiliary information.
[0197] In terms of hardware, servers can use general-purpose computer equipment, such as rack servers equipped with multi-core central processing units (CPUs) and graphics processing units (GPUs), and run general-purpose operating systems (such as Linux-based server operating systems). In terms of software, servers can run web application frameworks (such as service frameworks implemented in Python), speech recognition engines (such as speech recognition software based on deep neural networks), emotion recognition modules, and generative artificial intelligence model access interfaces (such as calling external large-scale language model services via HTTP / HTTPS protocols).
[0198] The terminal can be smart glasses, a mobile terminal, or a tablet terminal with audio acquisition and display capabilities. The terminal hardware includes a microphone, camera, display screen, wireless communication module, and local processor. The terminal application runs on the software to collect voice and image data from the user and the evaluated subject, transmit this data to a server via the network, and simultaneously present the auxiliary information returned by the server to the user on the display screen.
[0199] Users wear smart glasses or operate a mobile device to engage in natural language conversations with the person being evaluated. Users receive auxiliary information from the server via the device's display screen, allowing them to adjust their questioning style or response strategies during the conversation.
[0200] A server's program structure can include multiple functional modules, each interacting through well-defined data structures and interfaces. For example, a server might include a dialogue data management module, a speech recognition module, a multimodal emotion recognition module, a prompt generation module, a generative AI model interaction module, an analysis result integration module, a report generation module, an auxiliary information generation and output module, and a data storage control module. Information is exchanged between modules using structured data (e.g., key-value pairs, lists, nested objects).
[0201] In the dialogue data management module, the server performs session-level management of audio and image data received from the terminal. The server assigns a unique session identifier to each evaluated entity and maintains a dialogue history data structure for that session, including timestamps, speaking roles, and text content. This data structure may include fields such as: session ID, round number, role tag (user / evaluated entity), timestamp, original audio data index, original image data index, and transcribed text. Through this structured management, the server can quickly locate dialogues within a specific time period and their corresponding multimodal signals without increasing communication load, thus supporting subsequent fine-grained analysis.
[0202] In its speech recognition module, the server divides continuous audio data into segments of fixed duration or semantic boundaries and processes them using a deep neural network-based speech recognition engine. The server can use acoustic models based on convolutional neural networks and recurrent neural networks, or an end-to-end speech recognition model based on the Transformer architecture, to process audio frame sequences. The server extracts acoustic features such as Mel-frequency spectrum and Mel-frequency cepstral coefficients (MFCCs) from the audio signal, inputs these features into the neural network to obtain the probability distribution of characters or phonemes corresponding to each frame, and then generates a text sequence using decoding algorithms (such as CTC decoding and Beam Search decoding). Through this neural network-based feature extraction and decoding algorithm, the server can achieve high-precision text transcription in noisy environments, improving the overall accuracy of subsequent semantic analysis.
[0203] In the multimodal emotion recognition module, the server uses acoustic features extracted from speech signals (including pitch, energy, speech rate, formant distribution, etc.) and facial expression features extracted from image data (e.g., locating facial key points and calculating expression feature vectors using convolutional neural networks) to construct the input for the emotion recognition model. The server can employ a multimodal fusion neural network structure, for example, concatenating acoustic and visual feature vectors at the feature layer, and then inputting them into a multi-layer fully connected network with attention mechanisms or a Transformer encoder to perform weighted fusion of different modal features, outputting an emotion category probability distribution (e.g., happy, nervous, calm, confused, etc.) and an emotion intensity score. The server pre-trains this multimodal emotion recognition network by minimizing cross-entropy loss or mean squared error loss. In practical applications, the pre-trained model weights are used to perform forward inference on real-time input data, thereby inferring the emotional state of the evaluated object at various time segments.
[0204] In the prompt generation module, the server automatically constructs prompts adapted to the generative AI model based on the conversation topic and dialogue history. The server pre-stores multiple prompt templates, each corresponding to a different analysis task, such as "needs / preference analysis," "personal profile summary," and "evaluation metric setting." Based on the current conversation state, the server extracts key statements from the most recent rounds of the evaluated entity from the dialogue history and fills them into placeholders in the corresponding templates, forming the specific prompts.
[0205] For example, when the server is performing real-time demand / preference analysis, it can generate the following Chinese prompt: You are a generative AI model assistant for shopping mall guides.
[0206] Subject: Recommendations for healthy foods.
[0207] Objective: To analyze customers' current needs and preferences from their statements and output structured results.
[0208] Output format: 1) Current requirements (briefly stated in one sentence); 2) Preferences (taste, health concerns, price preferences, etc., in list format); 3) Recommendation direction (3 product categories, with a one-sentence explanation for each).
[0209] The following are the latest rounds of conversations between the customer and the sales associate: User: What kind of products are you looking for today? Customer: I've been interested in health foods lately, preferably not too sweet.
[0210] Please output according to the above format.
[0211] When the server generates a customer report after the meeting, it can generate the following prompt: You are a customer profile analysis assistant.
[0212] Please generate an analysis report for this customer based on the following complete dialogue, including: - Character profile (personality traits, areas of interest); - Long-term needs and interests; - Significant preferences (such as health, brand, price, taste, etc.); - Suggested customer service strategy (how to communicate with this customer and recommend products in the future).
[0213] The following is the dialogue: ...(Complete dialogue text)... In the generative AI model interaction module, the server calls externally or locally deployed generative AI model services via a network interface. The server can choose a large-scale language model based on the Transformer architecture. Such models typically include multi-layered self-attention encoders and decoders, employing multi-head attention mechanisms and feedforward networks to semantically encode the input text. The server encodes the input sequence, consisting of prompts and dialogue text, into a one-dimensional token sequence, performs word segmentation or sub-word splitting on this sequence, converts it into integer indices, and then forms a vector representation through embedding layers and positional encoding. The model performs self-attention calculations and non-linear transformations on the input sequence in multiple attention layers, and finally generates the next token with the highest probability through the output layer, thus progressively generating the analysis result text.
[0214] When training generative AI models, servers can employ supervised learning methods, using large-scale text corpora and dialogue data as training samples. They can use the cross-entropy loss function to calculate the difference between the model output and the target text, and update the model parameters using a gradient descent-based optimization algorithm (e.g., Adam). Servers can also expand the training sample space and improve the model's generalization ability to different expressions through data augmentation methods (e.g., synonym rewriting, sentence shuffling, mask prediction).
[0215] In the analysis results integration module, the server performs structured parsing of the text results output by the generative artificial intelligence model. The server can pre-define the output format in the prompt statements, requiring the model to use specific tags or paragraph order. After receiving the model output, the server extracts fields such as "current needs," "preferences," and "recommendation direction" through regular expression matching, keyword retrieval, or simple syntax parsing, constructing them into internal data objects. For example, the server can store the "needs" field as a short descriptive string, the "preferences" field as a list, and the "recommendation direction" field as a collection of categories and reasons. This structured result significantly reduces the parsing complexity of subsequent recommendation algorithm modules and improves overall computational efficiency.
[0216] In the report generation module, the server integrates the analysis results from the generative artificial intelligence model with the emotional state information output from the multimodal emotion recognition module. The server can use weighted rules to align emotional states at different time points with corresponding dialogue content and generate more detailed character analyses based on emotional patterns (e.g., "tension increases when discussing prices" or "emotions tend to be positive when discussing health topics"). The final report output by the server can include text descriptions, rating scales, and tag sets. By mapping this information to a unified report data structure, reports from different conversations are comparable and searchable.
[0217] In the auxiliary information generation and output module, the server retrieves matching items from a predefined knowledge base or product / service database based on real-time analysis results. The server can use keyword retrieval algorithms, vector space-based similarity search, or collaborative filtering-based recommendation methods to map preference tags such as "healthy food," "low sugar," and "ready-to-eat" to a set of entity entries and calculate a matching score. During calculation, the server can employ sparse matrix multiplication and an inverted index structure to reduce retrieval latency, thus achieving low-latency feedback during the dialogue. The server converts the top-scoring results into brief descriptive text and generates auxiliary information for terminal display, such as "Recommendation: Low-sugar, high-protein oat bars, suitable for customers concerned about weight management, currently with a buy-one-get-one-half-price promotion." After receiving auxiliary information from the server, the terminal presents it to the user on the display screen as overlaid text or cards. During presentation, the terminal can adjust the timing and position of the information display according to preset display priorities and time control strategies to avoid obstructing key areas of the user's view. The terminal can also send user feedback on the auxiliary information (such as whether to accept a recommendation) to the server. Upon receiving this feedback, the server can use it as training signals for subsequent recommendation algorithms, enabling online optimization.
[0218] Within the storage control unit, the server stores dialogue history data, analysis results, emotion state sequences, report information, and auxiliary information in a structured format within a database or other storage medium. The server can manage this data using a relational database management system or a document-oriented database, supporting fast queries and batch statistics through indexing and views. This unified data management strategy allows the server to improve data reading and writing efficiency without adding large amounts of redundant data, and provides a reliable data foundation for subsequent model retraining and system optimization.
[0219] Through the aforementioned structure and processing flow, the server, terminal, and user collaborate collaboratively within the system of this invention with a clear division of labor. The server, through multimodal feature extraction, deep neural network inference, structured result parsing, and efficient data management, achieves high-precision modeling and real-time inference of the evaluated individual's characteristics, emotional state, needs, and preferences. This processing is not simply an automation of manual analysis processes; rather, by introducing a learnable model structure, adjustable prompts, and a pipelined data processing architecture, it improves speech recognition accuracy, emotion recognition robustness, and computational efficiency of recommendation decisions from the perspective of the computer's internal processing mechanisms. This reduces the burden of manual rule maintenance and feature engineering, and by combining real-time auxiliary output with post-meeting report generation, it achieves a closed-loop data processing system, enhancing the overall technical performance of the computer system in complex interactive scenarios.
[0220] use Figure 12 The processing flow is explained.
[0221] Step 1: The user puts on the device and begins a conversation.
[0222] Users engage in natural language communication with the evaluated entity, such as asking about needs or introducing topics.
[0223] Input: None (Natural dialogue initiated by the user).
[0224] Output: Acoustic signal containing the user's speech and the speech of the evaluated object.
[0225] After the user speaks, the microphone in the terminal collects the analog acoustic signals from the environment into the audio input circuit, providing the original sound source for subsequent digitization and processing.
[0226] Step 2: The terminal collects audio data and digitizes it.
[0227] The terminal converts the analog acoustic signal obtained from the microphone into digital audio data, such as 16kHz sampling rate, 16-bit quantized PCM data, and slices the audio stream into fixed time windows.
[0228] Input: Continuous analog acoustic signal from the microphone.
[0229] Output: Segmented digital audio data blocks (audio frame sequence).
[0230] The terminal buffers these data blocks locally and adds a timestamp and speaking channel marker to each data block (e.g., distinguishing between channels closer to the user and those closer to the person being evaluated) so that the server can subsequently distinguish the speaker.
[0231] Step 3: The terminal transmits audio data to the server.
[0232] The terminal establishes a secure connection with the server through a wireless communication module (such as Wi-Fi or cellular network), and sends time-stamped audio data blocks to the server in sequence using a long connection method such as WebSocket or HTTPS.
[0233] Input: Locally buffered digital audio data blocks and their metadata (timestamps, channel tags).
[0234] Output: Audio data stream transmitted over the network to the server.
[0235] When sending audio data, the terminal can compress and encode it (such as Opus or AAC) to reduce bandwidth usage and embed session IDs and round markers in the data packets to ensure that the server can correctly reassemble the session.
[0236] Step 4: The server caches and reassembles the audio stream.
[0237] The server receives audio data packets sent by terminals over the network, stores them in a memory buffer according to session ID and timestamp, and reassembles the continuous audio stream in chronological order.
[0238] Input: Compressed or uncompressed audio data packets containing session ID, timestamp, and channel tag.
[0239] Output: A collection of audio clips organized by session and time sequence.
[0240] The server divides the continuous audio stream into speech recognition task units according to a preset segmentation strategy (such as a segment every 5 seconds or with silence detection as the boundary), providing input of a reasonable length for subsequent speech recognition modules.
[0241] Step 5: The server performs speech recognition and generates text.
[0242] The server calls the speech recognition engine, inputting each audio segment into a speech recognition model based on a deep neural network. The model first extracts acoustic features (such as Mel spectrum and MFCC) from the audio frames, then calculates the acoustic probability distribution corresponding to each time step through a multi-layer neural network (such as convolutional layers, recurrent layers, or Transformer encoders), and finally generates text through a decoding algorithm.
[0243] Input: A sequence of digital audio data divided into segments.
[0244] Output: A dialogue text fragment with time boundaries and its confidence score.
[0245] The server performs post-processing on the recognition results, including re-scoring the language model, removing noise words, and handling common colloquial abbreviations and typos, thereby improving the readability and semantic accuracy of the text.
[0246] Step 6: The server identifies the speaker and builds the dialogue history.
[0247] The server maps each transcribed text to a "user" or "evaluated subject" role based on channel markers, energy distribution, or an optional speaker recognition model. The server encapsulates text fragments with timestamps into dialogue units and appends them to the dialogue history list of the current session in chronological order.
[0248] Input: A transcribed text fragment with a timestamp and its channel information.
[0249] Output: Structured dialogue history data (a list containing fields such as role, time, and text).
[0250] While building the dialogue history, the server asynchronously writes the structured dialogue data to the database, providing persistent storage for subsequent report generation and model retraining.
[0251] Step 7: The server performs multimodal feature extraction and emotion recognition.
[0252] The server extracts acoustic features (such as fundamental frequency, volume, speech rate, and formant parameters) from audio data and facial feature points and expression vectors from image data corresponding to the same time window. The server concatenates or weights and fuses the acoustic and visual feature vectors at the feature layer and inputs them into a pre-trained multimodal emotion recognition neural network. This network uses, for example, multiple fully connected layers and self-attention mechanisms to weight the contributions of different modalities and outputs the probability of emotion category and emotion intensity score.
[0253] Input: Time-slice aligned audio signal segments and image frame sequences.
[0254] Output: Emotional labels (such as happy, tense, calm, etc.) and intensity values marked by time segments.
[0255] The server associates the emotional results of each time slice with the corresponding dialogue unit to form a joint representation of "text + emotion", providing richer input features for subsequent inference of character characteristics and needs preferences.
[0256] Step 8: The server generates prompts that are adapted to the generative artificial intelligence model.
[0257] Based on the current session's preset topic (e.g., "healthy food recommendations"), the latest rounds of dialogue history, and emotional state, the server selects the corresponding prompt template and dynamically fills the specific dialogue text into the placeholder positions in the template. Simultaneously, the server embeds the desired output format (e.g., paragraph structure or list format) and evaluation metrics (e.g., "needs," "preferences," "recommendation direction") into the prompt.
[0258] Input: Structured dialogue history data, conversation topic information, and corresponding emotion tags.
[0259] Output: A prompt statement in the form of a specific text document.
[0260] By using this templated prompt generation method, the server explicitly encodes complex analysis tasks into text, guiding the generative artificial intelligence model to output results according to a predetermined structure, thereby reducing the complexity of result parsing and improving overall processing efficiency.
[0261] Step 9: The server calls a generative artificial intelligence model to perform semantic analysis.
[0262] The server concatenates the prompt and the most recent dialogue text into a single text input sequence, which is then sent to the generative AI model via an API. The model, based on the Transformer architecture, performs word segmentation, embedding, and positional encoding on the input sequence. It then calculates the correlation between the tags in the sequence within a multi-layer self-attention network, iteratively updates the hidden states, and finally generates the analysis result text step-by-step according to the tags.
[0263] Input: A sequence of input text containing prompts and dialogue text.
[0264] Output: Text of the analysis results regarding the current needs, preferences, characteristics, or suggested directions of the evaluated individual.
[0265] After receiving the model output, the server parses the output text, extracts predefined fields such as "Current requirement: ...", "Preference: ...", "Recommendation direction: ...", and encapsulates them into a structured data object.
[0266] Step 10: The server integrates and analyzes the results with sentiment information.
[0267] The server aligns the demand and preference analysis output by the generative AI model with the emotion tags generated by the multimodal emotion recognition module along a timeline. The server calculates emotion trends around different topics and key phrases, such as the rise and fall of emotion intensity when mentioning different content like price, features, and health.
[0268] Input: Structured analysis results from the generative AI model, multimodal emotion tag sequences, and dialogue timestamps.
[0269] Output: Comprehensive analysis results with sentiment weights and context labels.
[0270] Based on this comprehensive result, the server infers, for example, "positive sentiment is evident in the topic of healthy food" and "anxiety occurs in the topic of price sensitivity," and marks these inferences as high-value features for subsequent report and recommendation strategy generation.
[0271] Step 11: The server performs information retrieval and recommendation calculations.
[0272] Based on keywords (such as "healthy," "low sugar," "high protein," and "moderate budget") and preference tags (such as "not too sweet" and "weight management focus") extracted from the comprehensive analysis results, the server performs retrieval operations in local or remote structured data storage. The server can use inverted indexes, vectorized embeddings, and similarity calculations to perform matching score calculations on candidate items (such as goods, services, or knowledge items), and use similarity measurement methods such as cosine similarity or dot product operations.
[0273] Input: Keywords, preference tags, and sentiment weights from the comprehensive analysis results.
[0274] Output: A list of candidate entries sorted by matching degree and their corresponding explanations.
[0275] The server sorts and filters candidate items, prioritizing those that highly match the needs of the evaluated object and elicit a positive emotional response, thereby generating a recommendation set.
[0276] Step 12: The server generates and outputs real-time auxiliary information.
[0277] The server transforms the sorted candidate items into short text prompts, including the item name, key features, and brief reasons for recommendation, and generates suggested responses based on the conversation context. The server then packages these texts into a concise structure suitable for terminal display, such as prompts displayed on multiple lines.
[0278] Input: The sorted list of recommended items and the comprehensive analysis results.
[0279] Output: A collection of auxiliary information text that can be displayed on the terminal.
[0280] The server sends this auxiliary information to the terminal in real time via the network, enabling low-latency assistance to the user during the conversation.
[0281] Step 13: The terminal receives and displays auxiliary information.
[0282] The terminal receives text data containing auxiliary information from the server and parses it locally. Based on a preset display layout, the terminal overlays the recommended content as text or simple icons on the screen of the smart glasses or mobile terminal, without obstructing the user's key visual areas.
[0283] Input: Auxiliary information text transmitted by the server.
[0284] Output: Recommended prompts and suggested messages that are visually displayed on the terminal screen.
[0285] As users continue the conversation, they can quickly understand the recommended direction and expression through the content displayed on the terminal, thereby adjusting their questions or responses.
[0286] Step 14: Users can refer to the supplementary information to respond and ask follow-up questions.
[0287] Users can read the recommended prompts and suggested dialogue on their devices and choose appropriate content to explain or ask questions to the person being evaluated, based on the progress of the conversation. For example, a user can introduce a "low-sugar, high-protein oat bar" based on the prompts and then ask more detailed questions based on the person being evaluated's previous statements.
[0288] Input: Accessibility information presented by the terminal and the current dialogue context.
[0289] Output: New natural language speech and further dialogue content.
[0290] The user's comments will be collected again by the terminal and sent to the server, entering a new analysis cycle, thus forming a closed loop of continuous iteration in the system.
[0291] Step 15: The server generates and stores the post-meeting report information.
[0292] After the session ends, the server summarizes the complete dialogue history, multimodal sentiment sequence, multiple analysis results from the generative AI model, and recommendation records to construct a comprehensive report on the evaluated object. The server can then invoke the generative AI model again to summarize the entire session using the following types of prompts: You are a customer profile analysis assistant.
[0293] Please generate an analysis report for this customer based on the following complete dialogue, including: - Character profile (personality traits, areas of interest); - Long-term needs and interests; - Significant preferences (such as health, brand, price, taste, etc.); - Suggested customer service strategy (how to communicate with this customer and recommend products in the future).
[0294] The following is the dialogue: ...(Complete dialogue text)... Input: complete dialogue text, multimodal sentiment sequence, and historical analysis results.
[0295] Output: Structured report text and related tag data.
[0296] The server stores the report information along with the tags in the database and associates it with the identifier of the evaluated object for quick loading and reference in subsequent sessions, thereby completing data closure and continuous optimization at the system level.
[0297] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0298] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0299] Traditional computer-based test-taking processes often rely on fixed-rule scoring procedures or simple keyword matching algorithms. Such technical solutions have the following problems: First, servers typically only perform superficial statistical analysis or template-based comparisons of text, making it difficult to achieve in-depth semantic understanding of complex natural language responses, thus failing to effectively uncover the potential abilities and comprehensive characteristics of test takers. Second, existing systems often use one-time input and output when calling generative AI models, lacking customized prompt design and structured result management mechanisms for test scenarios. Servers struggle to reliably convert model outputs into comparable and calculable ability indicators, limiting the depth of computer technology utilization in talent selection scenarios. Third, for horizontal comparisons and statistical analyses of multiple test takers, servers often require manual export and offline processing, lacking the ability to automatically perform filtering, comparison, and statistical operations based on structured analysis results within the same system, reducing the automation and efficiency of the overall data processing chain. Fourth, existing solutions rarely utilize feedback information to dynamically update prompts and analysis processes. Servers cannot automatically adjust the input structure to generative AI models based on historical analysis results and user evaluations, making it difficult to form a closed-loop optimization between model reasoning capabilities and actual use cases, hindering the improvement of system analysis accuracy and stability over time.
[0300] Therefore, a technical solution is needed that improves the computer system architecture and data processing flow, enabling the server to: perform standardized preprocessing of raw answer information; use generative artificial intelligence models to perform semantic analysis on the answers and output structured results; complete the calculation of ability indicators, generation of visualized data and screening of candidate groups in the same computing environment; and automatically adjust the prompts and input data composition based on feedback, thereby achieving end-to-end computer technology optimization of the candidate evaluation process.
[0301] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0302] In this invention, the server includes means configured to acquire multiple response information and register the response information in storage information; means configured to extract text information from the response information registered in the storage information and perform preprocessing including deleting useless symbols, whitespace normalization, and segmentation into basic string units to generate formatted text information that can be input into a generative artificial intelligence model; means configured to combine the formatted text information with prompt statements that specify evaluation indicators and output formats to form input data for a generative artificial intelligence model and obtain analysis results about the response information by running the generative artificial intelligence model on remote computing resources or local computing resources; means configured to store the analysis results as structured data and perform ability indicator calculation, inter-indicator comprehensive evaluation, and numerical data generation for visualization based on the structured data; and means configured to generate report information indicating the characteristics and potential abilities of test takers based on the visualized numerical data and the analysis results and organize it into display data that can be output to an information display device. This enables the formation of an integrated data processing pipeline on the same server side, encompassing raw text acquisition, standardized preprocessing, generative AI model inference, structured analysis result management, and visualization report output. By programmatically controlling the prompts and input data structures, the usability and stability of generative AI models in candidate evaluation scenarios are significantly improved, reducing manual intervention and repetitive export operations. This allows for automatic comparison and filtering of multiple candidates' ability data and enhances the processing efficiency, scalability, and analytical accuracy of the entire evaluation system at the computer technology level.
[0303] A "system" refers to a collection of technical devices consisting of at least one information processing device, a storage device, and a program for performing data acquisition, data processing, and result output, used to automatically analyze and evaluate the test taker's answers.
[0304] A "server" refers to an information processing device that undertakes core data processing and control functions in a system. It can be a single physical computer, a virtual machine cluster, or a cloud computing resource, used to perform operations such as data storage, preprocessing, model calling, and result generation.
[0305] "Stored information" refers to a collection of data held by a storage device, including response information, formatted text information, analysis results, structured data, report information, and metadata such as related identifiers, timestamps, and status.
[0306] "Response information" refers to the content submitted by the test taker in response to a predetermined question or task, which is at least in text form and can be obtained from documents, web forms or other input interfaces for subsequent analysis of the test taker's characteristics and abilities.
[0307] “Textual information” refers to the content extracted from the response information and represented as a sequence of text characters. It is not limited to natural language types and is used as input data in preprocessing and generative artificial intelligence model analysis.
[0308] "Preprocessing" refers to a series of normalization operations performed by the server on text information before it is input into the generative artificial intelligence model. These operations include, but are not limited to, deleting useless symbols, whitespace normalization, encoding standardization, and segmentation into basic string units, in order to improve the stability and accuracy of subsequent analysis.
[0309] "Formatted text information" refers to text data that has been preprocessed and meets the input requirements of generative artificial intelligence models, and has a unified encoding, unified separation rules and a structured representation.
[0310] "Generative artificial intelligence models" refer to reasoning models built on machine learning methods that can generate natural language text or other data forms based on input, including artificial intelligence models that employ deep learning structures (such as sequence models or attention mechanism models).
[0311] "Prompt statements" refer to the control text content constructed by the server and input into the generative artificial intelligence model. They are used to specify the analysis task, evaluation indicators, output format and constraints, thereby guiding the model to perform targeted analysis and generation of response information.
[0312] "Input data" refers to all the data provided to the generative artificial intelligence model, including formatted text information, prompts, and necessary context information and parameter settings, which are used to trigger the model to perform inference and generation.
[0313] "Remote computing resources" refer to external computing devices or computing service environments, such as cloud computing platforms or remote processing nodes, that are connected to servers via communication networks to run generative artificial intelligence models and return analysis results.
[0314] "Local computing resources" refers to processors, accelerators, and their associated storage resources that are located in the same physical or logical environment as the server, and are used to run generative artificial intelligence models and related data processing programs locally.
[0315] "Analysis results" refer to the output of generative artificial intelligence models after processing the response information based on input data, which are results about the characteristics, abilities or other evaluation dimensions of the test takers. These results can be natural language text or structured representations.
[0316] "Structured data" refers to the data format obtained after the analysis results have been formatted and organized. It is expressed at least in part as key-value pairs, records, tables, or hierarchical structures, which facilitates server calculation, retrieval, and storage management.
[0317] "Competency indicators" refer to numerical values or levels calculated based on structured data to quantitatively describe an applicant's performance in a specific competency dimension, including but not limited to problem-solving ability, communication ability, and leadership ability.
[0318] "Comprehensive evaluation" refers to the process by which a server aggregates and analyzes the overall performance of an examinee based on multiple competency indicators and predetermined weights or rules, in order to generate a single or a few comprehensive evaluation results.
[0319] "Numerical data for visualization" refers to a set of numerical values extracted or calculated from structured data and capability indicators that are suitable for use in chart presentation, such as scores, proportions, and rankings, for generating graphical representations on display devices.
[0320] "Report information" refers to a collection of content generated based on analysis results and numerical data used for visualization, which describes the characteristics, abilities, and potential strengths of test takers, including textual descriptions, scoring summaries, and accompanying explanatory data.
[0321] "Display data" refers to the output data format that has been shaped by the server and can be directly rendered by information display devices, including structured or semi-structured data used to generate charts, text paragraphs, lists, and interactive interface elements.
[0322] "Information display device" refers to a device capable of receiving and presenting display data to a user, including but not limited to displays, terminal screens, display interfaces of mobile computing devices, or other graphical user interface devices.
[0323] "Test taker" refers to the person who submits answers during the evaluation process and is analyzed and evaluated by the system. This person can be a job seeker, a test taker, or another individual being evaluated.
[0324] A "candidate group" refers to a collection of multiple test takers used for horizontal comparison, screening, and selection within a system.
[0325] "Selection results" refers to the result data automatically generated by the server based on the analysis results, ability indicators and selection conditions of multiple candidates, which is used to represent the screening conclusions and grouping situation.
[0326] "Evaluation information" refers to human feedback data obtained from users to reflect the rationality of analysis results or the quality of system output, including confirmation, disapproval, correction opinions and other forms of evaluation signals.
[0327] In one embodiment of the present invention, the server serves as the core information processing device, the terminal as the human-computer interaction and display device, and the user as the operating subject. These three components work together to automate the analysis and evaluation of the examinee's responses. The server can be composed of a general-purpose computer, a data center node, or a cloud computing instance. Typical configurations include a multi-core central processing unit, a graphics processing unit or other accelerators, main memory, and non-volatile memory. The server runs an operating system, database management program, network server program, natural language processing program, and generative artificial intelligence model invocation program. The terminal can be any computing terminal with a display screen and input device, such as a mobile terminal, tablet device, or desktop computing device. A web browser or client application runs on the terminal for communicating with the server and displaying the evaluation results.
[0328] In this invention, the server integrates data storage, text preprocessing, model inference, and result management within the same computing environment by installing database management software (e.g., relational database management program), web server software (e.g., HTTP server program), application server framework (e.g., general web framework), natural language processing library (e.g., library for word segmentation and regular expression processing), and generative artificial intelligence model interface library. The server maintains multiple data tables or datasets in the storage device to store raw answer information, preprocessed formatted text information, analysis results from the generative artificial intelligence model, capability index data, numerical data for visualization, and report information, respectively. The server uses a structured storage format for this data, such as table records or document structure, to facilitate subsequent indexing, filtering, and calculation.
[0329] The server processes user-uploaded responses using a natural language processing program. After the user selects a file and sends it to the server via the terminal interface, the server first extracts text information from various document formats using a file parsing program. For example, when the file is a document, the server calls the document parsing module; when the file is a portable document format, the server calls the corresponding text extraction module; when the file is a spreadsheet format, the server calls the spreadsheet parsing module to merge cell contents into a text string. The server then converts the obtained text information into a predetermined character encoding and stores it as raw text fields.
[0330] The server then preprocesses the text information using a natural language processing library. It removes control characters, redundant whitespace, and irrelevant symbols using a regular expression library, and unifies the encoding of different forms of whitespace and punctuation through character normalization. The server performs word segmentation or tokenization on the text. For Chinese text, the server can use a dictionary-based and statistical method-based word segmentation algorithm to divide a continuous sequence of Chinese characters into a sequence of terms. For English or other language text, the server can use a space- and punctuation-based word segmentation algorithm, and optionally perform part-of-speech tagging or sentence boundary detection. The server encodes the preprocessed results into a structured format, such as a data object consisting of sentence arrays and term arrays, and registers it as formatted text information for use as input to generative artificial intelligence models. Because the server unifies the character set and token boundaries during the preprocessing stage, the length and distribution of subsequent model inputs are more stable, thereby reducing invalid tokens and redundant sequences, lowering the computational load, and improving inference speed.
[0331] In this invention, the server controls the behavior of the generative artificial intelligence model by generating prompt statements. The server pre-stores several prompt statement templates in its storage device, each template corresponding to a specific analysis task and output format. When constructing input data, the server concatenates or replaces formatted text information with the corresponding template. For example, for problem-solving ability analysis, the server can generate the following prompt statement: "You are an experienced human resources expert and organizational psychology consultant."
[0332] You are given a job seeker's written answers to several questions. Please evaluate their problem-solving ability based solely on these written answers.
[0333] Require: 1. Use a scale of 1 to 5 to quantitatively score problem-solving ability; 2. Explain the reasoning for the score in no more than 200 words; 3. List three specific behaviors or practices that demonstrate problem-solving abilities in the summary; 4. Output the results in JSON format, with fields including: score, reason, and key_behaviors.
[0334] The job seeker's answer is as follows: [Job seeker's answer text] The server replaces "[Job seeker's answer text]" with the corresponding formatted text information of the interviewee, forming a complete prompt statement. The server then uses this prompt statement as a control command, passing it along with the text content to the generative artificial intelligence model.
[0335] In this invention, the generative artificial intelligence model can be a language generation model built based on deep learning. For example, the server can use a multi-layer encoder-decoder or autoregressive attention network structure, including multi-layer self-attention modules, feedforward fully connected layers, and normalization layers, with the number of parameters reaching hundreds of millions or billions. During the training phase, the server uses a large number of text samples for pre-training, enabling the model to learn language patterns and semantic relationships. During the fine-tuning phase, the server can use answer samples labeled with ability dimension tags or scores to perform supervised fine-tuning of the model, enabling the model to learn to output structured text containing fields such as "score," "reason," and "key_behaviors" when given prompts. During training, the server uses a cross-entropy loss function or a weighted loss function as an error metric, updates the network weights through backpropagation and gradient descent, and can introduce techniques such as learning rate scheduling, regularization, and gradient pruning to improve the model's stability and generalization ability. During inference, the server only performs forward propagation and does not update the weights, thereby ensuring the consistency and repeatability of the output.
[0336] When the server invokes the generative AI model, it serializes the input data into a specific data format. In this implementation, it can call the model service deployed in the cloud via a remote interface or directly call the locally deployed inference engine. The server sets parameters during the invocation, such as temperature, sampling strategy, and maximum output length, to control the randomness and length of the generated text. To reduce communication load and inference latency, the server can locally truncate and filter the formatted text information based on importance, for example, retaining only sentences containing key verbs and key characters, thereby reducing the number of unnecessary input tags, reducing model computation, and improving processing speed.
[0337] After receiving the text output from the generative AI model, the server uses a parsing module to convert the output into structured data. The server specifies the output format in the prompts, ensuring the model outputs data in a manner close to structured text. The server extracts the score field, reasoning description, and key behavior list using pattern matching or a parsing library, and checks the field completeness and numerical range. If formatting errors are encountered, the server can use correction rules to correct the output string, such as adding missing parentheses or removing redundant symbols, to improve the parsing success rate. This method of constraining the output format through prompts and having the server parse it allows the generative AI model's output to adapt to the database structure and subsequent calculations, significantly reducing human intervention.
[0338] The server stores the parsed analysis results as structured data in the database. This structured data includes at least the test-taker identifier, question identifier, ability dimension name, score, reasoning text, key behavioral items, and metadata such as analysis time and model version. The server performs further data processing based on this structured data, including ability index calculation and comprehensive evaluation. The server can define normalization rules and weight parameters for each ability dimension, synthesizing scores from multiple dimensions into a comprehensive score using linear or non-linear functions. The server can also calculate the mean, variance, and ranking of multiple test-takers on the same dimension to provide statistical support for cross-sectional comparisons.
[0339] In the visualization data generation module, the server converts structured data into numerical arrays and label sets required for charts. For example, the server constructs an array for radar charts with "ability dimension name - score - maximum score" and for bar charts with "examinee identifier - ability score". This numerical data, after being organized, is stored as "numerical data for visualization" or sent directly to the terminal. Because the server internally standardizes the data structure and indicator definitions, the terminal only needs to read the fields according to the specifications to generate the graphical interface, thus significantly reducing the complexity of the front-end logic.
[0340] In this invention, the terminal performs data display and interaction functions. The terminal obtains display data and visualized numerical data from the server via a communication network, and uses a front-end rendering program to present competency charts, textual comments, and lists of key behaviors on the display screen. The terminal can use a graphics rendering library to locally draw radar charts, bar charts, line charts, etc., and label different competency dimensions with different colors or shapes. Users can select a specific examinee or competency dimension on the terminal interface. The terminal sends a corresponding request to the server, which returns the corresponding original answer fragment and the reasoning text output by the generative artificial intelligence model. The terminal highlights these elements on the interface, allowing users to intuitively see the scoring criteria.
[0341] Users can not only view individual candidate reports through the terminal, but also perform candidate group screening. The server stores the ability indicators and comprehensive scores of multiple candidates in its storage device. When responding to terminal requests, the server performs conditional filtering and sorting operations on the specified candidate group, such as selecting candidates whose ability scores are above a certain threshold and whose comprehensive evaluation ranks among the top few. The server sends the selection results to the terminal in list form for display and further processing. Because the server calculates directly based on structured ability indicators without requiring manual export to external tools, the entire screening process is completed internally by the computer, reducing cross-system data transmission and manual processing errors, and achieving a closed-loop calculation process.
[0342] In this invention, the server can also dynamically update the prompt statement parameters based on user feedback and selection results. When a user confirms or raises objections to an analysis result on the terminal, the terminal uploads the feedback information to the server. The server associates this feedback with the corresponding analysis results to estimate the consistency between the current prompt statement and model output and the user's evaluation. The server can statistically analyze the consistency rate of a prompt template across different samples. If it falls below a predetermined threshold, the server automatically adjusts the weight descriptions or output requirements of the evaluation items in the template, such as strengthening the focus on logical structure or weakening the weight of writing style. The server can also learn a mapping rule based on historical data to map the original score output by the generative artificial intelligence model to a corrected score that better matches human evaluation, thereby reducing systematic bias. Through this feedback-driven prompt statement adjustment mechanism, this invention achieves dynamic optimization of the input composition of the generative artificial intelligence model, enabling the overall analysis accuracy to improve over time without changing the underlying model parameters.
[0343] The technical advantages of this invention lie in the fact that the server not only replaces some of the manual reading and scoring work, but more importantly, it improves the internal data flow and computational path of the computer through specific data structures, preprocessing algorithms, prompt statement construction rules, structured parsing, and indicator calculation processes. By performing standardization and truncation on the input side, the server improves the computational efficiency of generative artificial intelligence models; by implementing format constraints and parsing rules on the output side, it achieves a stable conversion from free text to structured data; and by using unified ability indicators and database structures in the intermediate layer, it improves the comparability and retrieval of data from multiple test takers within the same system. These improvements directly lead to increased processing speed, simplified storage management, reduced communication data volume, and improved overall evaluation accuracy. They represent improvements to computer technology itself, rather than merely the automation of business processes.
[0344] In other implementations, the server can employ different generative AI model structures. For example, the server can use an encoder-decoder architecture to encode the response text into a dense vector sequence, and then decode it conditionally with the task description and prompts. Alternatively, it can use a multi-task learning framework to simultaneously output scores and reasons for multiple capability dimensions within the same network, thereby reducing the time overhead of multiple inferences. The server can also combine traditional feature engineering methods with deep learning models, such as extracting statistical features like response length, sentence complexity, and keyword density, and concatenating them with the output features of the generative AI model before inputting them into a post-classifier or regressor to further improve scoring stability. The server can dynamically switch inference modes between remote and local computing resources based on the application scenario, such as automatically switching to a local lightweight model under high concurrency and using a remote large-scale model under low concurrency to balance response time and analysis quality.
[0345] In another implementation, the terminal can not only passively receive and display data, but also visually edit the prompt statement templates through a graphical interface. Users can fine-tune the text content of the prompt statements on the terminal, such as adding new evaluation items or modifying output descriptions. The terminal then synchronizes the new template to the server. The server manages these modifications, recording the version of the prompt statement used for each analysis result during subsequent analysis, facilitating backtracking and comparison of the effects of different templates. In this way, the present invention allows for the evolution of analysis strategies in a configurable manner without changing the core program code, improving the maintainability and scalability of the system during long-term operation.
[0346] In summary, this invention establishes a dedicated data processing and interaction mechanism for generative artificial intelligence models among the server, terminal, and user, forming a complete technical chain from answer information collection, text preprocessing, model inference, result structuring, ability index calculation, visualization, and feedback optimization. The data flow and algorithmic steps between modules within the server are meticulously designed, enabling the system to maintain high efficiency, stability, and accuracy when processing large amounts of test-taker text data, thus substantially improving the traditional test-taker evaluation process at the computer technology level.
[0347] use Figure 13 The processing flow is explained.
[0348] Step 1: Users upload their answers on the terminal. The user logs into the system interface on the terminal and selects the "Import Answers" function. The input is one or more files (e.g., document files, portable document files, plain text files, or spreadsheet files) containing the examinee's answers, stored locally on the terminal. After the user selects the target file through the file selection dialog box on the terminal interface, the terminal packages the selected file along with metadata such as the user identifier, examinee identifier, and question identifier into an HTTPS request body and sends it to the server. The terminal simultaneously displays upload progress and status indicators on the interface. The output is an upload request message sent to the server.
[0349] Step 2: The server receives and parses the upload request. The server, primarily a web service program, receives upload requests from terminals via an HTTP server component. The input is the request message sent by the terminal in step 1, containing the raw file's binary data and accompanying identification information. The server first verifies the user's identity and permissions, then parses the request header and body, separating the file content and metadata. The server determines the file format based on the file type field or extension and selects an appropriate parsing module for subsequent text extraction. After parsing, the server creates a record for each file in the database, assigning an internal response identifier and an examinee identifier. The output consists of the raw file entry stored in the database and its associated metadata entry.
[0350] Step 3: The server extracts text information from the file. The server invokes a file parsing program to read file data from the original file entries registered in step 2. The input consists of the original file's binary data and its type information stored on the storage device. The server selects the appropriate parsing library based on the file type: for document files, it calls the document parsing module; for portable document files, it calls the text extraction module; and for spreadsheet files, it calls the table parsing module. The server performs format parsing in memory, converting complex document structures into plain text strings while removing layout information and retaining only the text content. Data processing includes page separator handling, table cell merging, and paragraph joining. The server writes the obtained text content as "raw text information" into the corresponding field of the database and maintains its association with the answer identifier. The output is the raw text information persistently stored.
[0351] Step 4: The server cleans and standardizes the text information. The server reads the raw text information generated in step 3 from the database as input and calls natural language processing and regular expression libraries for preprocessing. The server performs data processing operations on the input text: deleting control characters and invisible characters; compressing multiple consecutive whitespace characters into a single space or a uniform newline character; removing meaningless symbols and special tags unrelated to semantics. The server also unifies character encoding to a predetermined encoding format to ensure consistency in subsequent processing. The server generates the "cleaned text information" in memory and then writes it back to the corresponding field in the database. The output is cleaned text information with unified encoding and reduced noise.
[0352] Step 5: The server performs word segmentation and tokenization on the cleaned text information. The server takes the cleaned text information output from step 4 as input and calls a natural language processing library to perform word segmentation or tokenization operations. First, the server selects the corresponding word segmentation algorithm based on the language identifier; for Chinese text, the server uses a dictionary-based and statistical word segmentation module to divide continuous character sequences into word sequence sequences; for text in other languages, the server uses spaces and punctuation for tokenization. The server can further perform sentence boundary detection, segmenting the entire text into a list of sentences. The server represents each sentence as an array of several word terms, forming a structured representation. Data processing includes adding position and sentence indices to each word term for subsequent truncation and filtering. The server stores this structured result as a dedicated field in the database called "formatted text information." The output is formatted text information that can be directly used as input for generative artificial intelligence models.
[0353] Step 6: Users select the analysis object and analysis type on the terminal. Users access the "AI Analysis" function page on the terminal interface. The input consists of a list of registered test takers and a list of questions in the system; the terminal retrieves these lists from the server and displays them on the interface. Users select one or more test takers using interface controls and choose the desired analysis type, such as "Problem-Solving Ability Analysis," "Leadership Analysis," or "Comprehensive Ability Report." The terminal converts the user's selections into structured request data containing test taker identifiers, question identifiers, and analysis type identifiers, and sends it to the server via HTTPS. The output is an analysis request message containing the analysis parameters.
[0354] Step 7: The server prepares the input text for the generative artificial intelligence model. The server receives the analysis request from step 6, using the candidate identifier and question identifier from the request as input, and retrieves the corresponding formatted text information from the database. Depending on the analysis type, the server concatenates or segments the text information from multiple questions in a predetermined order to form the candidate's complete answer text. Based on the maximum input length allowed by the generative AI model, the server performs data operations on the text, including: counting the total number of tokens, truncating sentences according to boundaries, retaining sentences containing key verbs or keywords, and discarding redundant parts. The server generates a length-limited, information-density "analysis input text." The output is an optimized analysis input text for each candidate.
[0355] Step 8: The server constructs prompt statements and generates model input data. Based on the aforementioned analysis input text, the server reads a template matching the current analysis type from the prompt template library. The input consists of an analysis type identifier, the analysis input text, and the prompt template content. The server performs text replacement and concatenation operations in memory, replacing placeholders in the template with the actual analysis input text. For example, the server generates the following prompt: "You are an experienced human resources expert and organizational psychology consultant."
[0356] You are given a job seeker's written answers to several questions. Please evaluate their problem-solving ability based solely on these written answers.
[0357] Require: 1. Use a scale of 1 to 5 to quantitatively score problem-solving ability; 2. Explain the reasoning for the score in no more than 200 words; 3. List three specific behaviors or practices that demonstrate problem-solving abilities in the summary; 4. Output the results in JSON format, with fields including: score, reason, and key_behaviors.
[0358] The job seeker's answer is as follows: [Job seeker's answer text] The server replaces "[Job seeker's answer text]" with the analysis input text generated in step 7 to obtain the complete prompt statement. The server combines the prompt statement and necessary model parameters (such as maximum output length and temperature parameters) into a "model input data" structure. The output is an input data object that can be directly passed to a generative artificial intelligence model.
[0359] Step 9: The server invokes the generative artificial intelligence model and obtains the analysis output. The server uses the model input data generated in step 8 as input and invokes the generative AI model inference service deployed on remote or local computing resources. The server serializes the input data into a request format acceptable to the model service and sends it via a network interface; if the model is local, the server invokes it directly through the local inference engine interface. The server sets timeout and retry policies during the invocation process to ensure reliability. Internally, the generative AI model performs forward propagation computation, encoding the prompts and response text into vector representations, and generating an output label sequence after multiple layers of attention operations and nonlinear transformations. The server receives the text output returned by the model as the "raw analysis output." The output is the raw analysis output text containing information such as ratings, reasoning descriptions, and behavior lists.
[0360] Step 10: The server parses and analyzes the output and generates structured results. The server takes the raw analysis output text from step 9 as input and uses pattern matching or parsing functions to perform structured processing on the text. Based on the output format agreed upon in the prompt statement, the server extracts the content of each field, such as locating the numerical value of "score," the textual description of "reason," and the multiple behavioral descriptions corresponding to "key_behaviors." Data processing includes: numerical conversion and range checking of the score values, splitting the behavior list by delimiters, and trimming the reason text. The server performs error detection during parsing; if the format is detected as not conforming to expectations, the server can correct it according to predetermined rules or re-initiate the model call if necessary. After parsing, the server organizes the fields into structured records and stores them in the analysis results table of the database. The output is the structured analysis result data corresponding to each answer.
[0361] Step 11: Server computing power metrics and generate visualized numerical data. The server reads the structured results generated in step 10 from the analysis results table as input and performs numerical calculations on the scores of multiple ability dimensions. Following pre-defined rules, the server averages or weights the scores of multiple questions for the same test-taker to obtain a comprehensive score for each ability dimension. The server can further normalize these scores, mapping them to a uniform score range. The server can also sort, calculate the average and variance of the scores of multiple test-takers, and generate indicators suitable for cross-sectional comparisons. The output of the data processing includes: an array of ability dimension scores for each test-taker, a comprehensive score, ranking information, and a numerical matrix used to generate charts. The server stores or caches these values as "visualized numerical data" for rapid response to terminal requests.
[0362] Step 12: The server generates report information and displays data. The server takes competency indicators and analysis results as input and calls the report generation module to construct the candidate's report information. Using template matching, the server inserts competency dimension names, scores, reasoning text, and key behaviors into a predefined report structure, generating a report that includes textual descriptions and structured fields. The server organizes the text paragraphs and chart data in the report into a display data structure suitable for terminal rendering, such as objects containing elements like titles, paragraph lists, chart data blocks, and interactive link identifiers. The output is the report information and corresponding display data for each candidate.
[0363] Step 13: Terminal acquires and displays reports and charts The terminal requests the display data generated in step 12 from the server, inputting the test taker's identifier and the user's viewing instructions. The server returns a response containing visualized numerical data and report information. Upon receiving the response, the terminal uses its front-end rendering engine to plot the competency dimension scores as a radar chart or bar chart, and displays the report text in segments on the screen. Based on the interaction identifiers in the displayed data, the terminal sets clickable areas for each competency dimension. When the user clicks on an item, the terminal requests the corresponding detailed analysis results and original answer fragments from the server, which are then displayed in a pop-up window or sidebar. The output is a visualized report interface presented on the screen.
[0364] Step 14: Users perform filtering and feedback operations on the terminal. Users view charts and reports of multiple candidates on the terminal interface, using this content as input to make screening decisions and provide evaluation feedback. Users can select options such as "Pass," "Keep," or "Eliminate" on the interface, and can also add "Approved," "Disapproved," or comment text to a specific analysis result. The terminal receives user actions, encapsulates the screening results and feedback information into a request, and sends it to the server. Upon receiving the request, the server writes this information to the selection result table and feedback table in the database, providing a basis for subsequent statistics and adjustments to prompts. The output is persistently stored selection result data and evaluation feedback data.
[0365] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0366] In existing computer-based personnel evaluation and safety management technologies, simple rule matching or keyword retrieval is usually performed only for text-based question and answer, making it difficult to fully utilize large-scale language models for deep semantic understanding and to uniformly model the contextual information of multi-turn dialogues. At the same time, sentiment analysis is often conducted independently using single-frame facial expression recognition or single-sentence emotion scoring, lacking temporal correlation and unified evaluation with text semantics and ability features, resulting in an insufficiently refined portrayal of the characteristics, abilities, and emotional changes of the evaluated individuals.
[0367] Furthermore, traditional systems typically treat "question-answering presentation," "emotion recognition," and "result display" as separate functional modules at the computer architecture level, failing to form a unified data flow and control flow centered on the server side. On the one hand, servers lack a mechanism to automatically generate adaptive prompts based on different task scenarios, making it impossible to dynamically drive generative AI models to output structured feature data for different problems; on the other hand, servers often do not perform unified timestamp alignment and statistical modeling on multimodal data (language information, moving image information, audio information) from terminals, making it difficult to construct the correlation between "features-capabilities-emotions-risks" at the computational level in a timely and accurate manner.
[0368] Furthermore, when written materials contain missing or ambiguous information, existing technologies often rely on manual questioning. Servers lack the ability to automatically identify "evaluation items that are not fully understood" and automatically construct additional questions and prompts. They are unable to adaptively generate new data collection and analysis tasks during the calculation process, resulting in insufficient automation and scalability of the system.
[0369] Therefore, at the computer technology level, a new system architecture and data processing flow are needed to enable servers to: (1) Using a generative artificial intelligence model as the core, perform deep semantic analysis on multi-turn text dialogues from the terminal and output machine-readable structured characteristic data; (2) Using an emotion recognition engine as an aid, time-series emotion estimation is performed on moving image information and audio information, and the results are precisely aligned with the text analysis results on the time axis; (3) Perform multimodal data fusion and statistical operations uniformly within the server, automatically calculate character characteristics, ability tendencies, emotional tendencies and risk levels, and generate evaluation reports that can be directly consumed by other computer programs. (4) Based on written information and existing evaluation results, additional questions and prompts are automatically generated on the server side to drive the generative artificial intelligence model to perform supplementary analysis, thereby improving the completeness and accuracy of the evaluation in a closed loop through computation. (5) In security or real-time monitoring scenarios, based on continuous input information and behavior records, the risk level is updated in real time on the server side and alarms are automatically triggered, enabling the computer system to evolve from a passive recording tool into an active decision support platform.
[0370] The purpose of this invention is to provide a novel system architecture that utilizes centralized server control for prompt statement generation, multimodal data alignment and fusion, and real-time risk assessment. This improves the computer's ability to analyze human characteristics and emotions at both the algorithm and system implementation levels, thereby enhancing the automation, robustness, and scalability of human-computer interactive assessment and security monitoring tasks.
[0371] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0372] In this invention, the server includes: a device for acquiring input language information and motion image information from an evaluated object for a predetermined topic, with the cooperation of a terminal, and providing the language information to a generative artificial intelligence model to perform dialogue processing according to the dialogue process; a device for generating prompt statements based on the language information for inputting the evaluated object's response into the generative artificial intelligence model, and sending the prompt statements along with the language information to the generative artificial intelligence model to obtain structured data containing at least characteristic information and ability information; an analysis device for extracting facial expression information and audio feature quantities based on the motion image information and audio information, and using an emotion recognition engine to generate emotion estimation data representing multiple emotion indicators and their temporal changes; and a device for... The apparatus comprises: a device for mapping the structured data to the emotion estimation data and integrating characteristic information, ability information, and emotion indicators corresponding to multiple questions to calculate an evaluation result that includes at least the person's characteristics, ability tendencies, emotion tendencies, and risk level of the evaluated object; a device for generating an evaluation report that includes at least ability indicators, emotion indicators, person's characteristics, and risk indicators based on the evaluation result and providing the evaluation report to the terminal in a predetermined document format; and a device for continuously acquiring new input information or behavioral record information related to the evaluated object from the terminal, repeatedly performing parsing processing in real time using the generative artificial intelligence model and the emotion recognition engine to update the risk level, and generating a warning message and sending it to the terminal when the risk level exceeds a predetermined threshold. This allows for the formation of an integrated multimodal data processing pipeline within the server, centered around a generative artificial intelligence model and an emotion recognition engine. This enables automatic generation of prompts, extraction of structured features, time-series emotion estimation, and real-time risk updates. As a result, the computer system can comprehensively evaluate the subjects being evaluated with higher accuracy and efficiency, and provide immediate alerts based on the calculation results in security scenarios. This, in turn, improves the overall processing capabilities and technical effectiveness of computers in personnel evaluation and risk management tasks.
[0373] A "system" refers to a collection of computer-based devices consisting of servers, terminals, and program modules running on them, used to perform personnel evaluation and risk management processes.
[0374] A "server" refers to a computing device used to centrally perform processing such as data reception, prompt generation, generative artificial intelligence model invocation, emotion recognition processing, data fusion, evaluation result calculation, and report generation.
[0375] A “terminal” refers to an electronic device that communicates with a server to present problems to the evaluated object, collect language information, moving image information and audio information, and receive and display evaluation reports or warning information.
[0376] "The evaluated object" refers to an individual whose characteristics, abilities, emotional tendencies, and risk levels are analyzed and evaluated in the system.
[0377] "Language information" refers to textual data obtained by the evaluated subject through text input or speech recognition, which reflects the content of the evaluated subject's responses during the dialogue process.
[0378] "Dynamic image information" refers to image sequence data acquired through an image acquisition device that includes changes in the facial expressions and postures of the evaluated object over time.
[0379] "Audio information" refers to sound data containing the speech characteristics of the evaluated object, acquired through an audio acquisition device.
[0380] "Generative artificial intelligence models" refer to artificial intelligence models that perform semantic understanding and content generation based on input language information and prompts, and output structured or semi-structured data for the analysis of character characteristics and abilities.
[0381] "Prompt statements" refer to instructional text generated by the server and sent to the generative artificial intelligence model to specify the analysis objectives, output format, and evaluation dimensions.
[0382] "Structured data" refers to machine-readable data output by generative artificial intelligence models, which contains characteristic information, capability information and other evaluation indicators, and is organized in a fixed field or hierarchical structure.
[0383] "Characteristic information" refers to parameterized information extracted from the language information and dialogue behavior of the evaluated object, used to represent personality traits, behavioral patterns and other characteristics of the person.
[0384] "Capability information" refers to parameterized information extracted from the responses of the evaluated individuals, used to represent capability dimensions such as problem-solving ability, communication ability, and collaboration ability.
[0385] An "emotion recognition engine" refers to a program module or device that analyzes moving image and audio information to identify different emotion categories and their intensity.
[0386] "Facial expression information" refers to data extracted from moving image information that represents the visual characteristics of the evaluated object, such as the state of facial muscles and changes in facial expressions.
[0387] "Audio features" refers to the set of numerical parameters extracted from audio information to represent acoustic features such as pitch, volume, speech rate, and timbre.
[0388] "Emotional indicators" refer to parameters used to indicate the emotional state (e.g., positive, negative, tense, calm, etc.) and intensity of the evaluated object at a specific point in time or within a time interval.
[0389] "Emotion estimation data" refers to a dataset generated by an emotion recognition engine based on facial expression information and audio features, representing various emotion indicators and their changes over time.
[0390] "Time information" refers to the timestamps or timelines corresponding to language information, moving image information, audio information, and analysis results, and is used to perform time alignment of multimodal data within the system.
[0391] "Evaluation results" refer to comprehensive assessment information calculated by the server based on the integration of structured data and sentiment estimation data, which includes at least the characteristics, abilities, sentiments, and risk levels of the individual.
[0392] An "evaluation report" is a document data generated based on evaluation results and used to provide information to the terminal, including at least ability indicators, emotional indicators, personality traits, and risk indicators.
[0393] "Competency indicators" refer to the numerical values or levels used in evaluation reports to quantitatively represent the performance of the evaluated object in various competency dimensions.
[0394] "Emotional indicators" refer to numerical values or levels used in evaluation reports to quantify the emotional state and changing characteristics of the evaluated object.
[0395] "Personal characteristics" refers to the comprehensive characteristic information used in the evaluation report to describe the personality traits, behavioral tendencies, etc. of the evaluated person.
[0396] "Risk indicators" refer to numerical values or levels used in evaluation reports to quantify the potential risk level of the evaluated object.
[0397] "Risk level" refers to the quantitative evaluation of the degree of adverse effects or security threats that the evaluated object may bring in a specific scenario, based on the results of multimodal analysis.
[0398] "Written information" refers to static materials that are primarily text-based, including but not limited to resumes, application materials, and questionnaire answers.
[0399] "Behavioral record information" refers to data continuously acquired during system operation that reflects the operational behavior or activity trajectory of the evaluated object.
[0400] "Supplementary analysis results" refer to the analytical data obtained by re-analyzing additional questions and corresponding answers using a generative artificial intelligence model, which is used to supplement the insufficient information in the original evaluation items.
[0401] "Statistical information" refers to data obtained by summarizing, averaging, and calculating variance of characteristic information, ability information, and emotion indicators obtained from multiple dialogue units in terms of time or frequency dimensions.
[0402] "Characteristic stability" refers to the degree to which the evaluation value of the same character's characteristic dimension remains relatively consistent across different time periods or different dialogue units.
[0403] "Consistency of emotional response" refers to the degree to which the patterns of change in emotional indicators are similar or predictable under different situations or problems.
[0404] "Warning messages" refer to notification data generated by the server and sent to the terminal when the risk level exceeds a predetermined threshold, used to prompt attention or action.
[0405] In various embodiments of this invention, the server, terminal, and user each assume different functional roles, achieving the acquisition, transmission, processing, and display of multimodal data through hardware and software collaboration, and constructing an integrated data processing system around a generative artificial intelligence model and an emotion recognition engine. The embodiments of this invention are described below in conjunction with typical hardware structures, software components, data structures, and algorithm flows.
[0406] I. Overall Hardware and Software Composition of the System The server can employ information processing devices with high-performance computing capabilities, such as rack-mounted computers configured with multi-core central processing units and graphics processing units. The server can run a Unix-like operating system and install deep learning frameworks and inference service environments.
[0407] The server can use the following software components: The server uses Python as its primary development language.
[0408] The server uses a web application framework to provide interface services to the terminal.
[0409] The server uses deep learning frameworks (such as tensor operation-based frameworks or tensor graph-based frameworks) to train and deploy generative artificial intelligence models.
[0410] The server uses natural language processing libraries (such as libraries for word segmentation, part-of-speech tagging, and dependency parsing) to perform auxiliary text preprocessing.
[0411] The server uses emotion recognition engines, such as facial expression recognition libraries based on convolutional neural networks and recurrent neural networks, speech emotion recognition libraries based on acoustic features, or emotion analysis application interfaces based on cloud services.
[0412] The server uses report generation components (such as template-based document generation libraries or markup language-based document conversion tools) to generate evaluation report documents.
[0413] The terminal can be a smartphone, tablet computer, wearable display device, desktop computer, or service robot, etc., and is equipped with a display device, input device, camera device, and audio acquisition device. The terminal can run a browser or dedicated application and exchange data with the server through secure communication protocols.
[0414] Users can be either the object being evaluated or the viewer of the evaluation results. Users view questions, enter answers, and receive system evaluation results and warnings through the interface on the terminal.
[0415] II. Structure and Learning Methods of Generative Artificial Intelligence Models The generative artificial intelligence model used in this invention can be built on a multi-layered Transformer architecture. The server may employ the following technical features: The server uses a multi-head self-attention mechanism to model the input sequence, which includes segmented and encoded language information and prompts.
[0416] The server sets up several encoding and decoding layers for generative artificial intelligence models. Each layer contains a self-attention sub-layer, a feedforward fully connected sub-layer, and a residual normalization structure.
[0417] The server assigns an embedding vector to each tag and overlays positional encoding so that the model can distinguish sequence information.
[0418] The server uses supervised learning during the training phase to jointly optimize multiple task objectives. The server can specify the following typical loss functions: The server uses cross-entropy loss for text generation tasks and multi-label binary cross-entropy or mean squared error loss for structured label prediction tasks.
[0419] The server updates the network weights using optimization algorithms (such as adaptive moment estimation or momentum gradient descent).
[0420] The server trains the model using mini-batch gradient descent, and performs joint training using a large amount of question-and-answer data, character attribute label data, and emotion label data.
[0421] During training, the server can use data augmentation techniques, such as synonym replacement, sentence transformation, and multilingual alignment, to improve the model's robustness to different expressions.
[0422] During the inference phase, the server uses frozen model weights, performs only forward computation, and does not update parameters. The server performs batch inference on the input through floating-point operations and tensor operations to improve throughput and response speed.
[0423] By explicitly adding labels for "characteristic dimension," "ability dimension," and "emotional tendency dimension" during the training phase, the server enables the generative AI model to automatically learn feature subspaces related to these dimensions within its internal representation space. Therefore, when the server provides specific prompts during inference, the model can focus its computations within these subspaces, quickly outputting structured data for the corresponding dimensions, thereby improving computational efficiency and analytical accuracy.
[0424] III. Structure and Feature Processing of Emotion Recognition Engines In this invention, the server uses an emotion recognition engine to process moving image and audio information.
[0425] The server can employ convolutional neural networks or convolutional-residual network structures in the image path to extract facial regions and expression features from frame images. The server can obtain key points through a pre-trained facial detection model and then output multiple emotion probabilities through a dedicated expression classification network.
[0426] The server can first extract acoustic features such as Mel frequency cepstral coefficients, fundamental frequency, energy, zero-crossing rate, and spectral centroid from the audio path, and then input these features as time series data into a recurrent neural network, a gated recurrent unit network, or a one-dimensional convolutional network to output the emotion category and intensity.
[0427] The server can map multimodal features into a unified sentiment index vector by fusing image and audio features near the same timestamp, using simple weighted averaging, attention-weighted fusion, or small fully connected fusion networks.
[0428] Internally, the server stores a sentiment indicator vector and its timestamp for each time slice, forming a time series of sentiment estimation data. The server can perform smoothing or peak detection on the time series to suppress noise and highlight significant sentiment shifts.
[0429] Through the above structure, the server can not only identify the basic emotions of the evaluated object at a single moment, but also obtain the trajectory of emotion changes over time, thus providing high temporal resolution input for subsequent characteristic and risk analysis.
[0430] IV. Data Structures and Internal Data Flow In this invention, the server manages multimodal data using a unified data structure.
[0431] The server establishes a session identifier for each dialogue session and assigns a question identifier and an answer identifier to each question and answer within that session.
[0432] The server establishes a text data structure for language information, including fields such as the original text, word segmentation sequence, embedding index, and language detection results.
[0433] The server establishes a video data structure for the moving image information, including fields such as frame timestamp, image path or cache reference, key point coordinates, and facial expression vector.
[0434] The server establishes an audio data structure for the audio information, including fields such as sampling rate, frame division information, and acoustic feature vectors.
[0435] The server creates a label data structure for the structured data output by the model, which includes feature dimension names, capability dimension names, corresponding scores, confidence levels, and brief descriptive text.
[0436] The server establishes a time-series structure for the sentiment estimation data, with each element recording a timestamp, multidimensional sentiment indicators, and intensity values.
[0437] Throughout the data flow process, the server consistently associates the aforementioned data using session identifiers and timestamps. During the fusion phase, the server aligns structured data and sentiment estimation data based on time information to enable comprehensive calculations on the same issue or within the same time interval. This data management approach, based on precise timestamps and unified identifiers, allows the server to efficiently perform multimodal data fusion internally, reducing redundant scanning and computation, thereby improving overall processing speed and reducing storage and communication overhead.
[0438] V. Prompt Statement Generation and Rule Design Without Manual Interference In this invention, the server automatically generates prompts to drive the generative artificial intelligence model to output results in the expected format.
[0439] The server selects a template from a predefined prompt template library based on the question type, target dimension, and output structure, and populates the template with dynamic content (including question text, user answer summary, required output field names, etc.) to generate specific prompt statements.
[0440] The server can be configured with the following example prompt statements: The server can generate the following prompt statement: "Please analyze the following candidates' answers and assess their problem-solving and teamwork skills. Output a score for problem-solving (0-10), a score for teamwork (0-10), and a brief explanation." "As a human resources assessment assistant, based on the following multi-round dialogue, please summarize the main personality traits, strengths, and potential risks of this person, with a word count of no more than 800 words." "Based on the following answers, extract 3 personality trait tags and 3 ability trait tags, and briefly explain the basis for extraction." In supplementary analysis scenarios, the server can automatically identify dimensions as "evaluation items with insufficient understanding" based on low confidence levels or large fluctuations in certain dimensions within the evaluation report, and construct targeted prompts and follow-up questions. For example, the server can generate: "Based on the current evaluation results, there is insufficient information regarding the candidate's leadership capabilities. Please design three open-ended questions to further explore the candidate's leadership performance." And the prompts used to analyze additional answers: "Based on the candidate's answers to the above leadership-related questions, please reassess their leadership level and give a score of 0-10, along with the reasons for the analysis." Through this rule-based and template-based prompt generation mechanism, the server no longer relies on manual writing of instruction text for each analysis. Instead, it automatically constructs appropriate prompts within the computer based on data characteristics and evaluation results, thus forming a non-manual and non-conventional model-driven approach that improves the system's automation and scalability.
[0441] VI. Server-Side Feature Fusion and Risk Calculation Algorithm In this invention, the server performs multi-dimensional fusion of structured data and sentiment estimation data from generative artificial intelligence models.
[0442] The server performs a weighted average of the ability scores and characteristic labels of the same evaluated object across multiple questions. The weights can be adaptively adjusted based on the importance of the questions, model confidence, and sentiment stability.
[0443] The server measures the stability of ability performance by using the variance of statistical ability scores and the stability of emotional responses by using the fluctuation range of statistical emotion indicators.
[0444] The server can build risk assessment modules based on a combination of rules and machine learning models. For example, the server can set the following rules: The risk score increases when the integrity-related trait score is low and strong negative emotions are present on key issues.
[0445] When emotional reactions fluctuate frequently and drastically within a short period of time, the risk of emotional instability increases.
[0446] At the machine learning level, the server can use gradient boosting trees or shallow neural networks to take feature rating vectors and sentiment feature vectors as input and output a normalized risk level.
[0447] When training a risk model, the server can use historical labeled data (such as risk event markers in past cases) and adopt a supervised learning approach to train the model by minimizing cross-entropy or mean squared error loss, thereby enabling the model to learn to identify complex risk patterns in a high-dimensional space.
[0448] By integrating the aforementioned features with risk calculation algorithms, the server does not simply add or subtract based on business rules. Instead, it uses multidimensional features and statistical information to construct mathematical and learning models, thereby forming an accurate quantification and prediction capability for individual risks within the computer, achieving error reduction and improved recognition rate.
[0449] VII. Evaluation Report Generation and Technical Presentation In this invention, the server organizes the evaluation results into a report suitable for both machine and human use.
[0450] The server internally defines report templates, including structural sections (such as summary, feature analysis, capability analysis, sentiment analysis, and risk assessment), chart sections (capability radar chart, sentiment time curve, and risk bar chart), and detailed descriptions.
[0451] The server directly maps structured data into numerical points in the chart and uses a graphics library to generate visualization elements, thereby avoiding repetitive calculations and drawing on the terminal and reducing the processing burden and number of communication round trips on the terminal.
[0452] When the server generates the document, it embeds the text content and charts as a unified document object, thus presenting the complete technical analysis results in a single file.
[0453] The server can support multiple output formats, such as printable documents and marked documents that can be embedded in other systems, allowing other computer systems to automatically parse and continue subsequent processing.
[0454] By completing most of the data processing and formatting on the server side, the terminal only needs to perform simple presentation, which greatly reduces the computing pressure on the terminal, improves the overall response speed of the system, and reduces the network bandwidth usage.
[0455] VIII. Technical Effects and Improvements in Computer Technology Through the aforementioned structured and modular implementation, the server brings about numerous improvements at the computer technology level: The server uses a unified data structure and timestamp alignment mechanism to enable language information, moving image information and audio information to be matched simultaneously in a single scan, thereby reducing multiple traversals and repeated parsing and improving processing efficiency.
[0456] The server uses a pre-trained generative AI model and emotion recognition engine to solidify the complex semantic understanding and emotion recognition process into efficient matrix operations and convolution operations, enabling rapid parallel processing of massive amounts of data, which is significantly better than the inefficient method of manual judgment.
[0457] The server standardizes and automates the model calling process through automatic prompt generation and dynamic supplementary analysis mechanisms, reducing the manual costs of parameter configuration and instruction writing, avoiding misuse of models or inconsistent outputs, thereby improving the stability and repeatability of the overall evaluation process.
[0458] The server constructs a risk assessment model by integrating characteristic vectors, ability vectors, and emotion vectors, transforming risk judgment from vague subjective impressions into measurable numerical calculations, thereby improving assessment accuracy. Furthermore, it can gradually optimize the judgment boundary by learning from historical data, reducing the false positive and false negative rates.
[0459] By completing chart drawing and report generation internally, the server reduces front-end and back-end data format conversion and complex logic, enabling data to be processed once and reused in multiple places, thereby improving system resource utilization and shortening response time.
[0460] Users do not need to understand the internal model and algorithm details in the system of this invention. They can obtain high-precision evaluation results and real-time risk warnings simply by interacting naturally through the terminal. The terminal only serves as a data acquisition and display terminal, and its computational load is significantly reduced.
[0461] Therefore, this invention achieves improvements at the computer technology level for personnel evaluation and risk management tasks by systematically designing generative artificial intelligence models, automatic generation of prompt statements, emotion recognition engines, multimodal data fusion, and risk calculation algorithms within the server, rather than simply automating the traditional manual assessment process.
[0462] use Figure 14 The processing flow is explained.
[0463] Step 1: Users view questions and enter answers on the terminal. Users view preset topics and corresponding questions sent by the server on the terminal interface, such as "Please describe an experience you had solving a difficult problem." Users type their answers in the text input box or answer verbally through the microphone, and the terminal uses local speech recognition to convert speech into text. The input consists of the user's original text answer and optional audio and video data; the output consists of language information (text), audio information, and animated image information temporarily stored in the terminal. The terminal performs basic text cleaning (removing extra spaces and control characters) and encodes and compresses the audio and video in preparation for subsequent transmission.
[0464] Step 2: The terminal sends multimodal response data to the server. The terminal packages the language information, audio information, and motion image information obtained in step 1 into a message with a session identifier and timestamp, and sends it to the server via a secure communication protocol. The input consists of locally stored multimodal data and session metadata; the output is a request message sent to the server. Before sending, the terminal encapsulates the text into structured fields and encodes the audio and video into a predetermined format, thereby reducing bandwidth consumption and ensuring correct server parsing.
[0465] Step 3: The server receives and preprocesses multimodal data. The server receives request messages from the terminal and parses them to obtain metadata such as language information, audio information, moving image information, and timestamps. The input is the raw multimodal data sent by the terminal; the output is a standardized data structure internal to the server, including cleaned text, audio segmented by frame, and image sequences extracted by frame. The server performs encoding standardization and illegal character filtering on the text, performs framing and feature extraction on the audio, and extracts keyframes and detects facial regions on the video, providing standardized input for subsequent emotion recognition and generative artificial intelligence model calls.
[0466] Step 4: The server generates prompts for semantic analysis. Based on the type of the current problem and the target evaluation dimensions (such as problem-solving ability, teamwork ability), the server selects an appropriate template from the prompt template library, fills in the user's response text and the required output fields, and generates a specific prompt statement. The input includes the problem type, user language information, and the prompt template; the output is a complete prompt statement text. In this step, the server performs data processing such as string concatenation and placeholder replacement, transforming the abstract evaluation target into natural language instructions that the generative artificial intelligence model can understand.
[0467] Step 5: The server invokes a generative artificial intelligence model to perform text semantic analysis. The server inputs the language information from step 3 along with the prompts generated in step 4 into the generative AI model. The input consists of the prompts and encoded user responses; the output is structured data containing at least characteristic and capability information, such as ratings, labels, and brief descriptions for each capability dimension. Internally, the server performs data computations such as vectorization and forward inference: it segments the text into tokens and maps them to embedding vectors, feeds them into a multi-layer attention network, generates result text or a sequence of tags with field names and values according to the output format specified in the prompts, and then parses it into structured data.
[0468] Step 6: The server performs emotion recognition on moving image and audio information. The server inputs the animated image information from step 3 into the image path of the emotion recognition engine, performing face detection and expression classification for each frame; simultaneously, it inputs audio information into the audio path of the emotion recognition engine, extracting acoustic features and classifying emotions. The input consists of video and audio frames segmented by time; the output is a sequence of emotion indicators with timestamps, where each time slice contains multi-dimensional emotion probability or intensity values. The server uses data calculations such as convolution operations, feature pooling, and classifier computation to obtain the timeline curves showing the changes in emotions such as "happiness," "tension," and "anger."
[0469] Step 7: The server performs time alignment and fusion of structured semantic data and sentiment estimation data. The server uses timestamps to align the structured data obtained in step 5 with the sentiment index sequence obtained in step 6. For each question's corresponding time interval, it calculates the average, peak, and volatility of the sentiment index within that interval. The input consists of structured characteristic / ability data and sentiment time series data; the output is a multimodal feature set aggregated by question, including ability scores, characteristic labels, and corresponding sentiment statistical features. In this step, the server performs data calculations such as window statistics, weighted averages, and standard deviation calculations to construct a feature vector for comprehensive evaluation of the individual.
[0470] Step 8: The server calculates a person's characteristics, abilities, and risk level. Based on the multimodal feature set generated in step 7, the server uses pre-defined rules and / or a trained risk assessment model to calculate the evaluated individual's profile, comprehensive scores for each ability dimension, emotional tendency, and risk level. The input is a multidimensional feature vector (including semantic scores and emotional statistics); the output is an evaluation result data structure containing a description of individual characteristics, a summary of ability tendencies, a summary of emotional tendencies, and a normalized risk score. The server performs matrix operations, weighted summation, analysis of variance, and model inference to generate numerical results reflecting the individual's overall state.
[0471] Step 9: Server-generated evaluation report content Based on the evaluation results obtained in step 8, and combined with a pre-set report template, the server constructs an evaluation report text suitable for both human reading and machine processing. The input consists of the evaluation result data structure and the report template; the output is a hierarchical report text and chart data. Internally, the server generates radar chart data for ability scores, line graph data for emotion changes, and summarizes the individual's characteristics using natural language, forming structured content such as "Summary Section," "Detailed Analysis Section," and "Risk Warning Section," organized together in text and chart formats.
[0472] Step 10: The server returns reports and immediate warning messages to the terminal. The server packages the report text and chart data generated in step 9 into a document or page structure and checks whether the risk level in step 8 exceeds a preset threshold. If it does, a brief warning message is generated and returned. The inputs are the report content and risk score; the output is a response message sent to the terminal, which includes the complete report and optional warning flags. The server reduces data size through compression and encoding operations and sends it over the network to improve transmission efficiency.
[0473] Step 11: The terminal displays the evaluation results and prompts the user. The terminal receives reports and warnings from the server, parses the document structure, and presents the capability radar chart, sentiment curve, and text analysis results on the screen. The input is the report document and warning markers returned by the server; the output is a graphical interface display and possible sound / icon alerts. Based on the warning markers, the terminal highlights risk information on the interface, such as using a red icon or a pop-up prompt, to alert the user to potential risks. In this step, the terminal performs specific actions such as document parsing, graph drawing, and interface refreshing, enabling the user to intuitively understand the server's calculation results.
[0474] Step 12: The user initiated a request for additional analysis based on the report results. After reading the report on the terminal, if the user finds insufficient information or has questions about a certain dimension (such as "leadership" or "stress tolerance"), they can select the corresponding dimension through the "In-Depth Analysis" button on the terminal and enter supplementary questions or instructions. The input consists of the dimension identifier selected by the user and optional text descriptions; the output is an append analysis request message sent to the server. The terminal encodes the user's intent into a structured request, making it easier for the server to automatically select an appropriate prompt template for subsequent processing.
[0475] Step 13: The server generates additional questions and new prompts, and performs supplementary analysis. After receiving the request from step 12, the server automatically determines the specific aspects that need to be supplemented based on the confidence level, variance, and other information for that dimension in the report. It then selects or synthesizes several additional questions from the question template library and generates prompts for analyzing the additional answers. The input is the user's additional analysis request and the current evaluation result; the output is a list of additional questions and corresponding prompts. The server then presents the additional questions to the user through the terminal. After receiving new answers, it repeats the processing flow from steps 4 to 8 to perform a more detailed analysis of that dimension, update the evaluation results, and generate a report fragment again, thus achieving an automatic closed-loop supplementary evaluation process.
[0476] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0477] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0478] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0479] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0480] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0481] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0482] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0483] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0484] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0485] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0486] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0487] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0488] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0489] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0490] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0491] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0492] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0493] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0494] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0495] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0496] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0497] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0498] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0499] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0500] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0501] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0502] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0503] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0504] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0505] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0506] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0507] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0508] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0509] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0510] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0511] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0512] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0513] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0514] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0515] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0516] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0517] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0518] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0519] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0520] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0521] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0522] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0523] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0524] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0525] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0526] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0527] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0528] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0529] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0530] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0531] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0532] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0533] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0534] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0535] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0536] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0537] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0538] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0539] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0540] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0541] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0542] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0543] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0544] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0545] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0546] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0547] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0548] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0549] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0550] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0551] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0552] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0553] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0554] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0555] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0556] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0557] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that executes specific processes by executing software, i.e., a program. Furthermore, processors can include, for example, FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are dedicated circuits with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0558] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0559] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0560] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0561] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0562] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0563] In addition, the following notes are provided in response to the above explanation.
[0564] Example 1 (Note 1) An information processing system, characterized in that it comprises: A data storage unit used to pre-store multiple topic information and to associate user authentication information, session history information and evaluation result information; An authentication control unit is used to receive authentication information sent by a user through a terminal device, perform authentication processing on the user based on the authentication information, and enable the user to be prompted with a list of topic information when authentication is successful. A session generation unit is used to receive identification information of selected topic information from the terminal device, generate session information corresponding to the user based on the identification information, and generate initial prompt statements for input to the generative artificial intelligence model. A dialogue control unit that, based on the initial prompt statement and the conversation history with the user, sends query data to an external generative artificial intelligence model, obtains question text for the conversation from the generative artificial intelligence model, and sends the question text to the terminal device to advance a multi-turn conversation with the user. A response management unit is used to sequentially acquire user response texts sent from the terminal device, accumulate the acquired response texts as the session information, and resend at least a portion of the session information as context information to the generative artificial intelligence model, thereby enabling the generative artificial intelligence model to generate response management unit for the next round of session question texts. This is used to perform text parsing processing, including word extraction processing, sentiment inference processing, and syntactic parsing processing, on the entire conversation information, and to generate text parsing units of structured data from the parsing results. An analysis text acquisition unit is used to generate prompt statements based on the structured data and analysis data containing the conversation information, send the analysis data with the prompt statements attached to it to the generative artificial intelligence model, and obtain analysis text containing character evaluation information from the generative artificial intelligence model. A document generation unit for generating evaluation document data based on the analysis text and the structured data, including at least a portion of user personality traits, thinking patterns, communication styles, and role suitability information, and converting the evaluation document data into an output format for storage or distribution.
[0565] (Note 2) The information processing system according to Appendix 1 is characterized in that, When generating the evaluation document data, the document generation unit extracts additional condition information representing unevaluated or insufficient information from the document review information, automatically generates additional prompt statements containing the additional condition information and sends them to the generative artificial intelligence model, obtains additional analysis text from the generative artificial intelligence model to supplement the unevaluated information, and integrates the additional analysis text into the evaluation document data.
[0566] (Note 3) The information processing system according to Appendix 1 is characterized in that, During the conversation, the dialogue control unit generates prompt statements corresponding to predefined evaluation perspective information for each acquired response text. The prompt statements are combined with the response text and sent to the generative artificial intelligence model. The generative artificial intelligence model sequentially obtains evaluation result information for each response text, associates and accumulates the obtained evaluation result information with the conversation information, and reflects the accumulated result in the evaluation document data.
[0567] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for conducting a dialogue between an information processing device and a generative artificial intelligence model around a pre-defined conversation topic, through an evaluated object, and for acquiring and parsing the content of the dialogue as chronological text information. An apparatus for generating prompt statements based on the text information and inputting them into the generative artificial intelligence model, and sending the prompt statements and the text information to the generative artificial intelligence model, thereby obtaining analysis results on the personality traits and hobbies of the evaluated object; An apparatus for performing emotion recognition processing on the speech information of the evaluated object contained in the text information and the acoustic feature information extracted from the speech signal corresponding to the speech information, and on the image information obtained by the expression detection device, so as to estimate the emotional state of the evaluated object. An apparatus for generating a report containing assessment information on the personality traits, hobbies, and emotional aspects of the evaluated subject, based on the analysis results and the inferences about the emotional state. A device for inferring the needs and preferences of the evaluated object in real time based on the analysis results obtained from the generative artificial intelligence model during the dialogue, and for displaying the corresponding information on the display device according to the inference results, so as to output auxiliary information to support the dialogue participants in responding to the activity. A device for storing the report information and the auxiliary information in a recordable and reusable form for use in supporting activities.
[0568] (Note 2) According to the information processing system described in Appendix 1, the information processing device is configured to: in order to deepen the analysis of the characteristics and preferences of the evaluated object, automatically generate a prompt statement containing additional perspectives and confirmation items based on the dialogue content and the report information, send the prompt statement to the generative artificial intelligence model to perform additional analysis, and update the report information accordingly.
[0569] (Note 3) According to the information processing system described in Appendix 1, the information processing device is configured to: include evaluation indicators and evaluation benchmarks for evaluating the characteristics and preferences of the evaluated object in the prompt statement in the analysis results obtained from the generative artificial intelligence model; and, according to the prompt statement, cause the generative artificial intelligence model to evaluate the response content of the evaluated object, and reflect the obtained evaluation results in the report information and the auxiliary information.
[0570] Example 2 (Note 1) An information processing system, characterized in that it comprises: A means configured to acquire multiple response information and register the response information into storage information; The means are configured to extract text information from the answer information registered in the stored information, perform preprocessing on the text information including deleting useless symbols, whitespace normalization and segmentation into basic string units, and convert the text information into formatted text information that can be input into a generative artificial intelligence model. A means configured to combine the formatted text information with prompts that specify evaluation metrics and output formats to form input data for the generative artificial intelligence model, and to obtain analysis results about the response information by running the generative artificial intelligence model on remote or local computing resources; It is configured to store the analysis results as structured data, and to perform the calculation of capability indicators, the comprehensive evaluation between indicators, and the generation of numerical data for visualization based on the structured data. A means configured to generate report information indicating the characteristics and potential abilities of test takers based on the numerical data used for visualization and the analysis results, and to organize it into display data that can be output to an information display device.
[0571] (Note 2) The information processing system according to Appendix 1 is characterized in that, The means include being configured to acquire the analysis results and report information of multiple candidates registered in the stored information, generate selection results for a candidate group by performing ability index comparisons, statistical calculations and screening based on selection criteria among the multiple candidates, and outputting the selection results in association with the displayed data.
[0572] (Note 3) The information processing system according to Appendix 1 is characterized in that, This includes means of updating the evaluation items, output format, and weights of the prompt statement based on the analysis results and the selection results or evaluation information obtained from the user, thereby dynamically changing the composition of the input data for the generative artificial intelligence model to improve the accuracy of the subsequent generation of the analysis results and the report information.
[0573] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: An apparatus for obtaining input language information and motion image information from the evaluated object with the cooperation of a terminal for a predetermined topic, and providing the language information to a generative artificial intelligence model to perform dialogue processing according to the dialogue process. An apparatus for generating prompt statements based on the language information for inputting the answer of the evaluated object into the generative artificial intelligence model, and sending the prompt statements and the language information together to the generative artificial intelligence model, thereby obtaining structured data containing at least characteristic information and ability information; An analytical device for extracting facial expression information and audio features based on the animated image information and audio information, and using an emotion recognition engine to generate emotion estimation data representing multiple emotion indicators and their temporal changes. An apparatus for mapping the structured data to the emotion estimation data based on time information, and integrating the characteristic information, ability information and emotion indicators corresponding to multiple questions, thereby calculating an evaluation result that includes at least the personality characteristics, ability tendency, emotion tendency and risk level of the evaluated object. An apparatus for generating an evaluation report based on the evaluation results, which includes at least ability indicators, emotion indicators, personality traits and risk indicators, and providing the evaluation report to the terminal in a predetermined document format; An apparatus for continuously acquiring new input information or behavioral record information related to the evaluated object from the terminal, repeatedly performing parsing processing in real time using the generative artificial intelligence model and the emotion recognition engine to update the risk level, and generating a warning message and sending it to the terminal when the risk level exceeds a predetermined threshold.
[0574] (Note 2) The information processing system according to Appendix 1 is characterized in that, The system further includes: a device for determining evaluation items that are not fully understood based on written information and the evaluation report, automatically generating additional questions for further questioning of the evaluation items and prompt statements for analyzing the answers to the additional questions, sending the prompt statements to the generative artificial intelligence model to obtain supplementary analysis results related to the evaluation items, and adding the obtained supplementary analysis results to the evaluation report.
[0575] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system further includes: a device for performing parsing processing on each dialogue unit using the generative artificial intelligence model and the emotion recognition engine on the language information, motion image information and audio information obtained through multiple dialogues, generating statistical information representing the time changes of the obtained characteristic information, ability information and emotion indicators, evaluating the stability of the characteristics and the consistency of the emotional response of the evaluated object based on the statistical information, and reflecting the evaluation result in the evaluation report.
Claims
1. An information processing system, characterized in that, include: processor; Emotion recognition engine; Generative artificial intelligence models; The processor is configured as follows: The candidate engages in a dialogue with the generative artificial intelligence model on a pre-defined theme, and the content of the dialogue is analyzed. The emotion recognition engine is used to analyze the candidate's speech, facial expressions, and tone of voice to estimate the candidate's emotions. Based on the dialogue parsing results and the emotion recognition results, a report is generated to evaluate the candidate's characteristics and emotional aspects.
2. The information processing system according to claim 1, characterized in that, The processor is also configured to generate and send prompt text requesting the generative artificial intelligence model to perform additional analysis in order to further confirm any parts that were not clearly identified in the written screening.
3. The information processing system according to claim 1, characterized in that, The processor is also configured to evaluate the candidate's responses based on prompt text in order for the generative artificial intelligence model to analyze the candidate's characteristics and reflect the analysis results in the report.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A