Information processing system, information processing apparatus, server device, program, or method
An information processing system analyzes audio and video during interviews to enhance communication support within organizations by offering real-time feedback on listening and behavior, addressing the inadequacies of existing technologies in labor relations support.
Patent Information
- Application Number
- JP2024021471
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-15
- Publication Date
- 2025-08-27
Smart Images

Figure 2025125423000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology disclosed in this application relates to an information processing system, an information processing device, a server device, a cloud, a program, or a method. [Background technology]
[0002] Currently, communication via networks is increasing due to the increase in telecommuting, and communication support technologies using information processing devices have been proposed. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] ICSoft "Telework Attendance Management" (https: / / www.ttc.cloud / telework) [Non-patent document 2] Cloco "Cloud Call Center System" (http: / / www.clocoinc.com / inbound) [Non-patent document 3] Sonic Garden's "Virtual Office Remotty" (https: / / ja.remotty.net / ?referrer=https%3A%2F%2Fwww.remotty.net%2F) [Non-patent document 4] Iguazu "Sococo" (http: / / www.iguazu-sococo.jp / ) [Non-Patent Document 5] Canon IT Solutions "Telework Supporter" (https: / / www.canon-its.co.jp / products / telework / ?1) Summary of the Invention [Problem to be solved by the invention]
[0004] In the prior art, there have been technologies that utilize information processing devices in communication between people. However, it cannot be said that the technology using information processing devices provides sufficient support for communication within an organization, particularly in labor relations. Therefore, various embodiments of the present invention provide an information processing system, an information processing device, a server device, a cloud, a program, or a method to solve the above-mentioned problems. [Means for solving the problem]
[0005] An example computer program of the present application is The system, A means for capturing audio and video during the interview between the first and second parties; a first extraction means for extracting a predetermined expression from the speech; a second extraction means for extracting a predetermined behavior from the video; a calculation means for calculating statistical data relating to the predetermined expression and the predetermined behavior; A computer program that functions as a
[0006] An example system of the present application comprises: an acquisition unit for acquiring audio and video during the interview between the first and second parties; a first extraction unit that extracts a predetermined expression from the speech; a second extraction unit that extracts a predetermined behavior from the video; a calculation unit that calculates statistical data related to the predetermined expression and the predetermined behavior; A system comprising:
[0007] An example method of the present application comprises: The system, capturing audio and video of the interview between the first and second parties; extracting a predetermined expression from the audio and a predetermined action from the video; calculating statistical data relating to the predetermined expressions and the predetermined behaviors; How to do it. [Effects of the Invention]
[0008] According to one embodiment of the present invention, communication can be supported more effectively. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram showing the overall configuration of a system according to an embodiment. [Figure 2] FIG. 2 is a block diagram illustrating the functions of a system according to an embodiment. [Figure 3] FIG. 3 is an example of a flow related to one system according to one embodiment. [Figure 4] FIG. 4 is an example of a screen whose display is controlled by one system according to one embodiment. [Figure 5A] FIG. 5A is an example of a flow of a system according to an embodiment. [Figure 5B] FIG. 5B is an example of data recorded by a system according to an embodiment. [Figure 5C] FIG. 5C is an example of data recorded by a system according to an embodiment. [Figure 6] FIG. 6 is an example of a screen whose display is controlled by one system according to one embodiment. [Figure 7] FIG. 7 is an example of a screen whose display is controlled by one system according to one embodiment. [Figure 8] 8 shows an example of a screen whose display is controlled by one system according to one embodiment. [Figure 9] FIG. 9 is an example of a screen whose display is controlled by one system according to one embodiment. [Figure 10] FIG. 10 is an example of a configuration of a system according to an embodiment.
[0010] 1. Summary of the invention A system according to an embodiment of the present invention relates to a technology for supporting communication between related parties. An example of the system according to the present invention will be described with reference to FIG.
[0011] In this application, the term "system" may be composed of one or more information processing devices. For example, a system of one embodiment may be only a server, or only a cloud, as described below, or may have either of these plus one or more terminal devices. Both the server and the cloud may be composed of one or more information processing devices. In this diagram, an example system may be a server and / or a cloud 003.
[0012] The terminal device may include one or more user terminals, one or more administrator terminals, and / or one or more service terminals. Note that the user terminal, administrator terminal, and service terminal may also be referred to as terminal or terminal device in this application document.
[0013] A user terminal may be a terminal used by a user. A user may be a person within an organization. For example, a user may include a higher-ranking person within an organization (the term may vary depending on the organization, such as a superior or a manager, but may be referred to as a "mentor" in this document) and a lower-ranking person within an organization (the term may vary depending on the organization, such as a subordinate or a member, but may be referred to as a "mentee" in this document). A user terminal used by a mentor may be referred to as a "mentor terminal." A user terminal used by a mentee may be referred to as a "mentee terminal." In this figure, user terminals 01A, 01B, 02A, and 02B may be used. The user terminals may be connected to a server and / or a cloud via a network. Note that user terminals used by users within one organization may be 01A and 01B. Alternatively, the mentor terminal may be 01A, and the mentee terminal may be 01B. Here, two user terminals are listed for one organization, but there may be more than two. In particular, there may be one or more mentor terminals, and one or more mentee terminals. Also, the user terminals used by users in the other organization may be 02A and 02B. The mentor terminal in such other organization may be 02A, and the mentee terminal may be 02B. Here, two user terminals in the other organization are given, but there may be two or more. In particular, there may be one or more mentor terminals, and one or more mentee terminals.
[0014] The administrator terminal may be a terminal used by an administrator in charge of managing users in the organization to which the users belong. For example, a person in a department responsible for managing users may be a person in a department called a human resources department or a labor department. The human resources department or labor department may need to understand the relationship between mentors and mentees early and objectively in order to manage the labor relations of personnel within the organization. For example, they may need to understand the relationship between mentors and mentees or understand the health (e.g., mental health) of mentors and / or mentees based on objective data. The exemplary system may provide such objective data in some cases. In this diagram, there may be administrator terminals 01C and 02C. The administrator terminals may be connected to a server and / or a cloud via a network. While 01C is described as an administrator terminal within the one organization, there may be two or more administrator terminals within the one organization. Although 01C is described as an administrator terminal within the other organization, there may be two or more administrator terminals within the other organization.
[0015] The service terminal may be a terminal used by a provider of the exemplary system disclosed herein. For example, the service terminal may be used to update data provided by the exemplary system, as described below, or to support data selection. In this figure, there may be service terminals 003A and 003B. The service terminal may be connected to a server and / or a cloud via a network. There may be one or more service terminals.
[0016] 2. Functions of the System According to an Embodiment The system in one example includes an acquisition unit that acquires data, a processing unit that processes the data, and a display control unit that controls the display of the data.
[0017] 2.1. Acquisition part The acquisition unit has a function of acquiring data. The acquisition unit may acquire data related to the mentor and / or data related to the mentee. For example, the acquisition unit may acquire data entered by the mentor and / or the mentee. Additionally or alternatively, the acquisition unit may acquire audio and / or video data related to the mentor and / or the mentee.
[0018] One or more acquisition units may exist in various locations within the system. For example, the acquisition unit may be located within a server or cloud, or within a user terminal.
[0019] The acquisition unit may be realized by an input device, a communication device, and / or a device that accesses data from a storage device.
[0020] 2.2. Processing Unit The processing unit has a function of processing the data acquired by the acquisition unit. For example, the processing unit may perform the analysis process described below.
[0021] The processing unit may be realized by a computing device described below. The processing unit may also be executed in an information processing device. For example, the processing unit may be executed in a server, a cloud, and / or a terminal device.
[0022] 2.3.Display control section The display control unit has a function of controlling the display of data. The displayed data may be data acquired by the acquisition unit, data processed by the processing unit, and / or data that the example system has in advance.
[0023] The display control unit may be realized by a computing device and / or a display device, which will be described later. The display control unit may also be executed within an information processing device. For example, the display control unit may be executed within a server, a cloud, and / or a terminal device.
[0024] 3. Working Example 3.1. Embodiment 1 An example system of the first embodiment will be described below with reference to FIG. 3. Note that prior to the advance preparation for the following interview, advance preparation for using the example system may be performed. The advance preparation may include, for example, setting an ID, account, password, and basic information. The basic information may include personnel information such as a person's name, email address, telephone number, job title, and place of work.
[0025] Advance preparation Step 1 One example system supports the setting of interview dates and times. For example, the system may accept input of interview dates and times along with calendar dates. Only candidate interview dates and times that are available for both the mentor and the mentee may be accepted for input. The mentor and mentee may be set in advance, or, if there are multiple candidates for the interview partner, the interview partner may be set first. The interview date and time may be set from either the mentee's terminal or the mentor's terminal.
[0026] In response to input of the date and time of the interview, the system in one example may set the corresponding date and time in the schedules of the mentor and mentee involved in the interview as the date and time of the interview.
[0027] Step 2 The system may assist in determining the topic of the interview. For example, the system may accept input of topics that the mentor would like to discuss, such as future goals, work, personal matters, and / or requests to their manager. The topics that the mentor would like to discuss may be acquired from the mentee's terminal or the mentor's terminal.
[0028] In one example, the system may notify the other parties of the interview of the content they wish to discuss, obtained from one of the parties. This has the advantage that the parties of the interview can share and view the content they wish to discuss with each other. While interviews generally tend to be led by the mentor, if the mentee inputs the topics they wish to discuss in the interview in advance and shares them with the mentee, this has the advantage that the mentor can at least proceed with the interview knowing the topics the mentee wants to discuss.
[0029] In addition, the example system may accept the input of personal notes in addition to the input and viewing of the content to be discussed. Such personal notes may be displayed on the terminal of the user who entered the personal note during the interview. In this case, there is an advantage that the user can conduct the interview while checking the personal note. Note that the personal note may be disclosed only to the person who entered it. In other words, the example system may display the personal note during the interview only to the user terminal that accepted the input of the personal note. With this configuration, there is an advantage that the user can easily enter personal notes without forgetting to include matters that the user does not wish to disclose to the other party in the interview.
[0030] Step 3 The exemplary system may generate and display advice regarding an interview. The advice regarding an interview may be generated based on the topic of the user's interview obtained as described above and / or data of the user's past interviews. For example, the exemplary system may include a machine-learned model, and the machine-learned model may generate advice for the user. The machine-learned model may use known techniques such as LLM (Large Scale Language Model) or ChatGPT.
[0031] For example, FIG. 4 shows an example of advice displayed by the system of the example.
[0032] The above are the steps before the interview, but the system of the example may perform some or all of the operations of the above processes, and in addition, the system of the example may perform various other processes.
[0033] Interview process Next, the flow of an actual interview will be explained using Figure 5A. The users conducting the interview (the mentor and mentee) may log in to the system at the start date and time of the interview and be able to use the system.
[0034] Step 1 An example system may accept input of the mentee's "energy level" and / or "busyness" before the start of the interview. Such "energy level" and / or "busyness" may be communicated and displayed to the mentor, the other party in the interview. By communicating and displaying the "energy level" and / or "busyness" input by the mentee to the mentor, there is an advantage in being able to share a common current situation. In this case, there is an advantage in that it may also lead to the relief of the mentee's psychological burden. For example, these data may be able to provide an opening for the conversation during the interview.
[0035] Step 2 An example system may control the initiation of a video conference. Here, the video conference function may be a function built into the example system, or may be a function of software outside the example system, and the initiation may be controlled by cooperation with such software. An example of the latter is a browser-based video conferencing system. Note that the example system may obtain audio and / or video from the video conferencing software during the video conference.
[0036] Step 3 The system may analyze the audio and / or video of the user during the video conference while capturing the audio and / or video, and may display the results of the analysis on the user terminal.
[0037] An example system may perform the following analytical processes based on audio and / or video: listening analysis, approval analysis, inappropriate word detection, negative attitude detection, advice generation, conversation ratio and / or timeline generation, topic and / or summary generation of interview content, and / or mirroring detection.
[0038] Active listening analysis For example, the exemplary system may extract listening words based on speech. The listening words may refer to words that express acceptance of what the other person in the interview has said. For example, the listening words may include expressions such as "That's right," "That's right," and "That's right." The listening words may be expressions that are stored in advance in the exemplary system.
[0039] The exemplary system may analyze the extracted listening words and perform statistical processing. For example, the exemplary system may calculate the total number of listening words extracted from the audio during a single interview. For example, the system may calculate the total number of listening words such as "that's right," "certainly," and "that's right." In this case, it is advantageous to collect data on how many times during the interview the audio indicates that the other party's utterance has been received. Additionally or alternatively, the exemplary system may calculate the total number of listening words extracted from the audio during a single interview for each type of listening word. For example, the system may calculate the total number of times the word "that's right" is uttered during a single interview. In this case, it is advantageous to collect data on how many times during the interview a specific expression indicating that the other party's utterance has been received. Furthermore, the exemplary system may calculate the total number of listening words extracted from the audio within a predetermined period of the interview. The predetermined period may be a period covering the same topic and / or summary. For example, an example system may calculate the total number of listening words extracted from audio during the same topic and / or summary.
[0040] Additionally or alternatively, for example, the exemplary system may extract a user's nod based on the video as an analysis process. For example, the presence or absence and / or degree of the user's nod may be determined based on the user's facial movement in the video. Such detection of the user's facial movement may be performed using known techniques.
[0041] The exemplary system may analyze the number of extracted nods and perform statistical processing. For example, the exemplary system may calculate the total number of nods extracted from video during a single interview. This has the advantage of collecting data on how many nods indicate acceptance of the other person's remarks during the interview. The exemplary system may also calculate the total number of nods extracted from video within a predetermined period of the interview. The predetermined period may be a period covering the same topic and / or summary. For example, the exemplary system may calculate the total number of nods extracted from video during the same topic and / or summary.
[0042] Additionally or alternatively, for example, the exemplary system may extract a smile of the user based on the video as an analysis process. For example, whether or not the user is smiling may be determined based on the movement of the user's face in the video. Such detection of whether or not the user is smiling may be performed using known techniques.
[0043] The exemplary system may analyze the number of extracted smiles and perform statistical processing. For example, the exemplary system may calculate the total number of smiles extracted from video during a single interview. This has the advantage of collecting data on how many times during an interview a smile indicates that the other person has acknowledged what they have said. The exemplary system may also calculate the total number of smiles extracted from video within a predetermined period of the interview. The predetermined period may be a period covering the same topic and / or summary. For example, the exemplary system may calculate the total number of smiles extracted from video during the same topic and / or summary.
[0044] The system of the example may determine whether or not the user is listening based on both the user's voice and attitude during the interview, based on whether the time period when listening words are extracted from the voice and the time period when nodding and / or smiling are extracted from the video are within a predetermined range of time difference.The system of the example may then calculate the total number of times listening is determined based on whether the extraction of listening words based on the voice and the extraction of nodding and / or smiling based on the video occur within a predetermined range of time difference.In this case, there is an advantage that data on more attentive listening attitudes can be collected based on both the user's voice and attitude.
[0045] The above-mentioned listening words, total number of listening words, total number of listening words within a specified period, number of nods, number of nods within a specified period, number of smiles, number of smiles within a specified period, number of listening sessions, and / or number of listening sessions within a specified period are sometimes referred to as "listening statistical data."
[0046] Such listening statistical data may be displayed on the user terminal in real time during the interview. In this case, there is an advantage that the parties in the interview can recognize the listening statistical data in real time and consider how to proceed with the interview. In this case, if the average number and / or ideal number of times within a group are displayed in real time as listening statistical data, these may each be displayed in real time according to the passage of time. Here, the average value within the group may be calculated based on data collected by an example system in the past.
[0047] ·Approval analysis Additionally or alternatively, for example, the exemplary system may extract positive words and / or words that develop the individual from the speech. Here, positive words may be words that affirm what the other person is saying. Examples include words such as "That's great," "Thank you," "Chance," and "Excellent." Examples include words such as "Take a risk," "Leave it to us," "Just do it," and "Challenge." Positive words and / or words that develop the individual may be expressions that are stored in advance in the exemplary system.
[0048] The exemplary system may analyze positive words and / or words that enhance the individual and perform statistical processing. For example, the exemplary system may calculate the total number of positive words and / or words that enhance the individual extracted from the speech during a single interview. For positive words, the total number of all positive words such as "That's nice," "Thank you," "Chance," and "Excellent" may be calculated. In this case, it is advantageous to collect data on how many times positive expressions were used during the interview. For words that enhance the individual, the total number of all words that enhance the individual such as "Take a risk," "Leave it to us," "Just do it," and "Challenge" may be calculated. In this case, it is advantageous to collect data on how many times expressions that enhance the individual were used during the interview. Additionally or alternatively, the exemplary system may calculate the total number of listening words extracted from the speech during a single interview for each type of positive word and / or word that enhances the individual. For example, it may calculate the total number of times the word "That's nice" was uttered during a single interview. In this case, it is advantageous to be able to collect data on how many times a particular expression is uttered during the interview. The exemplary system may also calculate the total number of positive words and / or enhancing words extracted from the speech within a predetermined period of the interview. The predetermined period may be a period covering the same topic and / or summary. For example, the exemplary system may calculate the total number of positive words and / or enhancing words extracted from the speech within the same topic and / or summary.
[0049] In addition, the above-mentioned positive words, the total number of positive words, the total number of positive words within a specified period, words that increase individuality, the total number of words that increase individuality, and / or the total number of words that increase individuality within a specified period are sometimes referred to as "approval statistical data."
[0050] Such approval statistical data may be displayed on the user terminal in real time during the interview. In this case, the parties to the interview can recognize the approval statistical data in real time and consider how to proceed with the interview. In this case, if the average number and / or ideal number of times within the group are displayed in real time as approval statistical data, these may each be displayed in real time according to the passage of time. Here, the average value within the group may be calculated based on data collected by an example system in the past.
[0051] ·Detection of inappropriate words Additionally or alternatively, the exemplary system may extract inappropriate words from the speech. The inappropriate words may be words that negate the other person's statements. For example, "lack of skills," "lack of ability," or "low goals" may be included. The inappropriate words may be expressions pre-stored in the exemplary system. Note that the collection of data on inappropriate words may function as data for inferring psychological safety, such as power harassment, overbearing attitudes, and negativity.
[0052] The exemplary system may analyze inappropriate words. As described above, the system may calculate the total number of inappropriate words uttered during an interview and / or during a predetermined time period of the interview for the mentor and / or mentee. Additionally or alternatively, the exemplary system may calculate the total number of listening words extracted from the audio during a single interview for each type of inappropriate word. For example, the system may calculate the total number of times the phrase "lacks skills" was uttered during a single interview. This advantageously allows data to be collected on how many times a particular expression was uttered during the interview. Additionally or alternatively, the exemplary system may detect the volume at which an inappropriate word was uttered. Storing the inappropriate word along with the volume advantageously allows data to be collected on the degree of emphasis of the inappropriate word. The exemplary system may also calculate the total number of inappropriate words extracted from the audio during a predetermined period of the interview. The predetermined period may be a period covering the same topic and / or summary. For example, an example system may calculate the total number of inappropriate words extracted from audio within the same topic and / or summary.
[0053] · Detecting negative attitudes Additionally or alternatively, the exemplary system may extract a negative attitude of the user based on the video as an analysis process. The negative attitude may be, for example, the user turning their head to the side, a high proportion of angry facial expressions, a high proportion of sad facial expressions, etc. The detection of the negative behavior based on the video may be determined based on known techniques.
[0054] The exemplary system may analyze the number of times a negative attitude is extracted and perform statistical processing. For example, the exemplary system may calculate the total number of times a negative attitude is determined to be extracted from video during a single interview. This has the advantage of being able to collect data on how many times a negative attitude is displayed in response to the other person's comments during the interview. The exemplary system may also calculate the total number of negative attitudes extracted from video within a predetermined period of the interview. The predetermined period may be a period covering the same topic and / or summary. For example, the exemplary system may calculate the total number of negative attitudes extracted from video during the same topic and / or summary.
[0055] In addition, inappropriate words, the total number of inappropriate words, the total number of inappropriate words within a specified period, the determined negative attitude, and / or the number of negative attitudes may also be referred to as "inappropriate word statistical data."
[0056] Such inappropriate word statistical data may be displayed on the user's terminal in real time during the interview. In this case, there is an advantage that the parties in the interview can recognize the inappropriate word statistical data in real time and consider how to proceed with the interview. In particular, when an inappropriate word from the mentor is detected, an example system may highlight the uttered inappropriate word itself and / or the detected negative attitude itself on the mentor terminal. Furthermore, an example system may highlight the inappropriate word and / or negative attitude itself on the mentor terminal by turning on and off a light, such as flashing the entire screen, in addition to or instead of the inappropriate word and / or negative attitude itself. This has the advantage that the mentor can recognize the utterance of inappropriate words and negative attitude and consider how to proceed with the interview.
[0057] The listening statistical data, the approval statistical data, and / or the inappropriate word statistical data may be referred to as “interview statistical data.” Note that, as explained above, the term “interview statistical data” may be used to refer to, for example, only listening words, or to refer to, for example, only trend statistical data.
[0058] Advice generation An example system may generate advice based on audio and / or video as an analytical process. An example system may generate advice based on interview statistical data as an analytical process. An example system may generate advice based on a machine-learned model. For example, an example system may generate advice based on a large-scale language model (LLM), ChatGPT, or the like. The technology for generating these pieces of advice may be publicly known.
[0059] The generated advice may be displayed in real time during the interview, which has the advantage that the parties in the interview can consider how to proceed with the interview in real time.
[0060] Generate conversation ratios and / or timelines An example system may calculate a conversation ratio during a meeting based on the audio. The conversation ratio may be based on the duration of each party's speech. For example, if the mentor speaks for 3 minutes, the mentee speaks for 6 minutes, and there is a 1 minute gap during a 10-minute meeting, the conversation ratio between the mentor and the mentee may be 3:6.
[0061] The conversation ratio may be displayed in real time during the interview, which has the advantage that the parties in the interview can consider how to proceed with the interview in real time.
[0062] Generate interview topics and / or summaries The system may generate a topic and / or a summary of the interview content based on the audio as an analysis process. The system may have a function for extracting the topic and / or summary during the interview based on the audio, and may extract the topic and / or summary based on the audio at predetermined intervals. Here, the predetermined interval may be a predetermined period of time or may be when the topic changes. The function for extracting the topic and / or summary itself may utilize publicly known technology.
[0063] An example system may acquire and store topics and / or summaries based on audio recorded during an interview. In this case, at the end of the interview, multiple topics and / or summaries may be stored along the interview timeline. FIG. 5B shows an example in which multiple topics are associated along the interview timeline in correspondence with the time elapsed since the start of the interview. That is, each topic within an interview may be associated with the start time to end time of that topic within the interview. In this way, topics and / or summaries within an interview may be associated with time periods within the interview. Furthermore, some or all of such associated lists may be displayed in real time during the interview or after the interview, as described below. Displaying such a list has the advantage of allowing the user to recall and review, along the timeline, the topics and / or summaries discussed in the interview, in particular.
[0064] Additionally, in the exemplary system, multiple topics and / or summaries within an interview may each be associated with interview statistical data. In this case, the total number of listening words, the total number of nods, the total number of smiles, the total number of positive words, the total number of encouraging words, the total number of inappropriate words, and / or the total number of negative attitudes, etc., used by the mentor and mentee while each topic and / or summary was being spoken during the interview may be associated with the mentor and mentee. In this case, the exemplary system advantageously manages interview statistical data associated with the topic and / or summary. For example, an associated list such as that shown in FIG. 5C may be generated. In this case, such data can be organized and managed. Furthermore, some or all of such associated lists may be displayed in real time during the interview or after the interview, as described below. Displaying such lists advantageously allows the user to recall or objectively understand what statements or attitudes were made in particular for which topics and / or summaries.
[0065] The system of the above example may display corresponding interview statistical data (total number of listening words, total number of nods, total number of smiles, total number of positive words, total number of encouraging words, total number of inappropriate words, and / or total number of negative attitudes) associated with the topic and / or summary on the user terminal. In particular, if the system of the example displays such interview statistical data on the mentor terminal in real time, the mentor has the advantage of being able to understand such objective data in real time.
[0066] Mirroring detection An example system may detect similarities in the behavior of a mentor and a mentee. For example, the example system may acquire the mentor's listening words, nods, and / or smiles from the mentor terminal, and simultaneously or within a predetermined time range, acquire the mentee's listening words, nods, and / or smiles from the mentee terminal. In this case, since the mentor and mentee are listening simultaneously or within a predetermined time range, it is advantageous to detect mirroring, a state in which the behavior of both parties during a meeting is similar. Thus, the inventors of the present application have discovered a method for detecting mirroring in an example system and have discovered that it can be used as an indicator of good communication.
[0067] Step 4 After the interview, the system may display interview statistical data as an analysis result on the user terminal. In particular, the system may display the interview statistical data on the mentor terminal.
[0068] In addition, the system may separate interview statistical data into those based on the mentee and those based on the mentor, and calculate the total number for each. The system may also display interview statistical data on a user terminal, separately for those based on the mentee and those based on the mentor. In this case, the interview statistical data separated for the mentee and the mentor may be displayed on the mentee terminal and / or the mentor terminal. When data separated for the mentor and the mentee is displayed on both the mentee terminal and the mentor terminal, the mentor and the mentee can each understand the extent to which they and the mentee have engaged in attentive listening, acknowledgement, and / or inappropriate language. When data separated for the mentor and the mentee is displayed only on the mentor terminal, the mentor can receive and understand the objective data without having to worry about the mentee's comments during the interview, since they are not displayed as objective data.
[0069] The above-mentioned interview statistical data may also be displayed along with the average number within the group and / or the ideal number of times. Displaying the average number within the group has the advantage of allowing users to understand their own position relative to the group by comparing it. The group may be various within an organization. For example, the group may be a group of people of similar rank within the organization, or a group of people in the same region within the organization. For example, if the mentor is a section manager and the mentee is an ordinary employee, the average number within the group in the interview statistical data for the mentor may be the average number of section managers, and the average number within the group in the interview statistical data for the mentee may be the average number of ordinary employees. Furthermore, when the ideal number of times is displayed, users can understand whether the number is close to the ideal number. The ideal number of times may be determined by a third party. For example, the ideal number of times may be determined by the human resources department of the organization to which the mentor belongs. It may also be determined by an external expert. Such ideal number of times may be stored in advance within the exemplary system, and such data may be displayed.
[0070] In addition, the system may display an evaluation of the average number and / or ideal number of meetings in the group on the mentor terminal and / or mentee terminal. For example, if the mentor's and / or mentee's interview statistical data is higher than the average number and / or ideal number of meetings in the group, a positive evaluation may be displayed, and if the data is lower than the average number and / or ideal number of meetings in the group, a negative evaluation may be displayed. This has the advantage of allowing the mentor and / or mentee to understand at a glance how they are doing compared to the average number and / or ideal number of meetings in the group.
[0071] For example, FIG. 6 shows an example of the interview statistical data displayed on the mentor terminal after the interview.
[0072] In this figure, Overview Chart 01 is a chart that allows you to understand the overall trends of the interview at a glance. The overview chart may show evaluations of the conversation rate (conversation ratio), attentive listening, comfortable atmosphere, positive words, and words that develop the individual. The overview chart may also overlay the average within the group, allowing comparison with the average within the group. Here, the conversation rate, attentive listening, positive words, and words that develop the individual may each show the proportion of the mentor who conducted the interview relative to the group (proportion of the total number uttered). The comfortable atmosphere may also be calculated from the deviation between the frequency of nodding and the frequency of smiling.
[0073] The AI attitude advice 02 and conversation advice 05 may display the above-mentioned advice.
[0074] Furthermore, the conversation ratio generated during the interview as described above may be displayed as conversation ratio 03. Such a display has the advantage of allowing the overall situation of the interview to be understood.
[0075] The diagram may also show a timeline of the conversation between the mentor (you in this diagram) and the mentee (the member in this diagram). Here, the parties in the conversation may distinguish between the mentor and the mentee based on whether the audio is acquired from the mentor terminal or the mentee terminal, and display the mentor and mentee separately. In this diagram, the mentor and mentee may be distinguished along the timeline by being displayed with a grid (mentor side) or diagonal lines (mentee side) based on the passage of time from the start of the interview. The mentor and / or mentee's smiling and / or nodding may also be determined based on video footage during the interview. In this diagram, the mentor and mentee may be distinguished along the timeline by the passage of time from the start of the interview, with solid lines in the graph indicating a smiling state and dotted lines indicating a nodding state. Note that the dotted nodding state in the above diagram may be based on audio and / or video. That is, the display may be based on audio only (detection of listening words only), or video only (detection of nodding attitude only), or the determination may be based on audio and video (detection of listening words and nodding attitude). When determining based on audio and video, the value of the dotted line may be determined and drawn based on the detection of listening words, or may be determined and drawn based on the detection of listening words and the nodding attitude. For example, when determining based on audio and video, the value of the dotted line may be determined and drawn based on the value of the detection of listening words and the value of the nodding attitude.
[0076] In addition, in this figure, lines indicating negative states may be displayed in addition to or instead of the solid and / or dotted lines in the graph, distinguishing between mentors and mentees along the timeline as time passes from the start of the interview. The lines indicating negative states may be based on the detection of inappropriate words and / or negative attitudes. Here, the lines indicating negative states may be based on audio and / or video. That is, they may be displayed based only on audio (only the detection of inappropriate words), or only on video (only the detection of negative attitudes), or they may be determined based on audio and video (based on the detection of inappropriate words and negative attitudes). When determining based on audio and video, the values of the lines indicating negative states may be determined and drawn based on the detection of inappropriate words, or based on the detection of negative attitudes, or based on the detection of inappropriate words and negative attitudes. For example, when determining based on audio and video, the values of the lines indicating negative states may be determined and drawn based on the values of the detection of inappropriate words and the negative attitudes.
[0077] FIG. 7 shows an example of the interview statistical data displayed on the mentor terminal after the interview.
[0078] In this diagram, the mentor terminal displays listening statistics data as listening 01, approval statistics data as approval 02, and irregular word statistics data as inappropriate words 03. This diagram is an example of displaying the mentor and mentee separately, along with the average and ideal number of times.
[0079] In addition, this figure also shows an example of displaying detected positive words, words that enhance individuality, and inappropriate words. In particular, positive words, words that enhance individuality, and inappropriate words that have been detected more than a predetermined number of times are highlighted (for example, in bold). In this way, by highlighting, the predetermined number in such cases may be a predetermined threshold value or may be a value stored in the system of this example.
[0080] The system may also be able to display the history of past interviews for each mentor and mentee. For example, Figure 8 shows an example of displaying past trends on a mentor terminal.
[0081] Furthermore, the exemplary system may display rankings as shown in FIG. 9 . That is, the exemplary system may calculate evaluation values within the same group using a predetermined formula and create and display rankings within the agreement group. Here, the predetermined formula may utilize the interview statistical data described above. More specifically, the evaluation value for each interview may be calculated and ranked using a formula such as multiplying a predetermined coefficient by some or all of the above interview statistical data. This has the advantage that by evaluating the interviews using the same formula, mentors can be motivated to conduct better interviews. In particular, the exemplary system may calculate values for rankings using a formula that utilizes the number of detected listening words, the number of nods, the number of smiles, the number of detected positive words, the number of detected words that promote individuality, the number of detected inappropriate words, and / or the number of negative attitudes. For example, the values for rankings may be calculated by multiplying each of these values by a coefficient.
[0082] The exemplary system may or may not store the captured video and / or audio. If the exemplary system stores the captured video and / or audio during the interview, it can be advantageously used to analyze the content of the interview later. On the other hand, if the exemplary system does not store the captured video and / or audio during the interview, it can be advantageously linked to protecting the privacy of the participants in the interview. In particular, if the exemplary system stores the captured video and / or audio during the interview in a portion of memory for analysis during the interview, but does not control the storage of the captured video and / or audio in memory at or after the end of the interview, it can be advantageously linked to protecting the privacy of the participants in the interview. In particular, if the mentee knows before the interview that the system will not store the video and / or audio during the interview, it can be advantageously trusted that such storage control will not be implemented, creating an atmosphere in which the mentee can freely express their opinions during the interview.
[0083] In the various systems described above, when a server and / or cloud is connected to the mentee terminal and the mentor terminal via a network, the video and / or audio acquired by the mentor terminal may be communicated from the mentor terminal to the server and / or cloud, which then acquires the video and / or audio, and the video and / or audio acquired by the mentee terminal may be communicated from the mentee terminal to the server and / or cloud, which then acquires the video and / or audio. The video and / or audio may then be analyzed within the server and / or cloud. This has the advantage that the server and / or cloud only needs to have the ability to analyze the video and / or audio, and the mentor terminal and / or mentee terminal do not need to have the ability to analyze the video and / or audio.
[0084] Alternatively, in the various systems described above, if the server and / or cloud are connected to the mentee and mentor terminals via a network, the video and / or audio captured by the mentor terminal may be analyzed within the mentor terminal, and the analysis results may be transmitted to the server and / or cloud, which then acquires the analysis results. The video and / or audio captured by the mentee terminal may be analyzed within the mentee terminal, and the analysis results may be transmitted to the server and / or cloud, which then acquires the analysis results. In this case, the analysis process within the mentor and / or mentee terminal may include generating some or all of the interview statistical data described above. In particular, when detecting attentive listening words, positive words, words that enhance individuality, inappropriate words, summaries, and / or topics, detecting nodding, smiling, and negative attitudes based on the audio and / or video, not transmitting the audio and / or video to the server and / or cloud has the advantage of keeping the audio and / or video within the mentor and / or mentee terminals, thereby systematically supporting privacy protection.
[0085] In the various systems described above, the server and / or cloud may have web server functionality, and the terminal device may access the web server and execute software on a browser. Software executed on a web browser within a terminal device has the advantage of reducing the burden of software development. In other words, in the early days of information processing devices, the information processing devices used by individuals were uniform and identical, so the burden of developing software to be installed on each individual's information processing device was not significant. However, in modern times, there are a wide variety of information processing devices, such as desktop PCs, laptops, smartphones, and tablets, and each information processing device has a complex diversity, with different operating systems and versions. Furthermore, individuals and companies select information processing devices suited to their intended work, resulting in a wide variety of information processing devices used by business people. Therefore, the burden of developing software for these diverse information processing devices is significant. Therefore, the software executed on the above-mentioned web browser may be software executable within the web browser, and differences between individual information processing devices may be absorbed by the web browser installed on each information processing device, providing software compatible only with the type and version of the web browser. In this case, there is an advantage that the burden of generating the software is reduced compared to creating software that can be installed on each individual information processing device.
[0086] On the other hand, in one example system, the system in the device that accesses the server and / or cloud may be software downloaded to and installed on the terminal device. In this case, software optimized for execution within each information processing device can be created, which has the advantage of increasing execution speed and reducing the utilization rate of the memory and CPU of each information processing device when executing the functions related to the present invention.
[0087] 3.2. Embodiment 2 The second embodiment is a system that includes a function related to the administrator terminal in addition to the system according to the first embodiment.
[0088] If the example system can quickly and easily obtain objective data about mentors and / or mentees, it has the advantage of being able to display labor management risks on a manager's terminal.
[0089] For example, the system of the example may be able to update the attentive listening words, positive words, words that develop individuality, and / or inappropriate words stored in the first embodiment through input from an administrator terminal. In this case, there is an advantage that the words can be changed according to the company culture or social situation.
[0090] In addition, the system may set a threshold for the number of inappropriate words based on input from an administrator terminal. In this case, the system may associate interviews in which inappropriate words were used at a rate equal to or greater than the threshold with the person who made the comment (particularly the mentor) and notify the administrator terminal. This has the advantage of being able to detect labor risks in advance based on the amount of inappropriate words used.
[0091] The exemplary system may also set a threshold for the number of times a mentee smiles during an interview based on input from an administrator terminal. In this case, the exemplary system may associate an interview in which the mentee smiles less than the threshold with the mentee and notify the administrator terminal. This has the advantage of allowing early acquisition of objective data on possible mental health risks based on the number of smiles. Note that the threshold for the number of smiles may be set based on the number of smiles the mentee has shown in past interviews. For example, the threshold may be based on a percentage decrease from the number of smiles shown in past interviews. In this case, the administrator terminal may be notified based on the decrease in the number of smiles compared to past events.
[0092] 3.3. Embodiment 3 The third embodiment is a system including a function related to a service terminal in addition to the system according to the first and / or second embodiment.
[0093] In one example, if the system acquires data from a service terminal, there is an advantage in that the burden on the labor department can be reduced.
[0094] For example, the system in one example may be able to update the attentive listening words, positive words, words that develop individuals, and / or inappropriate words stored in embodiment 1 through input from the service terminal. In this case, the attentive listening words, positive words, words that develop individuals, and inappropriate words in the dictionary (or its draft candidates) can be updated in accordance with the composition (gender, age, race) of the company staff at the service destination, the influence of current events, changes in the times, etc., which has the advantage of reducing the burden on the labor department for updating these words.
[0095] 3.4. Various Implementations The computer program according to the first aspect includes: The system, A means for capturing audio and video during the interview between the first and second parties; a first extraction means for extracting a predetermined expression from the speech; a second extraction means for extracting a predetermined behavior from the video; a calculation means for calculating statistical data relating to the predetermined expression and the predetermined behavior; It is a computer program that functions as a
[0096] The computer program according to the second aspect is the same as the first aspect, wherein "the predetermined expressions include listening words, positive words, words that develop the individual, and / or inappropriate words."
[0097] The computer program according to the third aspect is the one according to the first or second aspect, wherein "the predetermined behavior includes a smile, a grin, a nod, and / or a negative attitude."
[0098] A computer program according to a fourth aspect of the present invention is a computer program according to any one of the first to third aspects of the present invention, which comprises: A calculation means for calculating a ratio of the conversation time between the first party and the second party during the interview based on the audio.
[0099] The computer program according to the fifth aspect is one in any one of the first to fourth aspects above, which "controls the display unit of the system to display the speaking times of each of the first and second parties along a timeline based on the audio."
[0100] The computer program according to the sixth aspect is one of the first to fifth aspects above, in which "the statistical data includes, for each of the first party and the second party, the number of times each of the specified expressions and the specified actions occurs per unit time or per specified time."
[0101] The computer program according to the seventh aspect is one in any one of the first to sixth aspects above that "controls the display unit of the system to display the average number and / or ideal number of times in association with the number of times per unit time or per specified time for the specified expression and the specified behavior."
[0102] The computer program according to the eighth aspect is one of the first to seventh aspects above, which "does not control the system so that the audio from the interview is stored at the end of the interview."
[0103] The computer program according to the ninth aspect is one of the first to eighth aspects above, which "does not control the system so that the video of the interview is stored at the end of the interview."
[0104] The system according to the tenth aspect includes: an acquisition unit for acquiring audio and video during the interview between the first and second parties; a first extraction unit that extracts a predetermined expression from the speech; a second extraction unit that extracts a predetermined behavior from the video; a calculation unit that calculates statistical data related to the predetermined expression and the predetermined behavior; "A system that includes:
[0105] The system according to the eleventh aspect is the system according to the tenth aspect "including a memory."
[0106] The system according to the twelfth aspect is the system according to the tenth or eleventh aspect "including a processor."
[0107] The method according to the thirteenth aspect includes: The system, capturing audio and video of the interview between the first and second parties; extracting a predetermined expression from the audio and a predetermined action from the video; calculating statistical data relating to the predetermined expressions and the predetermined behaviors; "How to do it."
[0108] The method according to the fourteenth aspect is the method according to the thirteenth aspect "including a memory".
[0109] The method according to the fifteenth aspect is the method according to the thirteenth or fourteenth aspect "including a processor."
[0110] 4. Hardware configuration of the system The system according to the present invention may be composed of one or more information processing devices. As shown in Fig. 10, the information processing device 10 according to the present invention may include a bus 15, a calculation device 11, a storage device 12, and a communication device 16. In one embodiment, the information processing device 10 may also include an input device 13 and a display device 14. The information processing device 10 may also be connected directly or indirectly to a network 17.
[0111] The bus 15 may function to transmit information between the processing unit 11, the storage unit 12, the input unit 13, the display unit 14, and the communication unit 16.
[0112] An example of the arithmetic device 11 is a processor. This may be a CPU or an MPU. In one embodiment, the arithmetic device may include a graphics processing unit, a digital signal processor, or the like. In short, the arithmetic device 12 may be any device that can execute program instructions.
[0113] The storage device 12 is a device for recording information. This may be either an external memory or an internal memory, or either a main storage device or an auxiliary storage device. It may also be a magnetic disk (hard disk), an optical disk, a magnetic tape, a semiconductor memory, or the like. It may also have a storage device connected via a network or a storage device on a cloud connected via a network.
[0114] In this block diagram, registers, L1 caches, L2 caches, and the like, which store information in locations physically close to the arithmetic unit, may be included in the arithmetic unit 11, but in computer architecture design, these may be included in the storage device 12 as a device for recording information. In short, it is sufficient that the arithmetic unit 11, storage device 12, and bus 11 are configured to cooperate to execute information processing.
[0115] The storage device 12 may include a part or all of the programs capable of executing the processes according to the present invention. It may also be capable of appropriately storing data required for executing the processes according to the present invention. In one embodiment, the storage device 12 may also include a database.
[0116] Furthermore, although the above describes a case where the arithmetic unit 12 is executed based on a program stored in the memory device 13, as one form in which the above bus 11, arithmetic unit 12 and memory device 13 are combined, part or all of the information processing related to the present invention may be realized by a programmable logic device that can change the hardware circuit itself or a dedicated circuit in which the information processing to be executed is predetermined.
[0117] The input device 13 is used to input information, but may have other functions. Examples of the input device 14 include a keyboard, a mouse, a touch panel, and a pen-type pointing device.
[0118] The display device 14 has a function of displaying information. Examples include a liquid crystal display, a plasma display, and an organic EL display, but in short, any device that can display information will do. In addition, the display device 14 may partially include an input device 13 such as a touch panel.
[0119] The network 17 transmits information together with the communication device 16. That is, the network 17 has a function to transmit information from the information processing device 10 to another information terminal (not shown) via the network 17. The communication device 16 may use any connection method, such as IEEE1394, Ethernet (registered trademark), PCI, SCSI, USB, 2G, 3G, 4G, or 5G. The connection to the network 17 may be either wired or wireless.
[0120] The information processing device according to the present invention may be a general-purpose or dedicated type, and may be a workstation, desktop personal computer, laptop personal computer, notebook computer, PDA, mobile phone, smartphone, etc.
[0121] Although the present diagram illustrates a single information processing device 10, the system according to the present invention may be configured with multiple information processing devices, which may be internally connected or externally connected.
[0122] Furthermore, the system according to the present invention may be in various device formats. For example, the system according to the present invention may be a standalone system, a server-client system, a peer-to-peer system, or a cloud system. The system according to the present invention may be a standalone information processing device, may be composed of some or all information processing devices in a server-client system, may be composed of some or all information processing devices in a peer-to-peer system, or may be composed of some or all information processing devices in a cloud system.
[0123] Furthermore, when the system according to the present invention is configured with a plurality of information processing devices, the owners and managers of the respective information processing devices may be different.
[0124] Furthermore, the information processing device 10 may be a physical entity or a virtual entity. For example, the information processing device 10 may be virtually realized using cloud computing.
[0125] Although the above description has been given as a configuration implemented by the system of this example, these may also be a configuration implemented by one or more information processing devices within the system.
[0126] The inventions described in the examples of the present application are not limited to those described in the present application, and can be applied to various examples within the scope of the technical idea. Furthermore, a program according to an embodiment may itself have instructions for transmitting and / or receiving information, or may simply set information in a state in which another program can transmit and / or receive it.
[0127] Furthermore, the processes and procedures described in this application document may be realized not only by those explicitly described in the embodiments, but also by software, hardware, or a combination thereof. Furthermore, the processes and procedures described in this application document may be implemented as computer programs and executed by various computers. In particular, some or all of the functions realized by the term "system" used in this application document may be realized by a program executable on such an information processing device. Furthermore, these computer programs may be stored in a storage medium. Furthermore, these computer programs may be stored in a non-transitory or temporary storage medium.
Claims
1. The system, capturing means for capturing audio and video during the interview between the first and second parties; a first extraction means for extracting a predetermined expression from the speech; a second extraction means for extracting a predetermined action from the video; a calculation means for calculating statistical data relating to the predetermined expression and the predetermined behavior; A computer program that functions as a
2. The predetermined expressions include listening words, positive words, personal development words, and / or inappropriate words.
2. The computer program of claim 1.
3. The predetermined behavior includes a smile, a grin, a nod, and / or a negative attitude.
2. The computer program of claim 1.
4. The system, a calculation means for calculating a ratio of the conversation time between the first party and the second party during the interview based on the audio; The computer program according to claim 1, for causing the computer to function as
5. and controlling the display unit of the system to display speech time periods of the first party and the second party along a timeline based on the speech.
2. The computer program of claim 1.
6. The statistical data includes the number of times per unit time or per predetermined time of each of the predetermined expressions and the predetermined behaviors for each of the first party and the second party.
2. The computer program of claim 1.
7. controlling a display unit of the system to display an average number and / or an ideal number of times in association with the number of times per unit time or per predetermined time for the predetermined expression and the predetermined behavior; 2. The computer program of claim 1.
8. The system, not controlling that the audio from the interview is stored at the end of the interview; 2. The computer program of claim 1.
9. The system, not controlling that the video of the interview is stored at the end of the interview; 2. The computer program of claim 1.
10. an acquisition unit for acquiring audio and video during an interview between the first and second parties; a first extraction unit that extracts a predetermined expression from the speech; a second extraction unit that extracts a predetermined behavior from the video; a calculation unit that calculates statistical data related to the predetermined expression and the predetermined behavior; A system comprising:
11. The system of claim 10 comprising a memory.
12. The system of claim 10 comprising a processor.
13. The system, capturing audio and video of an interview between a first party and a second party; extracting a predetermined expression from the audio and a predetermined action from the video; calculating statistical data relating to the predetermined expressions and the predetermined behaviors; How to do it.
14. The method of claim 13 , wherein the system comprises a memory.
15. The method of claim 13 , wherein the system comprises a processor.
Citation Information
Cited By
Information processing device, information processing method, and program
JP7922860B1