Information processing program, information processing method and information processing device

The system addresses the variability in trainee psychological responses by tailoring synthetic voice training based on physiological data, ensuring effective training through personalized voice messaging.

JP2025117011APending Publication Date: 2025-08-12FUJITSU LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024011633
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing training systems using synthesized voice do not account for the varying psychological states of trainees, leading to ineffective training as the psychological response to speech elements can differ significantly among individuals.

Method used

An information processing system that estimates a trainee's psychological state based on physiological data and tailors the synthetic voice used in training to match the trainee's characteristics, using a correspondence model to identify appropriate speech elements for effective training.

Benefits of technology

Enables personalized training by generating voice messages that align with the trainee's psychological state, enhancing the effectiveness of training by addressing individual responses to speech elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025117011000001_ABST
    Figure 2025117011000001_ABST
Patent Text Reader

Abstract

To perform training using synthetic speech tailored to characteristics of a trainee.SOLUTION: An information processing program causes a computer to perform processing to estimate a psychological state of a trainee based on physiological data while the trainee is listening to a first voice message having a first element, identify a second element of synthetic voice to be used for the trainee's training based on the estimated psychological state of the trainee, and output a second voice message to be used for the trainee's training based on the identified second element.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing program, an information processing method, and an information processing device. [Background technology]

[0002] There is a technology called speech synthesis, which uses a computer to artificially generate voice messages that read out a scenario. Speech synthesis generates voice messages that are used in a variety of situations, such as reading out the news, answering questions at a call center, and providing facility guidance, and provides the listener with a synthesized voice. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-022447 Summary of the Invention [Problem to be solved by the invention]

[0004] For example, speech synthesis is used to generate speech messages, such as for conversations with trainees undergoing training or for notifications to trainees. Here, speech messages are generated based on various factors, such as the gender, age, vocal range, and speaking speed indicated by the synthesized speech. The psychological state experienced by trainees listening to a speech message varies depending on these factors. For example, some trainees may feel calmed by a speech message generated in a low male voice, while others may feel suspicious. As such, the psychological state experienced by trainees varies depending on the trainee's characteristics, and there have been cases where training appropriate for a trainee was not provided depending on the elements of the synthesized speech.

[0005] In one aspect, the present invention aims to provide an information processing program, an information processing method, and an information processing device that can perform training using synthesized voice that is tailored to the characteristics of the trainee. [Means for solving the problem]

[0006] In one aspect, an information processing program is provided that causes a computer to execute a process of estimating a psychological state of a trainee based on physiological data of the trainee listening to a first audio message having a first element, identifying a second element of a synthetic voice to be used in training the trainee based on the estimated psychological state of the trainee, and outputting a second audio message to be used in training the trainee based on the identified second element.

[0007] In one aspect, there is provided an information processing method in which a computer executes a process similar to the process based on the information processing program.

[0008] In one aspect, an information processing device that executes processing similar to the processing based on the information processing program is provided. [Effects of the Invention]

[0009] Training can be carried out using synthesized voice that is tailored to the trainee's characteristics. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of a configuration of an information processing system according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a model generation phase and a training phase according to an embodiment. [Figure 3] FIG. 3 is a functional block diagram illustrating an example of a functional configuration of an information processing device according to an embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of fraud data according to the embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of collaborator data according to the embodiment. [Figure 6] FIG. 6 is a diagram illustrating an example of helper psychological information according to the embodiment. [Figure 7] FIG. 7 is a diagram illustrating an example of trainee data according to the embodiment. [Figure 8] FIG. 8 is a diagram illustrating an example of trainee psychological information according to the embodiment. [Figure 9] FIG. 9 is a diagram illustrating an example of an element correspondence model between the types of elements "title of fraudster" and "type of fraud." [Figure 10] FIG. 10 is a diagram illustrating an example of an element correspondence model between the element types of "person" and "impression on person." [Figure 11] FIG. 11 is a diagram illustrating an example of an element correspondence model between the element types of "speech rate" and "mental state." [Figure 12] FIG. 12 is a diagram illustrating an example of a correspondence relationship model. [Figure 13] FIG. 13 is a diagram illustrating an example of a processing flow of the model generation phase of the information processing device according to the embodiment. [Figure 14] FIG. 14 is a diagram illustrating an example of a processing flow of the training phase of the information processing device according to the embodiment. [Figure 15] FIG. 15 is a diagram illustrating an example of a hardware configuration of an information processing device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, an information processing program, an information processing method, and an information processing device according to the present invention will be described in detail with reference to the accompanying drawings. Note that these embodiments are merely examples for carrying out the present invention, and are not intended to limit the present invention. [Example]

[0012] [System Description] An information processing system according to an embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram illustrating an example of the configuration of an information processing system according to an embodiment. The information processing system 10 includes a radar 11, an information processing device 12, a terminal 13, a network 15, and a network 16. The information processing system 10 in this embodiment is a system that performs specialized fraud prevention training to prevent trainees from being deceived by specialized fraud through the experience of simulated specialized fraud. For example, elderly people who are easily deceived by fraud have a high cognitive bias (confirmation bias) and may believe that they will not be deceived by specialized fraud. In such cases, the information processing system 10 allows the elderly trainees to experience simulated specialized fraud and actually be deceived, thereby increasing the trainees' awareness of fraud prevention.

[0013] The radar 11 is a device that measures physiological data of a person 14. In this embodiment, the radar 11 is a millimeter-wave radar. The physiological data is physiological data generated by human physiological activity. An example of the physiological data is vital data that indicates time-series changes in heart rate, respiratory rate, etc. The radar 11 measures the physiological data of the person 14 by emitting a transmission wave toward the person 14 and receiving the reflected wave. The special fraud prevention training using the information processing system 10 has a model generation phase in which a correspondence model to be used for training is learned, and a training phase in which training is carried out using the generated correspondence model. In the model generation phase, the radar 11 acquires physiological data of a collaborator who cooperates in generating the correspondence model. In the training phase, the radar 11 acquires physiological data of a trainee who receives training. The model generation phase and the training phase will be described in detail later.

[0014] The information processing device 12 is a device that learns a correspondence model in the model generation phase. The information processing device 12 is a device that executes training based on the generated correspondence model in the training phase. Examples of the information processing device 12 include a personal computer, a server computer, a smartphone, and a tablet terminal. The information processing device 12 converses in real time with a collaborator or trainee using a virtual agent. The virtual agent is an automatic conversation program that generates voice messages using AI (Artificial Intelligence) or machine learning and converses with humans. In this embodiment, the information processing device 12 converses with a collaborator using the virtual agent in the model generation phase, and converses with a trainee using the virtual agent in the training phase.

[0015] Terminal 13 is a device that outputs a voice message of the virtual agent to person 14. Examples of terminal 13 include a personal computer, a smartphone, a mobile phone, a tablet terminal, and a fixed-line telephone. In the model generation phase, terminal 13 outputs a voice message of the virtual agent received from information processing device 12 to the collaborator. In the training phase, terminal 13 outputs a voice message of the virtual agent received from information processing device 12 to the trainee.

[0016] Person 14 is a person who participates in the specialized fraud prevention training. Person 14 is a collaborator who cooperates in generating the correspondence model in the model generation phase. In one example, there are multiple collaborators. Person 14 is a trainee who receives training in the training phase.

[0017] The network 15 may be a wired or wireless communication network such as the Internet or an intranet. For example, the network 15 may be configured by connecting multiple intranets, or connecting the Internet and an intranet via a gateway or other device.

[0018] The network 16 is a communication network such as the Internet, an intranet, or a telephone line, whether wired or wireless.

[0019] Here, an example of the model generation phase and the training phase will be described with reference to FIG. 2. FIG. 2 is a diagram illustrating an example of the model generation phase and the training phase according to an embodiment. In the model generation phase 20, the information processing device 12 acquires information on fraud occurrences and extracts fraud elements from the fraud occurrence information. The fraud elements are elements that constitute fraud, such as the type of fraud, the title of the speaker impersonated by the fraudster, and the content of the utterance. The information processing device 12 acquires the collaborator's questionnaire responses. The information processing device 12 acquires the collaborator's physiological data for a predetermined period of time while listening to the voice message of the virtual agent. The voice message has speech elements related to the utterance, such as gender, age, and speaking rate, expressed by the synthesized voice. The information processing device 12 extracts characteristic elements of the collaborator from the collaborator's questionnaire responses and physiological data. The characteristic elements of the collaborator are elements that constitute the collaborator, such as the collaborator's attributes, such as age and gender, the collaborator's impression of a specific person, the collaborator's psychological characteristics, and the collaborator's psychological state regarding the speech elements, and are elements that represent the collaborator's characteristics. The information processing device 12 generates an element correspondence model that represents the correspondence between the elements based on the extracted fraud elements and characteristic elements. The information processing device 12 learns the correspondence model based on the generated element correspondence model.

[0020] In the training phase 21, the information processing device 12 acquires the trainee's responses to a questionnaire as a preliminary preparation for training. The information processing device 12 acquires the trainee's physiological data for a predetermined period of time while the trainee listens to the virtual agent's voice message. The information processing device 12 extracts the trainee's characteristic elements from the trainee's questionnaire and physiological data. The trainee's characteristic elements are elements that constitute the trainee, such as the trainee's attributes (e.g., age and gender), the trainee's impression of a specific person, the trainee's psychological characteristics, and the trainee's psychological state regarding speech elements, and are elements that represent the trainee's characteristics. The information processing device 12 updates the values of the element correspondence model based on the extracted characteristic elements. The information processing device 12 inputs the updated element correspondence model into the trained correspondence model and identifies fraud elements and speech elements of the voice message used for training. The information processing device 12 generates a scenario to be used for training based on the identified fraud elements, and generates a voice message that reads the scenario using a synthetic voice having the identified speech elements. The information processing device 12 performs special fraud prevention training using the generated voice message.

[0021] In conventional special fraud prevention training, fraudulent voice messages are generated and training is conducted regardless of the trainee's characteristics, such as the trainee's psychological state resulting from speech elements. However, when the speech elements of a voice message differ, the trainee's psychological state resulting from the voice message varies from trainee to trainee. As a result, for example, in the case of special fraud prevention training, the effectiveness of the training may vary, as trainees may be more susceptible to different speech elements. The information processing device 12 identifies speech elements of the synthetic voice used to train the trainee based on the trainee's psychological state when listening to a voice message containing speech elements, and outputs a voice message to be used for training the trainee based on the identified speech elements of the synthetic voice. In this way, the information processing device 12 identifies speech elements of the synthetic voice used for training based on the trainee's characteristic elements and generates a voice message to be used for training. This allows the information processing device 12 to conduct training using synthetic voice tailored to the trainee's characteristics. The information processing device 12 can generate special fraud situations that are likely to be deceived for each trainee, in other words, situations that require particular caution, thereby enabling effective special fraud prevention training.

[0022] [Function Configuration] 3 is a diagram illustrating an example of a functional block diagram showing the functional configuration of an information processing device according to an embodiment. The functional configuration of an information processing device 12 according to an embodiment will be described with reference to FIG. 3. The information processing device 12 has a communication unit 30, a control unit 31, and a storage unit 32.

[0023] The communication unit 30 controls communication with other devices and is realized by a communication interface such as a NIC (Network Interface Card).

[0024] The storage unit 32 stores various data or various programs executed by the control unit 31. The storage unit 32 is realized by, for example, a main storage device such as a random access memory (RAM) or an auxiliary storage device such as a hard disk drive (HDD). The storage unit 32 stores, for example, a fraud data DB 320, a helper data DB 321, a third voice information DB 322, a helper psychological information DB 323, an element correspondence model DB 324, a correspondence model DB 325, a trainee data DB 326, a first voice information DB 327, a trainee psychological information DB 328, a scenario generation model DB 329, and a second voice information DB 330.

[0025] The fraud data DB320 is a database that stores fraud data. For example, the fraud data is information about special frauds that have occurred in the past. For example, the fraud data is data generated from information about the occurrence of special frauds in a target area over a given period, extracted from a security information database, etc. For example, the fraud data is data that stores the fraud elements that constitute special frauds for each occurrence of special fraud.

[0026] Table 40, which is an example of fraud data, will now be described with reference to FIG. 4. FIG. 4 is a diagram illustrating an example of fraud data according to an embodiment. The "occurrence information ID" is an arbitrary identifier for identifying occurrence information of special frauds. For example, table 40 includes, for each occurrence information ID, items representing fraud elements such as "type of fraud," "occurrence date and time," "speaker's title," and "content of speech." The "type of fraud" indicates the type of special fraud, and examples include refund fraud, deposit and savings fraud, "I'm your son" fraud, cash card fraud, gambling fraud, and dating arrangement fraud. The "occurrence date and time" indicates the date and time the special fraud occurred. The "speaker's title" indicates the title of the speaker impersonated by the fraudster in the special fraud, and examples include a city hall employee, a bank employee, a son, and a lawyer. The "content of speech" indicates the content of speech spoken by the fraudster as a means of committing fraud, and examples include asking which bank the person uses, recommending a replacement cash card, and saying that they have had an accident and need money.

[0027] Returning to FIG. 3 , the collaborator data DB321 will be described. The collaborator data DB321 is a database that stores collaborator data, which is information acquired from a questionnaire about collaborators. An example of a collaborator is a victim who has been a victim of special fraud in the past. The questionnaire about collaborators is, for example, a questionnaire that asks about the collaborator's attributes, such as the collaborator's age, gender, family structure, address, and occupation. An example of a questionnaire about collaborators is a questionnaire that asks about the collaborator's impression of a specific person, such as what impression the collaborator has of city hall employees. An example of a questionnaire about collaborators is a questionnaire about the collaborator's psychological characteristics. An example of a questionnaire about psychological characteristics is the BIG5, which analyzes personality based on five elements, or the STAI (State-Trait Anxiety Inventory), which measures anxiety characteristics. An example of a questionnaire about psychological characteristics is a question that asks about one's own awareness of psychological characteristics, such as, "Are you a gullible person?", and the answer is based on a five-value scale, with 1 being the least gullible and 5 being the most gullible. Information such as the attributes of the collaborator obtained from a questionnaire, the collaborator's impression of a specific person, and the collaborator's psychological characteristics are examples of characteristic elements.

[0028] Here, with reference to FIG. 5, a table 50, which is an example of collaborator data, will be described. FIG. 5 is a diagram illustrating an example of collaborator data according to an embodiment. A "collaborator ID" is an arbitrary identifier for identifying a collaborator. For each collaborator ID, table 50 has items representing characteristic elements such as "gender," "age," "family composition," "impression toward city hall employees," and "gullibility." As an example, "gender," "age," and "family composition" are information about the collaborator's attributes obtained from a survey about the collaborator. As an example, "impression toward city hall employees" is information about the collaborator's impression of city hall employees obtained from a survey about the collaborator, and is expressed on three levels, for example, "positive," "negative," and "neutral." For example, a positive rating refers to a state in which the collaborator has a good impression of an object, such as a city hall employee, feeling familiarity or trust. For example, a negative rating refers to a state in which the collaborator has a bad impression of an object, such as feeling discomfort or distrust. For example, a neutral rating refers to a state that is neither positive nor negative. As an example, "gullibility" refers to the collaborator's own gullibility as answered by the collaborator in a survey about the collaborator, and is expressed on a five-value scale.

[0029] Returning to FIG. 3 , the third voice information DB322 will be described. The third voice information DB322 is a database that stores third voice information. The third voice information is information for generating a voice message of a virtual agent that converses with a collaborator during the model generation phase. During the model generation phase, the virtual agent converses with a collaborator while impersonating, for example, a fraudster in a special fraud case, a real estate agent, or a city hall employee. The third voice information, for example, is information about a scenario spoken by the virtual agent and speech elements of the virtual agent. For example, the scenario included in the third voice information may be a scenario related to an event other than special fraud, such as an everyday conversation with a collaborator. For example, the speech elements are speech-related elements such as gender, age, voice pitch, speaking rate, voice volume, clarity of pronunciation, intonation, and tone of voice expressed by the virtual agent's synthesized voice. For example, the information about the speech elements is a setting value for setting the speech elements. The voice message generated based on the third voice information is an example of a third voice message, and the speech element included in the third voice message is an example of a third element.

[0030] The collaborator psychological information DB323 is a database that stores collaborator psychological information obtained from the collaborator's physiological data. The information processing device 12 acquires the collaborator's physiological data measured by the radar 11 during a predetermined period in which the collaborator converses with the voice message of the virtual agent based on the third voice information. The collaborator psychological information is information regarding the collaborator's psychological state in response to the speech elements of the voice message, obtained by analyzing the collaborator's physiological data. The collaborator's psychological state in response to the speech elements is an example of a characteristic element.

[0031] Here, with reference to FIG. 6, a table 60, which is an example of collaborator psychological information, will be described. FIG. 6 is a diagram illustrating an example of collaborator psychological information according to an embodiment. The "voice ID" is an arbitrary identifier for identifying a voice message used in a conversation. For each voice ID, the table 60 includes items indicating the speech elements of the virtual agent with which the collaborator conversed and the collaborator's psychological state during a predetermined period of time during the conversation. The "speech elements" are the speech elements of the voice messages in the conversation associated with the voice ID, and include items such as "age," "gender," "speech rate," and "voice pitch." The "mental state" is the collaborator's psychological state estimated during a predetermined period of time during which the collaborator converses with the virtual agent. The "mental state" is expressed in three levels, for example, "positive," "negative," and "neutral." For example, a positive state indicates a state in which the collaborator feels familiarity and trust with the virtual agent with whom they are conversing, or a state in which the conversation creates a positive impression, such as a state in which the collaborator feels energized and lively through the conversation. Negative, for example, refers to a state in which the user feels a bad impression due to the conversation, such as discomfort or distrust toward the virtual agent, or confusion or embarrassment caused by the conversation. Neutral, for example, refers to a state that is neither positive nor negative. For example, the collaborator psychological information DB 323 stores collaborator psychological information for each collaborator ID.

[0032] Returning to Figure 3, the element correspondence model DB324 will be described. The element correspondence model DB324 is a database that stores element correspondence models. The element correspondence model is a model that represents the correspondence between related elements among elements such as fraud elements, characteristic elements, and speech elements contained in the fraud data, accomplice data, and accomplice psychological information. As an example, the element correspondence model is a model that represents the correspondence between elements using the co-occurrence probability that represents the rate at which each element occurs simultaneously, in other words, the probability of simultaneous occurrence. The element correspondence model will be described in detail later.

[0033] The correspondence model DB325 is a database that stores correspondence models. The correspondence model is a mathematical model that is trained based on the element correspondence model. When an element correspondence model is input, the trained correspondence model outputs a fraud element and an utterance element according to the input element correspondence model. As an example, the correspondence model is a machine learning model. Details of the correspondence model will be described later.

[0034] The trainee data DB326 is a database that stores trainee data, which is information obtained from a questionnaire about trainees. An example of a trainee is a person who receives training in special fraud, such as an elderly person. An example of the questionnaire about trainees is a questionnaire asking about the trainee's attributes, such as the trainee's age, gender, family composition, address, and occupation. An example of the questionnaire about trainees is a questionnaire asking about the trainee's impression of a specific person, such as what impression the trainee has of city hall employees. An example of the questionnaire about trainees is a questionnaire about the trainee's psychological characteristics. An example of the questionnaire about psychological characteristics is a questionnaire asking about the BIG5, STAI, and awareness of one's own psychological characteristics. The questionnaire about trainees may include items similar to those in the questionnaire about collaborators. The trainee's attributes, the trainee's impression of a specific person, and the trainee's psychological characteristics are examples of characteristic elements.

[0035] Here, with reference to FIG. 7, a table 70, which is an example of trainee data, will be described. FIG. 7 is a diagram illustrating an example of trainee data according to an embodiment. Table 70 has items representing characteristic elements such as "gender," "age," "family composition," "impression toward city hall employees," and "gullibility." As an example, "age," "gender," and "family composition" are information about the trainee's attributes obtained from a questionnaire about the trainee. As an example, "impression toward city hall employees" is information about the trainee's impression of city hall employees obtained from a questionnaire about the trainee, and is expressed on three levels, for example, "positive," "negative," and "neutral." As an example, "positive," "negative," and "neutral" represent states similar to "positive," "negative," and "neutral" in Table 50. As an example, "gullibility" is the trainee's own gullibility as answered in the questionnaire about the trainee, and is expressed on five levels.

[0036] Returning to FIG. 3 , the first voice information DB 327 will be described. The first voice information DB 327 is a database that stores first voice information. In the training phase, as a preparation for training, physiological data of the trainee during a predetermined period of conversation with the virtual agent is acquired, and speech elements of the virtual agent to be used in training are identified. The first voice information is voice information for generating a voice message of the virtual agent that will converse with the trainee in the preparation for training. In the preparation for training, the virtual agent converses with the trainee while impersonating, for example, a fraudster in a special fraud case, a real estate agent, or a city hall employee. The first voice information, for example, is information about a scenario spoken by the virtual agent in the preparation for training and about the speech elements of the virtual agent. For example, the scenario included in the first voice information may be a scenario related to an event other than special fraud, such as an everyday conversation with the trainee. For example, the information about the speech elements is a setting value for setting the speech elements. For example, the first voice information may include information about the same scenario and speech elements as the third voice information. A voice message generated based on the first voice information is an example of a first voice message, and a speech element included in the first voice message is an example of a first element.

[0037] The trainee psychological information DB 328 is a database that stores trainee psychological information obtained from the trainee's physiological data. The information processing device 12 acquires the trainee's physiological data measured by the radar 11 during a predetermined period in which the trainee converses with the voice message of the virtual agent based on the first voice information. The trainee psychological information is information related to the trainee's psychological state in response to the speech elements of the voice message, which is obtained by analyzing the acquired trainee's physiological data. The trainee's psychological state in response to the speech elements is an example of a characteristic element.

[0038] Here, with reference to FIG. 8, a table 80, which is an example of trainee psychological information, will be described. FIG. 8 is a diagram illustrating an example of trainee psychological information according to an embodiment. The "voice ID" is an arbitrary identifier for identifying a voice message used in a conversation with a virtual agent. For each voice ID, the table 80 has items indicating the speech elements of the virtual agent with which the trainee conversed and the trainee's psychological state during a predetermined period of time during which the trainee conversed with the virtual agent. The "speech elements" are the speech elements of the voice messages in the conversation associated with the voice ID, and examples of items include "age," "gender," "speech rate," and "voice pitch." The "psychological state" is the trainee's psychological state estimated during a predetermined period of time during which the trainee conversed with the virtual agent. For example, the psychological state is expressed in three stages: "positive," "negative," and "neutral." For example, "positive," "negative," and "neutral" indicate the same states as "positive," "negative," and "neutral" in the table 60.

[0039] Returning to Figure 3, the scenario generation model DB329 will be explained. The scenario generation model DB329 is a database that stores scenario generation models. The scenario generation model is a machine learning model that generates fraud scenarios to be used for training. As an example, the scenario generation model is a machine learning model that is generated by machine learning using the fraud elements that constitute special fraud as features and a fraud-related scenario as a correct answer label. When a fraud element identified by the correspondence model is input, the scenario generation model outputs a fraud scenario that corresponds to the input fraud element. As an example, the scenario generation model is a neural network, and is data that includes various parameters such as features obtained by machine learning.

[0040] The second voice information DB330 is a database that stores second voice information. The second voice information is voice information for generating a voice message of a virtual agent used in training during a training phase. An example of the second voice information is a fraud scenario generated by a scenario generation model. An example of the second voice information is information about a speech element of a virtual agent identified by a correspondence model. An example of the information about the speech element is a setting value for setting the speech element. A voice message generated based on the second voice information is an example of a second voice message, and a speech element included in the second voice message is an example of a second element.

[0041] Next, the control unit 31 will be described. The control unit 31 controls the overall processing of the information processing device 12. The control unit 31 is realized by a processor such as a CPU (Central Processing Unit), GPU (Graphical Processing Unit), or DSP (Digital Signal Processor) reading a program stored in a storage device, expanding the program in a main storage device such as a RAM, and executing the program. The control unit 31 may be realized by including an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array). The control unit 31 has a model generation unit 310, a voice information generation unit 311, and a training unit 312.

[0042] First, the model generation unit 310 will be described. The model generation unit 310 is a processing unit that learns a correspondence model in the model generation phase. The model generation unit 310 generates fraud data from information on the occurrence of special frauds, etc. As an example, the model generation unit 310 acquires information on the occurrence of special frauds via the communication unit 30. The model generation unit 310 extracts multiple character strings related to special frauds as fraud elements from the occurrence information using an existing natural language processing algorithm, and generates fraud data. The model generation unit 310 stores the generated fraud data in the fraud data DB 320.

[0043] The model generation unit 310 generates collaborator data from the collaborator's responses to the questionnaire. As an example, the model generation unit 310 acquires the collaborator's responses to the questionnaire via the communication unit 30. The model generation unit 310 extracts characteristic elements of the collaborator from the collaborator's responses to the questionnaire, and generates collaborator data. The model generation unit 310 stores the generated collaborator data in the collaborator data DB 321.

[0044] The model generation unit 310 acquires physiological data of the collaborator and generates psychological information of the collaborator. As an example, the model generation unit 310 references the third voice information DB 322 and acquires the third voice information. The model generation unit 310 generates a voice message by using a virtual agent to read out the scenario included in the third voice information in a synthesized voice having speech elements indicated by the third voice information.

[0045] The model generation unit 310 transmits the generated voice message to the terminal 13 via the communication unit 30. The collaborator converses with the virtual agent via the terminal 13. The collaborator's conversation with the virtual agent is an example of the collaborator listening to the voice message. The radar 11 measures physiological data of the collaborator during a predetermined period of time while the collaborator converses with the virtual agent. The model generation unit 310 acquires the physiological data measured by the radar 11 during the predetermined period of time while the collaborator converses with the virtual agent via the communication unit 30. The model generation unit 310 analyzes the acquired physiological data and estimates the collaborator's psychological state during the predetermined period of time while the collaborator converses with the virtual agent. As an example, the model generation unit 310 estimates the psychological state using a machine learning model that analyzes the physiological data. As an example, the machine learning model that analyzes the physiological data is a machine learning model generated by machine learning using human physiological data as features and human psychological states as correct labels. As an example, the collaborator's psychological state is expressed in three levels: positive, negative, and neutral. As an example, the model generation unit 310 associates the speech elements of the virtual agent used in the conversation with the psychological state of the collaborator to generate collaborator psychological information. The model generation unit 310 stores the generated collaborator psychological information in the collaborator psychological information DB 323.

[0046] The model generation unit 310 selects predetermined elements to be used in the element correspondence model from elements such as fraud elements, characteristic elements, and speech elements contained in the generated fraud data, accomplice data, and accomplice psychological information. As an example, the predetermined elements to be used in the element correspondence model are elements that are related to each other. Based on the selected elements, the model generation unit 310 generates an element correspondence model for each combination of related elements.

[0047] Here, an example of an element correspondence model will be described with reference to Figures 9 to 11. Figure 9 is a diagram illustrating an example of an element correspondence model between element types, "scammer's title" and "type of fraud." The element correspondence model is, for example, a matrix representing the probability of simultaneous occurrence of each element. The element correspondence model forms rows and columns for each element type. Matrix 90 has rows for elements representing "scammer's title" as an element type, and columns for elements representing "type of fraud" as an element type, with each component representing the probability of simultaneous occurrence between elements. Specifically, matrix 90 indicates that the probability that a fraudster simulating a city hall employee committed refund fraud is 75.2%, the probability that he committed deposit fraud is 1.0%, and the probability that he committed "it's my son" fraud is 1.0%. In this way, matrix 90 represents the correspondence between element types, "scammer's title" and "type of fraud." The element correspondence model represents the correspondence between elements by the probability of simultaneous occurrence between elements.

[0048] FIG. 10 is a diagram illustrating an example of an element correspondence model between the types of elements, "person" and "impression of person." In matrix 100, each element representing a "person" is a row, each element representing an "impression of person" is a column, and each component represents the probability of simultaneous occurrence between the elements. In this way, matrix 100 represents the correspondence between the types of elements, "person" and "impression of person." FIG. 11 is a diagram illustrating an example of an element correspondence model between the types of elements, "speaking rate" and "mental state." In matrix 110, each element representing a "speaking rate" is a row, each element representing a "mental state" is a column, and each component represents the probability of simultaneous occurrence between the elements. As an example, the probability of simultaneous occurrence of "speaking rate" and "mental state" is the proportion of the period during which the psychological state of the collaborator in each column is estimated during a given period during which the collaborator listens to a voice message containing the elements in each row. In this way, matrix 110 represents the correspondence between the types of speech elements, "speaking rate," and the collaborator's "mental state." The model generation unit 310 stores the generated element correspondence model in the element correspondence model DB 324 .

[0049] The model generation unit 310 learns a correspondence model based on the element correspondence model. The correspondence model is a mathematical model that represents the probability of simultaneous occurrence for a combination of multiple elements. For example, the model generation unit 310 learns a correspondence model based on the element correspondence model. n a so that maximizes i The values of and b are updated and the correspondence model is trained.

[0050] Z n =Σa i r(y jl, y km )+b...Equation (1) Z n is the overall probability of occurrence for a combination of multiple elements. y indicates an element, and (y jl, y km ) indicates a combination of elements. j and k are subscripts that indicate the matrix, and l and m are subscripts that indicate the elements of each matrix row. r(y jl, y km ) is (y jl, y km) is the probability of the simultaneous occurrence of the combination of elements y, and (y jl, y km ) is the correlation coefficient of the combination of elements y. For example, the closer the absolute value of the correlation coefficient is to 1, the higher the correlation between the combination of elements y, and the closer the absolute value is to 0, the lower the correlation. i is r(y jl, y km ) is a weight parameter, and b is a parameter.

[0051] Here, the correspondence model will be described in detail with reference to FIG. 12. FIG. 12 is a diagram for explaining an example of the correspondence model. As an example, matrices A, B, C, and D are element correspondence models having two elements in each row. As an example, element y in formula (1) is a row element of each matrix, and is an array consisting of components in the row direction. Specifically, in the case of matrix A, y A1 and y A2 is element y, and y A1 =[85.2,10.9,···], y A2 =[21.4, 72.1,...]. The model generation unit 310 selects an element correspondence model having the same elements from the element correspondence model DB 324. For example, the matrix 100 in FIG. 10 and the matrix 110 in FIG. 11 both have the same column elements, "positive", "negative", and "neutral", and are examples of matrices having the same elements. Returning to FIG. 12, if matrix A and matrix B, and matrix C and matrix D are matrices having the same column elements, then (y jl, y km ) indicates the combination of elements y in rows between matrices having the same elements. As an example, column 121 of table 120 indicates the combination of elements y in matrices A and B, and column 122 indicates the combination of elements y in matrices C and D. Z n indicates the probability of simultaneous occurrence of each combination when the combination of elements y in rows between matrices having the same elements is combined for each combination of matrices having the same elements. As an example, as shown in FIG. 12, when there are four matrices having two elements y, and two of them have the same elements, the model generation unit 310 generates Z nThe model generation unit 310 calculates Z n The correspondence model is trained so as to maximize the value of . The model generation unit 310 stores the generated correspondence model in the correspondence model DB 325. The correspondence model is, for example, data including information on combinations of various parameters and element y in equation (1) obtained by machine learning. In this way, the model generation unit 310 trains the correspondence model based on the correspondence between the third element and the psychological state of the helper estimated based on the physiological data of the helper listening to the third voice message having the third element.

[0052] Next, a description will be given of the voice information generating unit 311. The voice information generating unit 311 is a processing unit that performs advance preparations for training in the training phase.

[0053] The voice information generation unit 311 generates trainee data from the trainee's responses to the questionnaire. As an example, the voice information generation unit 311 acquires the trainee's responses to the questionnaire via the communication unit 30. The voice information generation unit 311 extracts characteristic elements from the trainee's responses to the questionnaire and generates trainee data. The model generation unit 310 stores the generated trainee data in the trainee data DB 326.

[0054] The voice information generation unit 311 estimates the psychological state of the trainee based on physiological data of the trainee listening to the first voice message having the first element. As an example, the voice information generation unit 311 acquires the first voice information by referring to the first voice information DB 327. The voice information generation unit 311 generates a voice message that reads out a scenario included in the first voice information in a synthetic voice having an utterance element indicated by the first voice information, using a virtual agent. The voice message generated based on the first voice information is an example of the first voice message, and the utterance element included in the first voice message is an example of the first element.

[0055] The voice information generation unit 311 outputs the generated voice message to the terminal 13 via the communication unit 30. The trainee converses with the virtual agent via the terminal 13. The trainee's conversation with the virtual agent is an example of the trainee listening to a voice message. The radar 11 measures physiological data of the trainee during a predetermined period of time while the trainee is conversing with the virtual agent. The voice information generation unit 311 acquires, via the communication unit 30, the physiological data measured by the radar 11 during the predetermined period of time while the trainee is conversing with the virtual agent. The voice information generation unit 311 analyzes the acquired physiological data and estimates the trainee's psychological state during the predetermined period of time while the trainee is conversing with the virtual agent. As an example, the voice information generation unit 311 estimates the psychological state using a machine learning model that analyzes the physiological data. As an example, the trainee's psychological state is expressed in three stages: positive, negative, and neutral. The voice information generation unit 311 associates the speech elements of the virtual agent used in the conversation with the trainee's psychological state to generate trainee psychological information. The model generating unit 310 stores the generated trainee psychological information in the trainee psychological information DB 328.

[0056] The voice information generation unit 311 selects predetermined elements for updating the values of the element correspondence model from among elements such as characteristic elements and speech elements included in the generated trainee data and trainee psychological information. As an example, the voice information generation unit 311 selects elements that reflect the characteristics of the trainee as the predetermined elements. As an example, elements that reflect the characteristics of the trainee are elements whose values vary depending on the trainee, such as the trainee's psychological state in response to speech elements or the types of fraud that the trainee is estimated to be susceptible to based on their psychological characteristics.

[0057] The voice information generation unit 311 updates the values of the components of the element correspondence model generated by the model generation unit 310 based on the selected element. As an example, the voice information generation unit 311 acquires an element correspondence model from the element correspondence model DB 324 and updates the values of the components of the element correspondence model including the selected element with the values of the trainee's elements. Specifically, when a psychological state for an utterance element such as the matrix 110 in FIG. 11 is selected as a predetermined element, the value of each component of the matrix 110 is updated with a value indicating the proportion of a period during which the psychological state in each column is estimated during a predetermined period during which the trainee listens to a voice message having the elements in each row. The voice information generation unit 311 stores the updated element correspondence model in the element correspondence model DB 324.

[0058] The speech information generation unit 311 identifies a second element of the synthetic speech used to train the trainee based on the estimated psychological state of the trainee. Specifically, the information processing device 12 inputs an element correspondence model including the updated element correspondence model into the trained correspondence model, and identifies a fraud element and an utterance element. The utterance element identified by the correspondence model is an example of the second element. As an example, the speech information generation unit 311 inputs each element y of the element correspondence model including the updated element correspondence model into equation (1), and calculates Z of the combination of each element. n Calculate a i and b use values obtained by learning. The voice information generation unit 311 uses the calculated Z n Z with the largest value n Specifically, the speech information generating unit 311 identifies the deception element and speech element to be used for training from the combination of elements that make up Z nIf the combination of elements constituting the message is (type of fraud: it's me fraud, title: son) and (title: son, speaking rate: fast), "type of fraud: it's me fraud" and "title: son" are identified as fraud elements, and "speaking rate: fast" is identified as a speech element. In this way, the voice information generation unit 311 inputs the correspondence between the first element, which is the speech element of the first voice message, and the trainee's psychological state into the correspondence model, and identifies the second element, which is the speech element of the second voice message. The voice information generation unit 311 stores the identified speech element in the second voice information DB 330.

[0059] The voice information generation unit 311 inputs the identified fraud elements into a scenario generation model to generate a scenario to be used for training. The voice information generation unit 311 stores the generated scenario in the second voice information DB 330.

[0060] The training unit 312 outputs a second voice message to be used for training the trainee based on the identified second element. Specifically, the voice information generation unit 311 refers to the second voice information DB 330 and acquires the second voice information to be used for training. The training unit 312 generates a voice message that reads out a scenario included in the second voice information using a synthetic voice having an utterance element indicated by the second voice information, using a virtual agent, and performs training. The voice message generated based on the second voice information is an example of the second voice message, and the utterance element included in the second voice message is an example of the second element.

[0061] The training unit 312 outputs a voice message from the communication unit 30 to the terminal 13, allowing the trainee to converse with the virtual agent. The radar 11 measures physiological data of the trainee using the radar 11 during a predetermined period while the trainee is conversing with the virtual agent. The training unit 312 acquires, via the communication unit 30, the physiological data measured by the radar 11 during the predetermined period while the trainee is conversing with the voice message. The training unit 312 analyzes the acquired physiological data and estimates the trainee's psychological state during the predetermined period while the trainee is conversing with the virtual agent. As an example, the training unit 312 estimates the trainee's psychological state using a machine learning model. Based on the trainee's psychological state, the training unit 312 estimates the trainee's risk of being deceived by a specialized fraud and provides feedback to the trainee. As an example, the training unit 312 estimates the risk of being deceived by a specialized fraud using a machine learning model. As an example, the machine learning model that estimates the risk of being deceived by a specialized fraud is a machine learning model generated by machine learning using a person's psychological state as a feature and the risk of the person being deceived by a specialized fraud as a correct answer label.

[0062] In this way, the information processing device 12 learns a correspondence model based on fraud elements, characteristic elements, and speech elements. Therefore, when the trainee's characteristic elements are input, the information processing device 12 can identify fraud elements and speech elements that correspond to the trainee's characteristics. The information processing device 12 can generate a voice message in which a fraud scenario that corresponds to the trainee's characteristics is spoken by a synthesized voice having speech elements that correspond to the trainee's characteristics. The information processing device 12 can conduct training using a voice message that reflects the trainee's characteristics. Specifically, for example, if a trainee wants to generate a voice that feels familiar, the information processing device 12 can identify the speaker's title and speech elements that the trainee finds familiar. Furthermore, for elements that do not need to reflect the trainee's characteristics, the values of the elements at the time of learning can be used, allowing tendencies toward common special frauds and the like to be reflected in the speech elements and scenarios. The information processing device 12 can generate voice messages that trainees are likely to be fooled by, thereby conducting effective training. Through training, trainees can learn about their own risk of committing special fraud, which helps prevent special frauds.

[0063] [Processing flow] 13 is a diagram illustrating an example of a process flow of the model generation phase of the information processing device 12 according to the embodiment. With reference to FIG. 13, an example of a process flow of the model generation phase of the information processing device 12 according to the embodiment will be described.

[0064] The model generation unit 310 acquires information on the occurrence of fraud, the collaborator's responses to the questionnaire, and the collaborator's physiological data via the communication unit 30 (step S10). As an example, the model generation unit 310 generates a voice message for the virtual agent based on the third voice information and transmits it to the terminal 13. The model generation unit 310 acquires the collaborator's physiological data measured by the radar 11 during a predetermined period in which the collaborator converses with the virtual agent.

[0065] The model generation unit 310 extracts fraud elements from the fraud occurrence information and generates fraud data (step S11). The model generation unit 310 extracts characteristic elements of the collaborator that represent the collaborator's attributes, the collaborator's impression of a specific person, the collaborator's psychological characteristics, etc. from the collaborator's questionnaire responses, and generates collaborator data. The model generation unit 310 estimates the collaborator's psychological state from the collaborator's physiological data, extracts characteristic elements of the collaborator that represent the collaborator's psychological state in response to utterance elements, and generates collaborator psychological information. The model generation unit 310 stores the generated fraud data in the fraud data DB 320, stores the generated collaborator data in the collaborator data DB 321, and stores the generated collaborator psychological information in the collaborator psychological information DB 323.

[0066] The model generation unit 310 generates an element correspondence model based on the generated fraud data, collaborator data, and collaborator psychological information (step S12). As an example, the model generation unit 310 acquires fraud data from the fraud data DB 320, acquires collaborator data from the collaborator data DB 321, and acquires collaborator psychological information from the collaborator psychological information DB 323. The model generation unit 310 selects relevant elements from elements such as fraud elements, characteristic elements, and utterance elements contained in the acquired fraud data, collaborator data, and collaborator psychological information, and generates an element correspondence model for each combination of relevant elements. The model generation unit 310 stores the generated element correspondence model in the element correspondence model DB 324.

[0067] The model generation unit 310 learns a correspondence model based on the generated element correspondence model (step S13). As an example, the model generation unit 310 acquires an element correspondence model from the element correspondence model DB 324. The model generation unit 310 learns a correspondence model based on the element correspondence model by learning Z n a so that maximizes i The model generation unit 310 updates the values of a and b and learns the correspondence model. The model generation unit 310 stores the learned correspondence model in the correspondence model DB 325.

[0068] 14 is a diagram illustrating an example of a process flow in the training phase of the information processing device 12 according to the embodiment. An example of a process flow in the training phase of the information processing device 12 according to the embodiment will be described with reference to FIG.

[0069] The voice information generation unit 311 acquires the trainee's responses to the questionnaire and physiological data of the trainee via the communication unit 30 (step S20). As an example, the voice information generation unit 311 generates a voice message for the virtual agent based on the first voice information and transmits it to the terminal 13. The voice information generation unit 311 acquires the trainee's physiological data measured by the radar 11 during a predetermined period when the trainee is conversing with the virtual agent.

[0070] The voice information generation unit 311 extracts trainee attributes, the trainee's impressions of specific people, and trainee characteristic elements that represent the trainee's psychological characteristics from the trainee's questionnaire responses, and generates trainee data (step S21). The voice information generation unit 311 estimates the trainee's psychological state from the trainee's physiological data, extracts trainee characteristic elements that represent the trainee's psychological state in response to utterance elements, and generates trainee psychological information. The voice information generation unit 311 stores the generated trainee data in the trainee data DB 326, and stores the generated trainee psychological information in the trainee psychological information DB 328.

[0071] The voice information generation unit 311 updates the element correspondence model based on the generated trainee data and trainee psychological information (step S22). As an example, the voice information generation unit 311 acquires trainee data from the trainee data DB 326 and acquires trainee psychological information from the trainee psychological information DB 328. The voice information generation unit 311 selects predetermined elements for updating the values of the element correspondence model from the characteristic elements and speech elements included in the acquired trainee data and trainee psychological information. The voice information generation unit 311 updates the values of the components of the element correspondence model generated by the model generation unit 310 based on the selected elements to reflect the characteristics of the trainee. The voice information generation unit 311 stores the updated element correspondence model in the element correspondence model DB 324.

[0072] The speech information generation unit 311 identifies the deception elements and speech elements to be used for training based on the correspondence model (step S23). As an example, the speech information generation unit 311 inputs the element y of the element correspondence model including the updated element correspondence model into the correspondence model, and calculates Z of the combination of each element. n The voice information generating unit 311 calculates the calculated Z n Z with the largest value nThe deception elements and speech elements to be used for training are identified from the combination of elements that make up the speech elements. The speech information generation unit 311 stores the identified speech elements in the second speech information DB 330. The speech information generation unit 311 inputs the identified deception elements into a scenario generation model to generate a scenario to be used for training. The speech information generation unit 311 stores the generated scenario in the second speech information DB 330.

[0073] The training unit 312 performs training based on the second audio information (step S24). As an example, the training unit 312 acquires second audio information from the second audio information DB 330. The training unit 312 generates an audio message in which the virtual agent reads out the scenario included in the second audio information in a synthesized voice having speech elements indicated by the second audio information. The training unit 312 transmits the audio message to the terminal 13, has the trainee converse with the virtual agent, and performs training. The training unit 312 acquires physiological data of the trainee measured by the radar 11 during a predetermined period during which the trainee converses with the virtual agent. The training unit 312 estimates the trainee's psychological state from the physiological data. The training unit 312 estimates the trainee's risk of being deceived by special fraud from the trainee's psychological state and provides feedback to the trainee.

[0074] [Variations] In this embodiment, the information processing device 12 acquires physiological data measured by a millimeter-wave radar, but physiological data measured by another sensing device may also be used. For example, the other sensing device may be a camera, a wearable device, or a microphone for collecting sound, and the physiological data may be physiological response data obtained from image data or sound data.

[0075] In this embodiment, the collaborator is a victim who has been a victim of special fraud in the past, but may also be a person who has never been a victim of special fraud in the past.

[0076] In this embodiment, the collaborator conversed with a virtual agent in the model generation phase, but the collaborator may also converse with a live person. In this case, collaborator data may be generated by linking it to the speech elements of the person who is the conversation partner.

[0077] In this embodiment, the trainee conversed with a virtual agent in the advance preparation for the training phase, but the trainee may also converse with a live person. In this case, trainee data may be generated by linking it to the speech elements of the person with whom the trainee converses.

[0078] In this example, an embodiment is described that targets special fraud prevention training, but the events that are the subject of the training and the people who are the targets of the training are not limited to those in this example. As an example, the information processing device 12 may conduct training on how to respond to customer harassment in a call center. As an example, the information processing device 12 may conduct evacuation training in the event of a disaster or fire. As an example, the information processing device 12 may conduct training on how to respond to an accident in the event of an accident.

[0079] Although the present embodiment describes an embodiment in which there is one trainee, there may be multiple trainees. The information processing device 12 may store trainee data in a storage unit for each trainee ID representing an arbitrary identifier for identifying a trainee. The information processing device 12 may store trainee psychological information and an updated element correspondence model for each trainee ID in a storage unit. The information processing device 12 may generate second audio information for each trainee ID based on the correspondence model, and perform training according to the characteristics of each trainee.

[0080] In this embodiment, the element y in equation (1) is a row element of each matrix and is an array composed of row-direction components, but the element y is not limited to the content of this embodiment. As an example, if the same element between matrices is a row of each matrix, the element y may be a column element of each matrix and an array composed of column-direction components. As an example, if the same element between matrices is a row of one matrix and a column of the other matrix, the element y may be an array composed of row or column components of different elements of each matrix.

[0081] [effect] As described above, the information processing device 12 estimates the psychological state of the trainee based on physiological data of the trainee listening to the first voice message having the first element, identifies the second element of the synthetic voice to be used for training the trainee based on the estimated psychological state of the trainee, and outputs the second voice message to be used for the trainee's training based on the identified second element. In this way, the information processing device 12 can identify the speech element of the synthetic voice according to the trainee's characteristics and perform training using a voice message using the speech element of the synthetic voice according to the trainee's characteristics.

[0082] In this embodiment, the physiological data is physiological data measured by radar. Radar is a non-contact sensor, and unlike contact sensors such as wearable devices, it does not need to be attached to the body of the person being measured. Therefore, the information processing device 12 can perform training in a situation where the trainee does not feel restricted or bothered by wearing a sensor. The information processing device 12 can acquire physiological data even during long training periods.

[0083] The information processing device 12 may input the correspondence between the first element and the trainee's psychological state into a correspondence model, which is a machine learning model, to identify the second element. This allows the information processing device 12 to identify the speech element according to the trainee's characteristics using the machine learning model.

[0084] The information processing device 12 may learn a correspondence model, which is a machine learning model, based on the correspondence between the third element and the psychological state of the helper estimated based on physiological data of the helper listening to the third voice message having the third element, thereby enabling the information processing device 12 to learn a correspondence model that estimates a speech element according to the characteristic element.

[0085] In this embodiment, the training is training to prevent the occurrence of special frauds. As a result, the information processing device 12 can generate voice messages that pose a high risk of deceiving the trainee according to the characteristics of the trainee, thereby improving the effectiveness of the training to prevent special frauds. The information processing device 12 can allow the trainee to experience special frauds, thereby preventing the occurrence of special frauds.

[0086] [Hardware configuration] Fig. 15 is a diagram illustrating an example of the hardware configuration of an information processing device according to an embodiment. An example of the hardware configuration of the information processing device 12 will be described with reference to Fig. 15. As shown in Fig. 15, the information processing device 12 includes, as an example, a CPU 150, a RAM 151, an input / output interface 152, a communication interface 153, and an HDD 154 as components, and these components are connected via a bus 155.

[0087] The CPU 150 is a processor that runs a process that executes the functions of the control unit 31. Specifically, the CPU 150 reads a program that executes the same functions as the control unit 31 from the HDD 154 or the like, loads it into the RAM 151 or the like, and executes the process that executes the functions of the control unit 31. The CPU 150 may obtain the program and data used to execute the program from a medium reading device, or may obtain the data via the communication interface 153. The CPU 150 may have one or more processor cores. The information processing device 12 may be provided with a processor other than a CPU, or may be provided with multiple types of processors.

[0088] The RAM 151 operates as a main storage device of the information processing device 12, and stores programs read from an auxiliary storage device such as the HDD 154, data used to execute the programs, etc. The information processing device 12 may include a memory other than the RAM, or may include multiple memories.

[0089] The input / output interface 152 is an interface that inputs signals to the information processing device 12 and outputs signals from the information processing device 12. As an example, the input / output interface 152 receives signals from input devices such as a keyboard or a mouse connected to the information processing device 12, and transmits signals such as images to output devices such as a display connected to the information processing device 12. A plurality of input devices and output devices may be connected to the information processing device 12 via the input / output interface 152.

[0090] The communication interface 153 is an interface for connecting the information processing device 12 to a network. For example, the standard of the communication interface 153 may be a wired LAN communication standard such as Ethernet, a wireless LAN communication standard such as Wi-Fi (registered trademark), a wireless mobile communication standard such as local 5G, or a telephone line standard.

[0091] The HDD 154 operates as an auxiliary storage device for storing programs for executing the functions of the OS (Operating System) and the control unit 31 of the information processing device 12, and data used by the programs. The information processing device 12 may be provided with an auxiliary storage device other than an HDD, such as an SSD (Solid State Drive), or may be provided with multiple auxiliary storage devices.

[0092] Although the embodiments of the present invention have been described above, the embodiments of the present invention are not limited to those described above. The present invention may be embodied in various different forms other than the above-described embodiments.

[0093] The processing procedures, processing methods, names of parts, various data, etc. shown in the embodiments and drawings are merely examples and may be changed arbitrarily unless otherwise specified.

[0094] The functional configurations and hardware configurations shown in the embodiments and drawings are merely examples, and therefore do not necessarily have to be configured or arranged as shown, and may be changed as desired unless otherwise specified. For example, the components of the functional configurations and hardware configurations may be distributed or integrated in any unit, such as a function.

[0095] The information processing program according to the embodiment is not limited to being executed by the information processing device 12. For example, the control program according to the embodiment may be executed across a plurality of devices.

[0096] Furthermore, the information processing program according to the embodiments may be distributed via a network such as the Internet. The information processing program according to the embodiments may be recorded on a computer-readable recording medium and sold. Examples of the computer-readable recording medium include a CD-ROM (Compact Disc Read Only Memory), a DVD (Digital Versatile Disc), a USB (Universal Serial Bus) memory, a floppy disk, and an MO (Magneto-Optical Disk). The information processing program according to the embodiments may be read from the computer-readable recording medium by a computer and executed by the computer. [Explanation of symbols]

[0097] 10 Information Processing Systems 11 Radar 12 Information processing equipment 13 Terminals 14 People 15 Network 16 Network 30 Communications Department 31 Control Unit 32 Storage section 310 Model Generation Unit 311 Voice Information Generation Unit 312 Training Department 320 Fraud Data DB 321 Collaborator Data DB 322 Third Audio Information DB 323 Collaborator psychological information DB 324 Element Correspondence Model DB 325 Correspondence Model DB 326 Trainee Data DB 327 First Audio Information DB 328 Trainee psychological information DB 329 Scenario Generation Model DB 330 Second Audio Information DB

Claims

1. estimating a psychological state of the trainee based on physiological data of the trainee listening to the first audio message having the first element; identifying a second element of a synthetic speech to be used in training the trainee based on the estimated psychological state of the trainee; and outputting a second voice message to be used in the training of the trainee based on the identified second element. An information processing program that causes a computer to execute a process.

2. The physiological data is physiological data measured by radar. The information processing program according to claim 1 .

3. The process of identifying the second element is a process of inputting a correspondence between the first element and the psychological state of the trainee into a machine learning model. The information processing program according to claim 1 .

4. The machine learning model is a machine learning model trained based on a correspondence between a psychological state of a participant who is listening to a third audio message having a third element, the psychological state being estimated based on physiological data of the participant, and the third element. The information processing program according to claim 3 .

5. The training is aimed at preventing the occurrence of special fraud. The information processing program according to claim 1 .

6. estimating a psychological state of the trainee based on physiological data of the trainee listening to the first audio message having the first element; identifying a second element of a synthetic speech to be used in training the trainee based on the estimated psychological state of the trainee; and outputting a second voice message to be used in the training of the trainee based on the identified second element. An information processing method in which processing is performed by a computer.

7. estimating a psychological state of the trainee based on physiological data of the trainee listening to the first audio message having the first element; identifying a second element of a synthetic speech to be used in training the trainee based on the estimated psychological state of the trainee; and outputting a second voice message to be used in the training of the trainee based on the identified second element. An information processing device having a control unit.

Citation Information

Patent Citations

  • Speech synthesizer, speech synthesis program and speech synthesis method

    JP2012022447A