system

JP7927909B1Active Publication Date: 2026-10-01SOFTBANK GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025044947
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2026-10-01
Estimated Expiration
2045-03-19

Smart Images

  • Figure 0007927909000001_ABST
    Figure 0007927909000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A system including a harassment determination means that receives a dialogue history as input and determines whether or not the dialogue history is harassment, and a video generation means that outputs a video that is automatically generated based on the dialogue history determined to be harassment by the harassment determination means.
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] The technology of the present disclosure relates to a system. [[Background Art]]

[0002] Patent Document 1 discloses a persona chatbot control method executed by at least one processor, the method comprising: receiving a user utterance; adding the user utterance to a prompt including an instruction associated with a description of a character of a chatbot; encoding the prompt; and inputting the encoded prompt into a language model to generate a chatbot utterance responsive to the user utterance. [[Prior Art Documents]] [[Patent Documents]]

[0003] [[Patent Document 1]] Japanese Unexamined Patent Application Publication No. 2022-180282 [[Summary of the Invention]] [[Problem to be Solved by the Invention]]

[0004] In modern corporate sales activities, harassment issues have become an important problem. However, it is difficult especially for people born in the Showa era to understand what constitutes power harassment and sexual harassment. Under such circumstances, specific means for preventing harassment before it occurs are demanded. [[Means for Solving the Problem]]

[0005] The present invention provides a harassment determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment, and a video generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the harassment determination means. This makes it possible to prevent harassment from occurring in the first place and to provide concrete material for reflection and improvement in the case of harassment that has occurred. [Brief explanation of the drawing]

[0006] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Embodiment 1 of Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1 of Form Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 of Embodiment 2. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2 of Form Example 2. [Figure 15] This is a sequence diagram showing the processing flow of the data processing system in Embodiment 3 of Example 3. [Figure 16] This is a sequence diagram showing the processing flow of the data processing system in Application Example 3 of Form Example 3. [Figure 17] This is a sequence diagram showing the processing flow of the data processing system in Example 1 of the Form 1 when an emotion engine is combined. [Figure 18] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1 of Form Example 1 when an emotion engine is combined. [Figure 19] This is a sequence diagram showing the processing flow of the data processing system in Example 2 of the Form 2 when an emotion engine is combined. [Figure 20] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2 of Form Example 2 when an emotion engine is combined. [Figure 21] This is a sequence diagram showing the processing flow of the data processing system in Example 3 of the Form 3 when an emotion engine is combined. [Figure 22] This is a sequence diagram showing the processing flow of the data processing system in Application Example 3 of Form Example 3 when an emotion engine is combined. [Figure 23] This is a sequence diagram showing the processing flow of a data processing system in another embodiment. [Modes for carrying out the invention]

[0007] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0008] First, the terms used in the following description will be explained.

[0009] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Further, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), and TPU (TENSOR PROCESSING UNIT (Registered Trademark)).

[0010] In the following embodiments, labeled RAM (Random Access Memory) is a memory that temporarily stores information, and is used as a work memory by the processor.

[0011] In the following embodiments, labeled storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0012] In the following embodiments, labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. Communication I / F manages communication between a plurality of computers. Examples of communication standards applied to communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (Registered Trademark), and Bluetooth (Registered Trademark).

[0013] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0014] [First Embodiment]

[0015] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0016] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0017] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0018] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0019] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0020] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0021] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0022] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0023] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0024] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0025] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0026] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.

[0027] "Example of form 1"

[0028] One embodiment of the present invention provides an automated harassment analysis system. This system receives a conversation history from a user as input and has a harassment determination means that determines whether or not the conversation history constitutes harassment. The harassment determination means analyzes the conversation history using natural language processing technology and determines whether or not harassment such as power harassment or sexual harassment has occurred.

[0029] "Example of form 2"

[0030] Furthermore, the system of the present invention includes a video generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the harassment determination means. The video generation means generates a video that reflects the content of the dialogue history determined to be harassment and presents it to the user. For example, if the dialogue history determined to be harassment is related to power harassment, it generates a video that points out the problems of power harassment.

[0031] "Example of form 3"

[0032] The system of the present invention can prevent harassment from occurring and provide concrete material for reflection and improvement in cases of harassment that have occurred. Specifically, by having the user input their conversation history into the system, the presence or absence of harassment is determined, and if harassment is found, a video reflecting the content is generated and presented. As a result, the user can recognize that their words and actions may constitute harassment and take concrete actions to improve them.

[0033] The following describes the processing flow for each example of the form.

[0034] "Example of form 1"

[0035] Step 1: The system receives the user's interaction history.

[0036] Step 2: The system's harassment detection mechanism analyzes the conversation history. This analysis uses natural language processing technology to determine whether or not harassment, such as power harassment or sexual harassment, has occurred.

[0037] "Example of form 2"

[0038] Step 1: The system receives the conversation history that has been determined to be harassment by the harassment detection method.

[0039] Step 2: The system's video generation mechanism automatically generates a video based on the conversation history that was identified as harassment. The video content reflects the content of the harassment.

[0040] Step 3: Present the generated video to the user.

[0041] "Example of form 3"

[0042] Step 1: The user enters their conversation history into the system.

[0043] Step 2: The system determines whether harassment occurred, and if so, generates a video reflecting the details of the harassment.

[0044] Step 3: Present the generated video to the user, allowing them to recognize that their words and actions may constitute harassment.

[0045] Step 4: Based on the presented video, users take specific actions to improve the harassment situation.

[0046] (Example 1)

[0047] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0048] In today's workplace, harassment is a serious problem, and its detection and response are crucial. However, traditional methods have made it difficult to efficiently analyze dialogue history and accurately determine the presence or absence of harassment. Furthermore, there has been a lack of visual representations of harassment, making it difficult to understand and share the problem.

[0049] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0050] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; an analysis means that analyzes the dialogue history using natural language processing technology; and an evaluation means that evaluates the analysis results using a generative AI model and determines whether or not harassment has occurred. This enables efficient analysis of dialogue history and accurate determination of harassment.

[0051] "Dialogue history" refers to a record of conversations and messages exchanged between users, and is data saved in text format.

[0052] "Information processing means" refers to a device or program that has the function of analyzing input data and making decisions based on specific conditions.

[0053] "Generation means" refers to a device or program for automatically creating new visual information or content based on input information.

[0054] "Natural language processing technology" refers to technologies for understanding and analyzing human language using computers, and includes text tokenization and grammatical analysis.

[0055] "Analysis means" refers to a device or program for analyzing input data in detail and understanding its structure and meaning.

[0056] A "generative AI model" is a model that uses artificial intelligence technology to learn from data and is trained to perform a specific task.

[0057] "Evaluation means" refers to a device or program that makes a judgment based on analyzed data according to specific criteria.

[0058] To implement this invention, the user must first input their conversation history. The user uses their device to input the content of workplace conversations and messages in text format. The inputted conversation history is then sent from the device to the server.

[0059] The server uses Python's natural language processing libraries, NLTK and spaCy, to analyze the received dialogue history. Using this software, the server tokenizes the text data, tags it with parts of speech, and analyzes its grammatical structure.

[0060] Next, the server evaluates the analysis results using a generative AI model. This model is built using deep learning frameworks such as TENSORFLOW® and PyTorch, and is pre-trained on a dataset related to harassment. The model detects elements of power harassment and sexual harassment contained in the dialogue history and determines whether or not they exist.

[0061] For example, if a user enters a conversation history such as "My boss forces me to work unreasonable overtime every day," the server analyzes this data and uses a generative AI model to determine that "it contains elements of power harassment." This result is sent back to the terminal and displayed to the user.

[0062] An example of a prompt message is, "Please determine if this conversation history contains elements of harassment." By using this prompt message, the server can efficiently analyze the conversation history and accurately determine whether or not harassment has occurred.

[0063] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0064] Step 1:

[0065] Users input conversation history in text format using a terminal. This input data includes the content of workplace conversations and messages. The entered conversation history is sent from the terminal to the server.

[0066] Step 2:

[0067] The server uses natural language processing libraries such as NLTK and spaCy to analyze the dialogue history received from the terminal. The server tokenizes the input text data, tags it with parts of speech, and analyzes its grammatical structure. This analysis extracts structural information from the text.

[0068] Step 3:

[0069] The server inputs the analyzed data into a generating AI model. This model is built using TensorFlow and PyTorch and has been pre-trained on a dataset related to harassment. Based on the input data, the model detects elements of power harassment and sexual harassment and determines whether or not they exist. This determination result is then output.

[0070] Step 4:

[0071] The server returns the judgment result obtained from the generated AI model to the terminal. The judgment result indicates whether or not the conversation history contains elements of harassment.

[0072] Step 5:

[0073] The terminal displays the judgment results received from the server to the user. For example, a message such as "This conversation contains elements of power harassment" might be displayed. This allows the user to understand the content of the conversation history and take necessary measures.

[0074] (Application Example 1)

[0075] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server," and the smart device 14 will be referred to as a "terminal."

[0076] Workplace harassment is a serious problem that negatively impacts employees' mental health and workplace productivity. However, early detection and appropriate response to signs of harassment are difficult. Traditional methods often rely on victims reporting the issue themselves, and problems may go undetected until they become severe. An effective system is needed to improve this situation and prevent harassment proactively.

[0077] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0078] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a conversion means that converts speech into text; and an analysis means that analyzes the text converted by the conversion means and detects signs of harassment. This makes it possible to detect signs of harassment in the workplace environment in real time and notify administrators.

[0079] "Dialogue history" refers to a record of conversations between users, which is saved as audio or text data.

[0080] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not it constitutes harassment.

[0081] "Generation means" refers to a device or program that has the function of automatically creating relevant information and content based on the dialogue history that has been determined to be harassment by the determination means.

[0082] "Conversion means" refers to a device or program that has the function of converting audio data into text data.

[0083] "Analysis means" refers to a device or program that has the function of analyzing text data and detecting signs of harassment.

[0084] A "notification means" is a device or program that has the function of sending warnings or information to administrators or relevant parties based on signs of harassment detected by an analysis means.

[0085] The system for implementing this invention mainly consists of a server and a user terminal. The server receives dialogue history as input and is equipped with a determination means for determining whether or not the dialogue history constitutes harassment. This determination means analyzes the dialogue history using natural language processing technology to determine whether or not harassment has occurred. Specifically, it uses the Google® Cloud Natural Language API to analyze the text data.

[0086] The user terminal is equipped with a conversion mechanism for converting speech to text. This conversion mechanism uses the Google Cloud Speech-to-Text API to convert speech data into text data. The converted text data is sent to a server and analyzed by a determination mechanism.

[0087] If the server detects signs of harassment as a result of its analysis, it will send a warning to the administrator using a notification system. This notification system uses Firebase Cloud Messaging to provide real-time notifications.

[0088] For example, if the statement "You're always useless" is made during a conversation at work, this statement may be judged as potentially being harassment, and a notification will be sent to the manager. An example of a prompt to the generating AI model would be, "Please determine whether the following conversation constitutes harassment: 'You're always useless.'"

[0089] In this way, the system can detect signs of harassment in the workplace in real time, enabling a rapid response.

[0090] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0091] Step 1:

[0092] The user terminal captures workplace conversations as audio data. This audio data is collected in real time through the user terminal's microphone.

[0093] Step 2:

[0094] The user's device converts the acquired audio data into text data using the Google Cloud Speech-to-Text API. This conversion process analyzes the audio signal and generates the corresponding text. The input is audio data, and the output is text data.

[0095] Step 3:

[0096] The user's terminal sends the converted text data to the server. The server analyzes the received text data using the Google Cloud Natural Language API. The analysis evaluates linguistic features within the text and detects signs of harassment. The input is text data, and the output is a determination of whether or not harassment occurred.

[0097] Step 4:

[0098] If the server detects signs of harassment based on its analysis, it will send a notification to the administrator using Firebase Cloud Messaging. The notification will include detailed information such as the nature and time of the harassment. The input is the analysis result, and the output is the notification to the administrator.

[0099] Step 5:

[0100] The administrator will take necessary actions based on the received notification. This includes confirming with relevant parties and conducting further investigations. The administrator's actions will be based on the system output.

[0101] (Example 2)

[0102] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 will be referred to as the "terminal".

[0103] In today's workplace, harassment is a serious problem, and early detection and appropriate response are essential. However, traditional methods have the drawback of subjective harassment assessments, making it difficult to implement appropriate education and countermeasures. Therefore, a system is needed that objectively assesses harassment and provides educational visual information based on that assessment.

[0104] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0105] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a visual information generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the information processing means; and a presentation means that presents the visual information generated by the visual information generation means to the user. This makes it possible to objectively determine harassment and provide educational visual information based on its content.

[0106] "Dialogue history" refers to a record of conversations between users, which is saved in text format.

[0107] "Information processing means" refers to a device or program that has the function of analyzing input data and making decisions based on specific conditions.

[0108] "Visual information generation means" refers to a device or program that has the function of automatically creating visual content based on analysis results.

[0109] "Presentation means" refers to a device or program for displaying or providing generated visual information to a user.

[0110] "Harassment" refers to inappropriate words or actions towards others in the workplace or other environments, and is an act that causes mental or physical distress.

[0111] A description of embodiments for carrying out this invention will be given.

[0112] The server receives the conversation history provided by the user. The user inputs the conversation history in text format via their terminal and sends it to the server. This conversation history is used as data for harassment assessment.

[0113] The server analyzes the dialogue history using a generative AI model. Specifically, it uses natural language processing techniques to determine whether the content of the dialogue constitutes harassment. In this process, a generative AI model with excellent natural language processing capabilities is used as the general-purpose model. The server inputs the prompt "Please determine whether this dialogue history constitutes harassment" into the generative AI model.

[0114] If harassment is detected, the server automatically generates visual information based on the detection result using a visual information generation system. This visual information includes educational content and aims to inform users about the problems and countermeasures against harassment. General video editing software may be used to generate the visual information.

[0115] For example, if a user enters a conversation history stating that "a superior excessively reprimanded a subordinate," the server will determine this to be harassment. Next, the server generates and presents visual information on the theme of "the impact of inappropriate behavior in the workplace and countermeasures." This visual information includes specific examples and countermeasures, allowing the user to deepen their understanding.

[0116] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0117] Step 1:

[0118] The user uses a terminal to input the conversation history in text format and sends it to the server. The input data is a record of the conversation between users and serves as the basis for determining whether harassment occurred. The server receives this conversation history and prepares for the next processing step.

[0119] Step 2:

[0120] The server inputs the received dialogue history into the generating AI model. Specifically, it uses the prompt "Determine whether this dialogue history constitutes harassment" to have the generating AI model analyze the dialogue history. The generating AI model uses natural language processing techniques to analyze the content of the dialogue and determine whether or not harassment is present. The output of this step is the result of the harassment determination.

[0121] Step 3:

[0122] The server receives the judgment result from the generating AI model, and if it determines that harassment has occurred, it generates visual information using a visual information generation means. Specifically, based on the judgment result, it creates visual information that includes educational content. Video editing software may be used to generate the visual information. The output of this step is the generated visual information.

[0123] Step 4:

[0124] The server presents the generated visual information to the user. The user can view the visual information through their device and learn about harassment issues and countermeasures. The output of this step is the visual information presented to the user.

[0125] (Application Example 2)

[0126] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0127] Workplace harassment is a serious problem that harms employees' mental health and leads to decreased productivity. However, detecting and preventing harassment is difficult, and prompt and appropriate measures are necessary, especially in situations where real-time response is required. Traditional methods often only identify harassment incidents after they have occurred, making immediate response difficult.

[0128] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0129] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the determination means; and a presentation means that presents the video generated by the generation means to the user. This makes it possible to detect harassment in the workplace environment in real time and immediately generate and present educational videos.

[0130] "Dialogue history" refers to data that records the content of a conversation and is saved in audio or text format.

[0131] A "determination means" is a device or program that has the function of analyzing input data and making a judgment based on specific conditions.

[0132] "Generation means" refers to a device or program that has the function of automatically creating new content based on input information.

[0133] "Presentation means" refers to a device or program that has the function of providing generated information or content to the user visually or audibly.

[0134] A "generative AI model" is an algorithm or program that uses artificial intelligence technology to generate new information or content from data.

[0135] A "prompt" is an instruction or question given to a generative AI model to obtain a specific output.

[0136] The system for implementing this invention mainly consists of a server and a user terminal. The server receives dialogue history as input and converts the audio data into text using a speech recognition API (e.g., Google Cloud Speech-to-Text) to determine whether the dialogue history constitutes harassment. Next, it analyzes the converted data using a natural language processing library (e.g., spaCy) to determine the possibility of harassment.

[0137] Based on the determined results, the server uses a generation AI model (e.g., OpenAI's GPT-4) to generate a video that points out the problems of harassment and proposes solutions. In this process, the generation AI model receives prompt text to obtain specific output. An example of a prompt text would be, "This conversation has been determined to be power harassment. Please generate a video that points out the problems of power harassment and proposes solutions."

[0138] The generated video is displayed on the user's device. The user's device is a smartphone or smart glasses, and its role is to visually provide the generated video to the user. This allows the user to understand the harassment issue in real time and take appropriate action.

[0139] For example, if a workplace conversation includes a statement like, "You're always so slow, you need to work harder," the server will identify this as harassment and input a prompt into a generation AI model to create an educational video. This video is then immediately presented to the relevant parties via the user's terminal.

[0140] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0141] Step 1:

[0142] The server receives audio data sent from the user's terminal. This audio data is a recording of a workplace conversation. The server converts this audio data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is audio data, and the output is text data.

[0143] Step 2:

[0144] The server analyzes text data using a natural language processing library (e.g., spaCy). The purpose of the analysis is to determine whether there are signs of harassment in the text. The input is text data, and the output is a judgment result indicating the possibility of harassment. Specifically, it analyzes keywords and context within the text and scores the likelihood of harassment.

[0145] Step 3:

[0146] Based on the judgment result, the server inputs a prompt message into the generating AI model (e.g., OpenAI's GPT-4). The prompt message is: "This conversation has been determined to be harassment. Please generate a video that points out the problems with harassment and proposes solutions." The input is the judgment result and the prompt message, and the output is the content of the video generated by the generating AI model. Specifically, the generating AI model analyzes the prompt message and generates appropriate video content.

[0147] Step 4:

[0148] The server uses a video generation tool to create the actual video based on the generated video content. The input is the video content generated by the AI ​​model, and the output is the completed video file. Specifically, the video generation tool converts text-based content into a visual video.

[0149] Step 5:

[0150] The server sends the completed video file to the user's terminal. The user's terminal then displays this video to the user. The input is a video file, and the output is the display of the video to the user. Specifically, the user's terminal plays the video and provides the user with visual information.

[0151] (Example 3)

[0152] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0153] In today's workplace, harassment remains a serious problem, and there is a need to prevent it and to implement concrete corrective measures when it occurs. However, traditional methods have the challenge of making it difficult to recognize harassment and obtain specific feedback for improvement.

[0154] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[0155] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a visual information generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the information processing means; and a presentation means that presents the generated visual information to the user. This enables the user to recognize the possibility that their words and actions constitute harassment and to take concrete corrective measures.

[0156] "Dialogue history" refers to a record of conversations a user has had in the past, and is data entered in text format.

[0157] "Information processing means" refers to technical means for analyzing input dialogue history and determining whether or not harassment has occurred.

[0158] "Visual information generation means" refers to a technical means that automatically generates visual information reflecting the content of harassment based on a conversation history that has been determined to be harassment.

[0159] A "presentation means" is a technical means that provides generated visual information to a user, enabling the user to visually confirm it.

[0160] "Harassment" refers to inappropriate words or actions towards others in the workplace or in society, and particularly includes inappropriate behavior and sexual harassment in the workplace.

[0161] A description of embodiments for carrying out this invention will be given.

[0162] The user first inputs their conversation history into the system via their terminal. The conversation history is in text format and includes past conversations. The user copies and pastes the conversation history into the input form on their terminal and clicks the "Start Analysis" button.

[0163] The server receives the conversation history sent by the user. The received data is preprocessed for natural language processing. Specifically, text cleaning and tokenization are performed. The server inputs the preprocessed text data into a natural language processing engine, which is an AI model for generating conversations. This engine analyzes the conversation content and determines whether or not harassment has occurred. The result of the determination is output as a score indicating whether or not harassment is likely.

[0164] Based on the assessment results, the server extracts specific details of any harassment that may have occurred. Using this information, it creates a scenario for generating visual information. The server then uses video editing software to generate visual information that reflects the extracted harassment. This visual information is designed to allow users to visually confirm their own behavior. The generated visual information is sent to the user's device for viewing.

[0165] As a concrete example, if a user wants to "check whether what they said in yesterday's meeting was appropriate," they input the conversation history into the system. An example of a prompt would be, "Please analyze what I said in yesterday's meeting and assess the possibility of harassment." Based on this prompt, the server performs the analysis, generates visual information as needed, and provides it to the user. The flow of specific processing in Example 3 will be explained using Figure 15.

[0166] Step 1:

[0167] The user enters their conversation history using a terminal. The input is in text format; the user copies and pastes the conversation history into the input form on the terminal and clicks the "Start Analysis" button. The entered conversation history is sent to the server.

[0168] Step 2:

[0169] The server preprocesses the dialogue history received from the user. Specifically, it performs text cleaning (removing unnecessary spaces and special characters) and tokenization (dividing words and phrases into parsable units). This preprocessing prepares the data in a format that can be easily analyzed by the generative AI model.

[0170] Step 3:

[0171] The server inputs pre-processed text data into a generative AI model. The generative AI model uses natural language processing techniques to analyze the dialogue and determine whether or not harassment has occurred. The result is output as a score indicating the likelihood of harassment. This score quantifies the risk of harassment.

[0172] Step 4:

[0173] Based on the assessment results, the server extracts specific details if harassment is detected. Using this extracted information, it creates a scenario for generating visual information. This scenario serves as a blueprint for determining what kind of visual information to generate.

[0174] Step 5:

[0175] The server uses video editing software to generate visual information that reflects the extracted harassment content. The generated visual information is designed to allow users to visually review their own words and actions. The visual information is sent to the user's device.

[0176] Step 6:

[0177] Users view visual information sent to their devices. This visual information includes specific examples of harassment and suggestions for improvement, allowing users to review their own behavior based on it.

[0178] (Application Example 3)

[0179] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0180] Harassment remains a serious problem in modern workplaces and educational institutions. Traditional methods are often insufficient in preventing harassment from occurring in the first place, and responses after it happens are frequently inadequate. In particular, the lack of real-time warnings and concrete improvement measures makes effective improvement difficult.

[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[0182] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the determination means; and a presentation means that issues a warning to the user in real time and presents visual information including specific advice for improvement. This enables the prevention of harassment and a rapid response after it occurs.

[0183] "Dialogue history" refers to a record of conversations between users, which is saved in audio or text format.

[0184] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not its content constitutes harassment.

[0185] "Generation means" refers to a device or program that has the function of automatically creating visual information based on the dialogue history that has been determined to be harassment by the determination means.

[0186] "Presentation means" refers to a device or program that has the function of displaying generated visual information to the user and issuing warnings in real time.

[0187] "Visual information" refers to visual content such as videos and images presented to users, including details of harassment and measures to improve it.

[0188] "Real-time" refers to the instantaneous processing and response that occurs at the very moment the user's interaction is taking place.

[0189] "Advice" refers to information that suggests specific actions or behavioral changes that users should take to improve their behavior and address harassment.

[0190] The system for implementing this invention mainly consists of a server and a user terminal. The server receives dialogue history as input and is equipped with a determination means for determining whether or not the dialogue history constitutes harassment. This determination means uses a generative AI model to analyze the input text data and evaluate the likelihood of harassment.

[0191] The user's device uses a speech recognition API (e.g., Google Speech-to-Text) to convert audio data into text data. This text data is sent to a server and analyzed by a judgment mechanism. If harassment is detected as a result of the judgment, the server uses a generation mechanism to generate visual information that reflects the content of the harassment. This visual information is created using video generation software (e.g., Adobe Premiere Pro).

[0192] The generated visual information is sent to the user's device and displayed to the user in real time via a presentation tool. This allows the user to recognize if their words or actions may constitute harassment and receive specific advice for improvement.

[0193] For example, workplace conversations may be recorded, and a warning message such as "That way of speaking is inappropriate" may appear. In this case, the video may also include advice such as, "This way of speaking may hurt the other person. Next time, try rephrasing it like this."

[0194] An example of a prompt for a generative AI model is: "Analyze the following conversation and determine if it may be harassment. Conversation: 'Your way of doing things is terrible. Do it properly.'"

[0195] The flow of the specific processing in Application Example 3 will be explained using Figure 16.

[0196] Step 1:

[0197] The user's device accepts voice input. When the user speaks, the device's microphone records the audio. The recorded audio data is converted into text data using a speech recognition API (e.g., Google Speech-to-Text). This converted text data becomes the input for the next process.

[0198] Step 2:

[0199] The server receives text data sent from the user's terminal. The server analyzes this text data using a generative AI model to determine the possibility of harassment. In this analysis, prompt sentences are input into the generative AI model to determine whether or not harassment is present. The output of this step is the result of the harassment determination.

[0200] Step 3:

[0201] Based on the assessment results, the server generates visual information using a generation mechanism if harassment is detected. This visual information is created using video generation software (e.g., Adobe Premiere Pro). The generated visual information includes specific improvement measures and advice. This visual information serves as input for the next process.

[0202] Step 4:

[0203] The server transmits the generated visual information to the user's terminal. The user's terminal displays this visual information to the user in real time using a presentation mechanism. Through the presented visual information, the user can recognize if their words or actions may constitute harassment and receive specific advice for improvement.

[0204] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0205] "Example of form 1"

[0206] One embodiment of the present invention provides a system incorporating an emotion engine. This system receives a user's dialogue history as input and determines whether or not that dialogue history constitutes harassment. This determination is made by an emotion engine that recognizes the user's emotions. Specifically, it extracts emotions from the user's dialogue history and determines whether those emotions are related to harassment. For example, if a user expresses anger or dissatisfaction, it is determined that this may constitute harassment.

[0207] "Example of form 2"

[0208] Furthermore, the system provides a mechanism to adjust harassment detection based on emotions recognized by the emotion engine. Specifically, it strengthens harassment detection when a user's emotions are strongly expressed or when certain emotions are repeatedly expressed. For example, if a user repeatedly expresses anger, it may be determined to be a strong form of harassment, and this result is provided as feedback to the user.

[0209] "Example of form 3"

[0210] Furthermore, the system provides a means to generate videos that reflect the user's emotions based on the history of conversations that have been identified as harassment. Specifically, it automatically generates videos that reflect both the content of the harassment and the user's emotions. For example, if a user expresses anger, a video reflecting that anger is generated and presented to the user. This allows the user to visually understand the relationship between their emotions and the harassment.

[0211] The following describes the processing flow for each example of the form.

[0212] "Example of form 1"

[0213] Step 1: The system receives the user's interaction history.

[0214] Step 2: The emotion engine extracts the user's emotions from the conversation history.

[0215] Step 3: Determine whether the emotions extracted by the emotion engine are related to harassment.

[0216] "Example of form 2"

[0217] Step 1: Receive the user's conversation history and the emotion recognition results from the emotion engine.

[0218] Step 2: Strengthen the criteria for determining harassment when a user's emotions are strongly expressed or when certain emotions are repeatedly expressed.

[0219] Step 3: Provide feedback to the user regarding the enhanced harassment assessment results.

[0220] "Example of form 3"

[0221] Step 1: Receive the conversation history and user sentiment that were identified as harassment.

[0222] Step 2: Combine the content of the harassment with the user's emotions and automatically generate a video that reflects both.

[0223] Step 3: Present the generated video to the user.

[0224] (Example 1)

[0225] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0226] In today's workplace, harassment is a serious problem, and early detection and countermeasures are essential. However, traditional methods often rely on subjective and inaccurate assessments of harassment. Furthermore, victims may face psychological resistance to reporting their experiences. This leads to challenges such as harassment being overlooked or responses being delayed.

[0227] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0228] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; an analysis means that analyzes the dialogue history using natural language processing technology; and an emotion recognition means that extracts emotions and determines whether or not the emotions are related to harassment. This makes it possible to objectively and quickly determine whether or not harassment has occurred and to respond promptly.

[0229] "Dialogue history" is a record of linguistic information exchanged between users during communication with others.

[0230] "Information processing means" refers to a device or program that has the function of analyzing input data and making a decision based on specific conditions.

[0231] "Generation means" refers to a device or program that has the function of creating and outputting new visual information based on input information.

[0232] "Natural language processing technology" is a technology that enables computers to understand and analyze human language.

[0233] "Analysis means" refers to a device or program that has the function of examining input data in detail and understanding its structure and meaning.

[0234] An "emotion recognition means" is a device or program that has the function of extracting emotions from input data and determining whether those emotions are related to specific conditions.

[0235] "Visual information" refers to information that can be recognized visually, such as images and videos.

[0236] As an embodiment for carrying out this invention, the automated harassment analysis system is configured as follows.

[0237] The user enters the conversation history using a terminal. This conversation history is a record of the linguistic information exchanged between the user and others. The entered conversation history is transmitted to the server via the internet.

[0238] The server analyzes the received dialogue history using information processing tools. These tools include Python's natural language processing libraries, NLTK and spaCy. This allows the server to analyze the dialogue history text, extract keywords, and understand the context.

[0239] Furthermore, the server uses emotion recognition to extract emotions from the conversation history. This emotion recognition uses emotion analysis tools such as Google Cloud's Natural Language API. The server then determines whether the extracted emotions are related to harassment.

[0240] For example, if a user enters a conversation history such as "My boss insults me every day," the server analyzes this history and extracts the keyword "insult." Next, the emotion engine identifies the emotion of "anger," and based on this information, it determines that there is a high probability of harassment.

[0241] An example of a prompt to input into a generative AI model would be: "Please determine if the following conversation history constitutes harassment: 'My boss insults me every day.'" Using this prompt, the generative AI model analyzes the content of the conversation history and determines whether or not harassment occurred.

[0242] The flow of the specific processing in Example 1 will be explained using Figure 17.

[0243] Step 1:

[0244] The user enters the conversation history using a terminal. The entered conversation history is in text format and includes linguistic information exchanged between the user and others. This data is transmitted to the server via the internet.

[0245] Step 2:

[0246] The server passes the received dialogue history to an information processing system. Here, the server uses Python's natural language processing libraries, NLTK and spaCy, to parse the dialogue history text. Specifically, the server extracts keywords from the text and performs data processing to understand the context. As a result of this analysis, important keywords and phrases contained in the dialogue history are output.

[0247] Step 3:

[0248] The server passes the analyzed data to the emotion recognition system. Here, the server uses emotion analysis tools such as Google Cloud's Natural Language API to extract emotions from the conversation history. Specifically, the server identifies emotions from words and phrases in the text and determines whether those emotions are related to harassment. As a result of this process, the extracted emotions and their relevance are output.

[0249] Step 4:

[0250] The server integrates the results obtained from the information processing and emotion recognition means to ultimately determine whether the dialogue history constitutes harassment. Specifically, the server evaluates the possibility of harassment based on the extracted keywords and emotions. As a result of this determination, it outputs whether or not the dialogue history constitutes harassment.

[0251] Step 5:

[0252] The server returns the judgment result to the user. The user's terminal is notified of the judgment result and displayed as a detailed report. Specifically, the user can check whether the conversation history constitutes harassment and take appropriate action if necessary.

[0253] (Application Example 1)

[0254] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server," and the smart device 14 will be referred to as a "terminal."

[0255] Workplace harassment is a problem that negatively impacts employees' mental health and workplace productivity. However, early detection of signs of harassment and implementation of appropriate countermeasures are difficult. In particular, there is a need for a system that can detect and warn about harassment in real time.

[0256] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0257] In this invention, the server includes a determination means that receives dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs information automatically generated based on the dialogue history determined to be harassment by the determination means; and a processing means that converts the user's voice data into text and analyzes it using natural language processing technology. This enables early detection of harassment in the workplace and prompt warnings.

[0258] "Dialogue history" refers to a record of conversations between users, which is saved as audio or text data.

[0259] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not its content constitutes harassment.

[0260] "Generation means" refers to a device or program that has the function of automatically creating and outputting relevant information and warnings based on the dialogue history that has been determined to be harassment by the determination means.

[0261] "Processing means" refers to a device or program that has the function of converting the user's voice data into text and analyzing that text using natural language processing technology.

[0262] "Analysis tool" refers to a device or program that has the function of extracting emotions from text data and evaluating the possibility of harassment.

[0263] "Notification means" refers to a device or program that has the function of issuing a warning to a user in the event of potential harassment.

[0264] A "portable information terminal" is an electronic device that is portable and capable of processing information, such as a smartphone or tablet.

[0265] The system that realizes this invention mainly consists of a server and a user's mobile information terminal. The server receives dialogue history as input and is equipped with a determination means for determining whether or not the dialogue history constitutes harassment. The determination means analyzes the dialogue history using natural language processing technology and evaluates the possibility of harassment. Specifically, it converts the user's voice data into text and performs analysis using a natural language processing library (e.g., spaCy, NLTK).

[0266] Furthermore, the server is equipped with analytical tools to extract emotions from text data and determine the possibility of harassment. Using emotion analysis libraries (e.g., TextBlob, VADER), it evaluates the user's emotions, and if anger or dissatisfaction is detected, it determines that there is a possibility of harassment.

[0267] The user's mobile device is equipped with a notification system to receive notifications from the server. If harassment is suspected, a warning will be displayed on the device to alert the user. This enables early detection and rapid response to harassment in the workplace.

[0268] As a concrete example, if a supervisor makes an inappropriate remark to a subordinate during a workplace conversation, the server analyzes the remark and, if it determines that it may constitute harassment, displays a warning on the subordinate's mobile device. An example of a prompt to be input to the generating AI model would be, "Analyze the content of this conversation and determine if it may constitute harassment."

[0269] The flow of a specific process in Application Example 1 will be explained using Figure 18.

[0270] Step 1:

[0271] The server receives audio data from the user's mobile device. This audio data is a recording of a conversation at work. The server uses speech recognition software to convert this audio data into text data.

[0272] Step 2:

[0273] The server parses the converted text data using a natural language processing library (e.g., spaCy, NLTK). Through this parsing, grammatical structures and keywords in the text are extracted, and basic data for understanding the content of the dialogue is generated.

[0274] Step 3:

[0275] The server extracts emotions from the parsed text data using a sentiment analysis library (e.g., TextBlob, VADER). In this step, emotional expressions contained in the text are evaluated, and it is determined whether negative emotions such as anger and discontent are included.

[0276] Step 4:

[0277] The server determines the possibility of harassment based on the result of the sentiment analysis. Specifically, when the negative emotion exceeds a predetermined threshold, it is determined that there is a possibility of harassment. This determination result can be further improved in accuracy by using a generative AI model.

[0278] Step 5:

[0279] When the server determines that there is a possibility of harassment, it transmits a warning to the user's portable information terminal. The terminal notifies the user of this warning and calls the user's attention. This allows the user to recognize the risk of harassment in real time.

[0280] (Example 2)

[0281] Next, Example 2 of the second embodiment will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart device 14 is referred to as a "terminal".

[0282] In today's workplace, harassment is a serious problem, and it is essential to properly detect it and take countermeasures. However, traditional methods have challenges in making accurate judgments because the determination of harassment is subjective and does not take into account changes in emotions.

[0283] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0284] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the information processing means; and an adjustment means that analyzes emotions and adjusts the harassment determination based on the analysis results. This makes it possible to make an objective and emotionally conscious determination of harassment.

[0285] "Dialogue history" refers to a record of conversations that users have had, and is data saved in audio or text format.

[0286] "Information processing means" refers to technical means for analyzing dialogue history and determining whether or not harassment occurred.

[0287] "Generation means" refers to a technical means for automatically creating and outputting visual information based on dialogue history that has been determined to be harassment.

[0288] "Adjustment measures" refer to technical means for analyzing user emotions and modifying or strengthening harassment judgments based on the results.

[0289] "Visual information" refers to visually presented information such as videos and images that reflect the content of the harassment.

[0290] "Emotion" refers to the user's psychological state and includes changes in feelings and emotions inferred from voice and text.

[0291] In an embodiment of this invention, a server, a terminal, and a user cooperate to operate the system. First, the user uses a terminal to record everyday conversations. The terminal is equipped with speech recognition software, which converts the speech data into text data. The converted text data is sent to the server as a conversation history.

[0292] The server analyzes the received dialogue history using information processing tools. This analysis employs natural language processing technology to detect specific keywords and phrases, thereby determining whether harassment has occurred. Furthermore, the server uses an emotion engine to analyze the user's emotions and adjusts the harassment determination based on the results. For example, if a user repeatedly expresses "anger," the server takes that emotion into account and strengthens its determination.

[0293] If harassment is detected, the server automatically generates visual information using a generation mechanism. This visual information is a video created using a generation AI model to represent a scenario corresponding to the content of the conversation. The generated video is sent to the user's device and presented to the user. By watching this video, the user can deepen their understanding of the nature of the harassment and its impact.

[0294] As a concrete example, let's look at a prompt message to input into a generation AI model: "If a user repeatedly expresses anger in conversations with their boss, please generate a video that points out the possibility of workplace harassment based on that conversation history." Using this prompt message, the server can generate an appropriate video and provide it to the user.

[0295] The flow of the specific processing in Example 2 will be explained using Figure 19.

[0296] Step 1:

[0297] A user records a conversation using a terminal. The terminal converts audio data into text data using speech recognition software. The input for this step is audio data, and the output is text data. As a specific operation, the user records a conversation at the workplace and converts the recording into text.

[0298] Step 2:

[0299] A server receives text data transmitted from a terminal. The server analyzes the text data using information processing means and determines whether harassment is present. The input for this step is text data, and the output is a harassment determination result. As a specific operation, the server utilizes natural language processing technology to detect specific keywords and phrases.

[0300] Step 3:

[0301] The server analyzes the user's emotion from the text data using a sentiment engine. The harassment determination is adjusted based on this analysis result. The input for this step is text data, and the output is an adjusted determination result. As a specific operation, the server estimates emotion from the tone of the audio and the content of the text, and strengthens or corrects the determination.

[0302] Step 4:

[0303] When the server determines that harassment has occurred, it automatically generates visual information using generation means. The input for this step is the adjusted determination result, and the output is visual information (video). As a specific operation, the server utilizes a generative AI model to convert a scenario corresponding to the content of the conversation into a video.

[0304] Step 5:

[0305] The server sends the generated visual information to the user's device and presents it to the user. The input in this step is visual information, and the output is the presentation to the user. Specifically, the user watches a video on their device to deepen their understanding of the content and impact of harassment.

[0306] (Application Example 2)

[0307] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0308] In modern workplaces and educational institutions, harassment is a serious problem, and early detection and countermeasures are essential. However, traditional methods can lead to delays in detecting harassment or failure to implement appropriate measures. In particular, there is a challenge in determining harassment while considering emotional changes.

[0309] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0310] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs a video that is automatically generated based on the dialogue history determined to be harassment by the determination means; and an adjustment means that recognizes the user's emotions and adjusts the harassment determination based on those emotions. This enables early detection of harassment and the implementation of appropriate countermeasures.

[0311] "Dialogue history" refers to a record of conversations between users, which is saved as text data.

[0312] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not it constitutes harassment.

[0313] "Generation means" refers to a device or program that has the function of automatically creating and outputting a video that points out the problematic aspects based on the history of a conversation that has been determined to be harassment.

[0314] "Adjustment measures" refer to devices or programs that recognize the user's emotions and have the function of strengthening or mitigating the determination of harassment based on those emotions.

[0315] "Emotions" refer to the psychological states and reactions that users exhibit during a conversation, including states such as anger, sadness, and joy.

[0316] The system for carrying out this invention includes a server and a user terminal. The server is equipped with a determination means for receiving and analyzing dialogue history. The determination means uses natural language processing technology to determine whether the dialogue history constitutes harassment. Specifically, it uses the Google Cloud Natural Language API to analyze text data and extract emotions and intentions.

[0317] The user's device, such as a smartphone or computer, is responsible for sending conversation history to the server. It also receives generated video from the server and presents it to the user. The generation method uses the Adobe Premiere Pro API to automatically generate video based on conversation history identified as harassment. This video includes content that points out the problem and prompts the user to take action.

[0318] Furthermore, the server is equipped with adjustment mechanisms to recognize user emotions in real time. The emotion engine enhances harassment detection when user emotions are strongly expressed or when certain emotions are repeatedly expressed. This enables more accurate detection.

[0319] For example, if workplace conversations are recorded and the user repeatedly expresses "anger," the server analyzes the conversation and determines that it may constitute workplace harassment. It then generates a video highlighting the harassment issues and notifies the user.

[0320] Examples of prompts for a generative AI model include the following:

[0321] "Analyze the following conversation history and determine if harassment is likely. Strengthen the assessment if the user's emotions are strongly expressed. If harassment is detected, generate a video highlighting the problem."

[0322] The flow of a specific process in Application Example 2 will be explained using Figure 20.

[0323] Step 1:

[0324] The user's terminal sends the conversation history as text data to the server. The input is the user's conversation history, and the output is the data transmission to the server. In this step, the user's terminal collects the conversation history and sends it to the server over the network.

[0325] Step 2:

[0326] The server analyzes the received dialogue history using the Google Cloud Natural Language API. The input is the text data of the dialogue history, and the output is the result of the sentiment and intent analysis. In this step, the server uses natural language processing techniques to extract sentiment and intent from the dialogue history.

[0327] Step 3:

[0328] The server determines the possibility of harassment based on the analysis results. The input is the analysis results of emotions and intentions, and the output is the harassment determination result. In this step, the server uses a determination means to evaluate the analysis results and determine whether or not harassment occurred.

[0329] Step 4:

[0330] If harassment is detected, the server generates a video highlighting the problem using the Adobe Premiere Pro API. The input is the harassment detection result and the dialogue history, and the output is the generated video. In this step, the server uses a generation mechanism to automatically create a video based on the dialogue history.

[0331] Step 5:

[0332] The server sends the generated video to the user's terminal. The input is the generated video, and the output is the transmission of the video to the user's terminal. In this step, the server transmits the video to the user's terminal over the network.

[0333] Step 6:

[0334] The user's device displays the received video to the user. The input is the video transmitted from the server, and the output is the video displayed to the user. In this step, the user's device plays the video and provides information to the user visually.

[0335] (Example 3)

[0336] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0337] In modern workplaces and society, harassment remains a serious problem, and preventing its occurrence is essential. However, it is difficult to determine what kind of words and actions constitute harassment in individual conversations and communications, and when harassment does occur, concrete corrective measures are required. Furthermore, there is a lack of visual means to understand the emotions associated with harassment, so these challenges need to be addressed.

[0338] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[0339] In this invention, the server includes: information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; data generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the information processing means; and data generation means that analyzes the emotions contained in the dialogue history and generates a video that reflects those emotions. This makes it possible for users to recognize the possibility that their words and actions constitute harassment and to visually understand concrete measures to improve their behavior.

[0340] "Dialogue history" refers to a record of conversations and communications between users, and is data saved in text format.

[0341] "Information processing means" refers to a device or program that has the function of analyzing input dialogue history and determining whether its content constitutes harassment.

[0342] "Data generation means" refers to a device or program that has the function of automatically creating images and other visual content based on the determination results of information processing means.

[0343] "Emotional analysis" is the process of identifying the user's emotions contained in the conversation history and evaluating the type and intensity of those emotions.

[0344] "Video" refers to content that includes visual information and is presented to users in the form of videos, animations, and other visual media.

[0345] A description of embodiments for carrying out this invention will be given.

[0346] The server receives the conversation history entered by the user. This conversation history is a record of the conversations and communications the user has had, and is provided in text format. The server analyzes this conversation history using natural language processing technology. Specifically, a common cloud-based natural language processing API can be used as the natural language processing software. This allows the server to understand the content of the conversation and detect signs of harassment.

[0347] Next, the server uses information processing means to determine whether the conversation history constitutes harassment. Based on this determination, the server uses data generation means to generate a video that reflects the content of the harassment. Video editing software can be used for this video generation.

[0348] Furthermore, the server uses sentiment analysis tools to analyze the user's emotions contained in the conversation history. Based on the results of the sentiment analysis, the server generates a video that reflects the user's emotions. This video is intended to help the user visually understand the relationship between their emotions and harassment.

[0349] As a concrete example, consider a scenario where a user enters a conversation history stating, "I used harsh words towards my subordinate during yesterday's meeting." The server analyzes this history and determines whether the harsh language constitutes harassment. If it does, the server generates a video based on the content that reflects the user's anger and presents it to the user.

[0350] An example of a prompt to be input to the generation AI model is, "Based on this dialogue history, determine whether harassment occurred and, if necessary, generate a video that reflects the emotions." The flow of the specific processing in Example 3 will be explained using Figure 21.

[0351] Step 1:

[0352] The user enters their conversation history. The user enters their conversation history into the system interface. The entered data is sent to the server in text format.

[0353] Step 2:

[0354] The server receives the dialogue history and performs natural language processing. The server analyzes the received text data using natural language processing software. This analysis helps understand the content and context of the dialogue and prepares it to detect signs of harassment. The input is text data, and the output is the analysis result.

[0355] Step 3:

[0356] The server determines whether harassment has occurred. Based on the results of natural language processing, the server determines whether harassment is present in the dialogue history. Specifically, it evaluates whether aggressive words or inappropriate expressions are included. The input is the analysis result, and the output is the determination result regarding the presence or absence of harassment.

[0357] Step 4:

[0358] If the server detects harassment, it prepares data for video generation. When harassment is detected, the server creates a script or template for video generation based on the details. The input is the detection result, and the output is the data for video generation.

[0359] Step 5:

[0360] The server analyzes the user's emotions. The server uses an emotion analysis tool to identify the user's emotions as they appear in the conversation history. The input is the conversation history, and the output is the result of the emotion analysis.

[0361] Step 6:

[0362] The server generates a video that reflects emotions. The server generates a video that reflects the content of the harassment and the user's emotions. Specifically, it uses video editing software to create visually represented content. The input is data for video generation and the results of emotion analysis, and the output is the generated video.

[0363] Step 7:

[0364] The server generates a video and presents it to the user. The server sends the generated video to the user, making it viewable. The input is the generated video, and the output is its presentation to the user.

[0365] (Application Example 3)

[0366] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0367] Harassment in the workplace and online communication is a problem that negatively impacts individual mental health and workplace productivity. However, there are insufficient systems to detect signs of harassment early and provide appropriate feedback. Furthermore, there is a lack of means for users to visually understand the connection between their own feelings and harassment and to take concrete actions toward improvement.

[0368] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[0369] In this invention, the server includes means for receiving a conversation history as input and determining whether the conversation history constitutes harassment; means for outputting a video automatically generated based on the conversation history determined to be harassment by the harassment determination means; and means for analyzing the user's emotions and generating a video that reflects those emotions. This allows the user to receive specific feedback to improve their communication style and prevent harassment from occurring.

[0370] "Dialogue history" refers to a record of conversations between users, which is saved in text format.

[0371] A "harassment determination tool" is a device that analyzes the input conversation history and determines whether or not its content constitutes harassment.

[0372] A "video generation method" is a system that automatically creates videos to provide visual feedback based on conversation history and user emotions that have been identified as harassment.

[0373] "User emotions" refers to the user's psychological state and emotional expression in their dialogue history, and this is analyzed and reflected in the video generation process.

[0374] A "generative AI model" refers to an algorithm or system that uses artificial intelligence technology to generate new content from input data.

[0375] "Feedback" refers to information and advice provided to users that helps improve communication styles.

[0376] The system for implementing this invention consists of a server and a user terminal. The server receives the conversation history transmitted from the user and analyzes its content using a harassment detection means. Specifically, it processes the text data using the natural language processing library spaCy and evaluates the possibility of harassment. Based on the evaluation results, it generates a video that reflects the content of the harassment using OpenAI's GPT generative AI model.

[0377] The video generation process includes analyzing the user's emotions. These emotions are extracted from the conversation history and visually represented by the video generation tool. Video editing libraries such as MoviePy are used for video generation. The generated video is sent to the user's device and presented to them. This allows the user to receive specific feedback to improve their communication style.

[0378] As a concrete example, consider email exchanges in the workplace. If a user sends an email saying, "This project isn't progressing at all, it's all your fault," the server analyzes this text and detects the possibility of harassment. An example of a prompt would be, "Evaluate whether the following text constitutes harassment: 'This project isn't progressing at all, it's all your fault'," which would be input into the GPT model. The generated video provides visual feedback to help the user reflect on their own behavior and make improvements.

[0379] The flow of the specific processing in Application Example 3 will be explained using Figure 22.

[0380] Step 1:

[0381] The user sends the conversation history to the server using their device. The input is the user's conversation history, sent to the server in text format. The server receives this data and prepares for the next processing step.

[0382] Step 2:

[0383] The server analyzes the received dialogue history using the natural language processing library spaCy. The input is the text data of the dialogue history, and the output is the analysis results indicating the possibility of harassment. The server detects signs of harassment by tokenizing the text and analyzing its grammatical structure.

[0384] Step 3:

[0385] Based on the analysis results, the server inputs prompt text into the OpenAI GPT, a generative AI model. The input is prompt text that reflects the analysis results, and the output is an evaluation of whether or not harassment occurred. Specifically, the server generates prompt text and sends it to the GPT model.

[0386] Step 4:

[0387] The server receives evaluation results from the GPT model and, if it determines that harassment has occurred, generates a video using a video generation method. The input is the evaluation results and user sentiment data, and the output is the generated video. The server uses a video editing library such as MoviePy to create a video that reflects the user's sentiment.

[0388] Step 5:

[0389] The server sends the generated video to the user's terminal and presents it to the user. The input is the generated video, and the output is the video playback on the user's terminal. The user watches this video and receives feedback to improve their communication style.

[0390] (Other examples)

[0391] Next, other embodiments will be described. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0392] Workplace harassment is a serious problem that negatively impacts employees' mental health and the work environment. However, detecting and addressing harassment is difficult, and identifying harassment from dialogue history is particularly challenging. Traditional methods require humans to manually review the content of conversations, which is time-consuming and labor-intensive, and can lack accuracy due to reliance on subjective judgment. Therefore, there is a need for a system that can automatically analyze dialogue history and identify harassment.

[0393] The identification process performed by the identification processing unit 290 of the data processing device 12 in other embodiments is realized by the following means.

[0394] In this invention, the server includes: an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that generates a prompt to instruct the generation of visual information based on the dialogue history determined to be harassment by the information processing means, and automatically generates visual information based on the prompt using a generation AI model; and an analysis means that analyzes the dialogue history using natural language processing technology and inputs the analysis results into the generation AI model. This makes it possible to automatically detect harassment from dialogue history and represent it visually.

[0395] "Dialogue history" refers to a record of conversations between users, and is data saved in text format.

[0396] "Information processing means" refers to software or hardware functions that analyze input dialogue history and determine whether or not harassment occurred.

[0397] A "generative AI model" refers to a model that uses artificial intelligence technology to generate new information or content based on input data.

[0398] A "prompt" refers to a set of instructions entered into a generative AI model to tell it to perform a specific action.

[0399] "Visual information" refers to images and graphics generated to visually represent the content of harassment.

[0400] "Natural language processing technology" refers to the technology that enables computers to understand and analyze human language, and is used for analyzing text data and extracting information.

[0401] The following describes "modes for carrying out the invention."

[0402] ---

[0403] This invention is a system that analyzes dialogue history, automatically detects harassment, and visually represents it. The system mainly consists of three elements: a server, a terminal, and a user.

[0404] The server receives the dialogue history sent by the user. The dialogue history is data stored in text format and refers to a record of conversations between users. The server analyzes the received dialogue history using Python's natural language processing libraries, NLTK and spaCy. The purpose of the analysis is to detect specific keywords and phrases contained in the dialogue history and to assess the possibility of harassment.

[0405] Based on the analysis results, the server generates prompt messages for input into the AI ​​model. These prompt messages include instructions to generate visual information. A specific example of such a prompt message might be, "Based on this dialogue history, please generate an image that visually represents the content of the harassment."

[0406] The generated prompt text is input into generative AI models such as OpenAI's GPT-3® or Google's BERT. The server hosts the generative AI model and automatically generates visual information based on the prompt text. This visual information reflects the content of the harassment and is provided to the user.

[0407] The terminal receives visual information sent from the server and displays it to the user. The user can review the displayed visual information and provide feedback as needed. This feedback helps improve the system.

[0408] This system allows users to quickly obtain a visual representation of harassment based on their conversation history. This enables early detection and response to workplace harassment.

[0409] The flow of specific processing in other embodiments will be explained using Figure 23.

[0410] Step 1:

[0411] The server receives the dialogue history sent by the user. The input is the dialogue history in text format. The server saves this dialogue history to a database in preparation for subsequent processing. The output is the saved dialogue history data.

[0412] Step 2:

[0413] The server analyzes the saved dialogue history using natural language processing techniques. The input is the dialogue history data saved in step 1. The server uses Python's NLTK and spaCy to extract keywords and phrases from the dialogue history and detect patterns that indicate potential harassment. The output is a list of keywords and a harassment likelihood assessment as a result of the analysis.

[0414] Step 3:

[0415] The server generates prompt text for input to the AI ​​model based on the analysis results. The input is the analysis results obtained in step 2. Using these results, the server creates the prompt text, "Based on this dialogue history, please generate an image that visually represents the content of the harassment." The output is the generated prompt text.

[0416] Step 4:

[0417] The server inputs the generated prompt text into the generative AI model. The input is the prompt text generated in step 3. The server uses generative AI models such as OpenAI's GPT-3 or Google's BERT to automatically generate visual information based on the prompt text. The output is the generated visual information.

[0418] Step 5:

[0419] The server sends the generated visual information to the terminal. The input is the visual information generated in step 4. The server sends this information to the user's terminal for the user to review. The output is the visual information sent to the terminal.

[0420] Step 6:

[0421] The terminal receives visual information sent from the server and displays it to the user. The input is the visual information sent in step 5. The terminal displays this information on the screen so that the user can visually confirm it. The output is the visual information displayed to the user.

[0422] Step 7:

[0423] The user reviews the displayed visual information and provides feedback as needed. The input is the visual information displayed in step 6. The user evaluates the content of the visual information and sends feedback to the server, including suggestions for improvement and comments. The output is the feedback sent to the server.

[0424] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0425] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0426] Other examples of generative AI include Gemini® (registered trademark) (Internet search). <url: https: gemini.google.com ?hl="ja">) are some examples.

[0427] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0428] [Second Embodiment]

[0429] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0430] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0431] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0432] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0433] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0434] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0435] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0436] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0437] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0438] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0439] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0440] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.

[0441] "Example of form 1"

[0442] One embodiment of the present invention provides an automated harassment analysis system. This system receives a conversation history from a user as input and has a harassment determination means that determines whether or not the conversation history constitutes harassment. The harassment determination means analyzes the conversation history using natural language processing technology and determines whether or not harassment such as power harassment or sexual harassment has occurred.

[0443] "Example of form 2"

[0444] Furthermore, the system of the present invention includes a video generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the harassment determination means. The video generation means generates a video that reflects the content of the dialogue history determined to be harassment and presents it to the user. For example, if the dialogue history determined to be harassment is related to power harassment, it generates a video that points out the problems of power harassment.

[0445] "Example of form 3"

[0446] The system of the present invention can prevent harassment from occurring and provide concrete material for reflection and improvement in cases of harassment that have occurred. Specifically, by having the user input their conversation history into the system, the presence or absence of harassment is determined, and if harassment is found, a video reflecting the content is generated and presented. As a result, the user can recognize that their words and actions may constitute harassment and take concrete actions to improve them.

[0447] The following describes the processing flow for each example of the form.

[0448] "Example of form 1"

[0449] Step 1: The system receives the user's interaction history.

[0450] Step 2: The system's harassment detection mechanism analyzes the conversation history. This analysis uses natural language processing technology to determine whether or not harassment, such as power harassment or sexual harassment, has occurred.

[0451] "Example of form 2"

[0452] Step 1: The system receives the conversation history that has been determined to be harassment by the harassment detection method.

[0453] Step 2: The system's video generation mechanism automatically generates a video based on the conversation history that was identified as harassment. The video content reflects the content of the harassment.

[0454] Step 3: Present the generated video to the user.

[0455] "Example of form 3"

[0456] Step 1: The user enters their conversation history into the system.

[0457] Step 2: The system determines whether harassment occurred, and if so, generates a video reflecting the details of the harassment.

[0458] Step 3: Present the generated video to the user, allowing them to recognize that their words and actions may constitute harassment.

[0459] Step 4: Based on the presented video, users take specific actions to improve the harassment situation.

[0460] (Example 1)

[0461] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0462] In today's workplace, harassment is a serious problem, and its detection and response are crucial. However, traditional methods have made it difficult to efficiently analyze dialogue history and accurately determine the presence or absence of harassment. Furthermore, there has been a lack of visual representations of harassment, making it difficult to understand and share the problem.

[0463] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0464] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; an analysis means that analyzes the dialogue history using natural language processing technology; and an evaluation means that evaluates the analysis results using a generative AI model and determines whether or not harassment has occurred. This enables efficient analysis of dialogue history and accurate determination of harassment.

[0465] "Dialogue history" refers to a record of conversations and messages exchanged between users, and is data saved in text format.

[0466] "Information processing means" refers to a device or program that has the function of analyzing input data and making decisions based on specific conditions.

[0467] "Generation means" refers to a device or program for automatically creating new visual information or content based on input information.

[0468] "Natural language processing technology" refers to technologies for understanding and analyzing human language using computers, and includes text tokenization and grammatical analysis.

[0469] "Analysis means" refers to a device or program for analyzing input data in detail and understanding its structure and meaning.

[0470] A "generative AI model" is a model that uses artificial intelligence technology to learn from data and is trained to perform a specific task.

[0471] "Evaluation means" refers to a device or program that makes a judgment based on analyzed data according to specific criteria.

[0472] To implement this invention, the user must first input their conversation history. The user uses their device to input the content of workplace conversations and messages in text format. The inputted conversation history is then sent from the device to the server.

[0473] The server uses Python's natural language processing libraries, NLTK and spaCy, to analyze the received dialogue history. Using this software, the server tokenizes the text data, tags it with parts of speech, and analyzes its grammatical structure.

[0474] Next, the server evaluates the analysis results using a generative AI model. This model is built using deep learning frameworks such as TensorFlow and PyTorch and is pre-trained on a dataset related to harassment. The model detects elements of power harassment and sexual harassment in the dialogue history and determines whether or not they exist.

[0475] For example, if a user enters a conversation history such as "My boss forces me to work unreasonable overtime every day," the server analyzes this data and uses a generative AI model to determine that "it contains elements of power harassment." This result is sent back to the terminal and displayed to the user.

[0476] An example of a prompt message is, "Please determine if this conversation history contains elements of harassment." By using this prompt message, the server can efficiently analyze the conversation history and accurately determine whether or not harassment has occurred.

[0477] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0478] Step 1:

[0479] Users input conversation history in text format using a terminal. This input data includes the content of workplace conversations and messages. The entered conversation history is sent from the terminal to the server.

[0480] Step 2:

[0481] The server uses natural language processing libraries such as NLTK and spaCy to analyze the dialogue history received from the terminal. The server tokenizes the input text data, tags it with parts of speech, and analyzes its grammatical structure. This analysis extracts structural information from the text.

[0482] Step 3:

[0483] The server inputs the analyzed data into a generating AI model. This model is built using TensorFlow and PyTorch and has been pre-trained on a dataset related to harassment. Based on the input data, the model detects elements of power harassment and sexual harassment and determines whether or not they exist. This determination result is then output.

[0484] Step 4:

[0485] The server returns the judgment result obtained from the generated AI model to the terminal. The judgment result indicates whether or not the conversation history contains elements of harassment.

[0486] Step 5:

[0487] The terminal displays the judgment results received from the server to the user. For example, a message such as "This conversation contains elements of power harassment" might be displayed. This allows the user to understand the content of the conversation history and take necessary measures.

[0488] (Application Example 1)

[0489] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0490] Workplace harassment is a serious problem that negatively impacts employees' mental health and workplace productivity. However, early detection and appropriate response to signs of harassment are difficult. Traditional methods often rely on victims reporting the issue themselves, and problems may go undetected until they become severe. An effective system is needed to improve this situation and prevent harassment proactively.

[0491] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0492] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a conversion means that converts speech into text; and an analysis means that analyzes the text converted by the conversion means and detects signs of harassment. This makes it possible to detect signs of harassment in the workplace environment in real time and notify administrators.

[0493] "Dialogue history" refers to a record of conversations between users, which is saved as audio or text data.

[0494] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not it constitutes harassment.

[0495] "Generation means" refers to a device or program that has the function of automatically creating relevant information and content based on the dialogue history that has been determined to be harassment by the determination means.

[0496] "Conversion means" refers to a device or program that has the function of converting audio data into text data.

[0497] "Analysis means" refers to a device or program that has the function of analyzing text data and detecting signs of harassment.

[0498] A "notification means" is a device or program that has the function of sending warnings or information to administrators or relevant parties based on signs of harassment detected by an analysis means.

[0499] The system for implementing this invention mainly consists of a server and a user terminal. The server receives dialogue history as input and is equipped with a determination means for determining whether or not the dialogue history constitutes harassment. This determination means analyzes the dialogue history using natural language processing technology to determine whether or not harassment has occurred. Specifically, it uses the Google Cloud Natural Language API to analyze text data.

[0500] The user terminal is equipped with a conversion mechanism for converting speech to text. This conversion mechanism uses the Google Cloud Speech-to-Text API to convert speech data into text data. The converted text data is sent to a server and analyzed by a determination mechanism.

[0501] If the server detects signs of harassment as a result of its analysis, it will send a warning to the administrator using a notification system. This notification system uses Firebase Cloud Messaging to provide real-time notifications.

[0502] For example, if the statement "You're always useless" is made during a conversation at work, this statement may be judged as potentially being harassment, and a notification will be sent to the manager. An example of a prompt to the generating AI model would be, "Please determine whether the following conversation constitutes harassment: 'You're always useless.'"

[0503] In this way, the system can detect signs of harassment in the workplace in real time, enabling a rapid response.

[0504] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0505] Step 1:

[0506] The user terminal captures workplace conversations as audio data. This audio data is collected in real time through the user terminal's microphone.

[0507] Step 2:

[0508] The user's device converts the acquired audio data into text data using the Google Cloud Speech-to-Text API. This conversion process analyzes the audio signal and generates the corresponding text. The input is audio data, and the output is text data.

[0509] Step 3:

[0510] The user's terminal sends the converted text data to the server. The server analyzes the received text data using the Google Cloud Natural Language API. The analysis evaluates linguistic features within the text and detects signs of harassment. The input is text data, and the output is a determination of whether or not harassment occurred.

[0511] Step 4:

[0512] If the server detects signs of harassment based on its analysis, it will send a notification to the administrator using Firebase Cloud Messaging. The notification will include detailed information such as the nature and time of the harassment. The input is the analysis result, and the output is the notification to the administrator.

[0513] Step 5:

[0514] The administrator will take necessary actions based on the received notification. This includes confirming with relevant parties and conducting further investigations. The administrator's actions will be based on the system output.

[0515] (Example 2)

[0516] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0517] In today's workplace, harassment is a serious problem, and early detection and appropriate response are essential. However, traditional methods have the drawback of subjective harassment assessments, making it difficult to implement appropriate education and countermeasures. Therefore, a system is needed that objectively assesses harassment and provides educational visual information based on that assessment.

[0518] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0519] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a visual information generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the information processing means; and a presentation means that presents the visual information generated by the visual information generation means to the user. This makes it possible to objectively determine harassment and provide educational visual information based on its content.

[0520] "Dialogue history" refers to a record of conversations between users, which is saved in text format.

[0521] "Information processing means" refers to a device or program that has the function of analyzing input data and making decisions based on specific conditions.

[0522] "Visual information generation means" refers to a device or program that has the function of automatically creating visual content based on analysis results.

[0523] "Presentation means" refers to a device or program for displaying or providing generated visual information to a user.

[0524] "Harassment" refers to inappropriate words or actions towards others in the workplace or other environments, and is an act that causes mental or physical distress.

[0525] A description of embodiments for carrying out this invention will be given.

[0526] The server receives the conversation history provided by the user. The user inputs the conversation history in text format via their terminal and sends it to the server. This conversation history is used as data for harassment assessment.

[0527] The server analyzes the dialogue history using a generative AI model. Specifically, it uses natural language processing techniques to determine whether the content of the dialogue constitutes harassment. In this process, a generative AI model with excellent natural language processing capabilities is used as the general-purpose model. The server inputs the prompt "Please determine whether this dialogue history constitutes harassment" into the generative AI model.

[0528] If harassment is detected, the server automatically generates visual information based on the detection result using a visual information generation system. This visual information includes educational content and aims to inform users about the problems and countermeasures against harassment. General video editing software may be used to generate the visual information.

[0529] For example, if a user enters a conversation history stating that "a superior excessively reprimanded a subordinate," the server will determine this to be harassment. Next, the server generates and presents visual information on the theme of "the impact of inappropriate behavior in the workplace and countermeasures." This visual information includes specific examples and countermeasures, allowing the user to deepen their understanding.

[0530] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0531] Step 1:

[0532] The user uses a terminal to input the conversation history in text format and sends it to the server. The input data is a record of the conversation between users and serves as the basis for determining whether harassment occurred. The server receives this conversation history and prepares for the next processing step.

[0533] Step 2:

[0534] The server inputs the received dialogue history into the generating AI model. Specifically, it uses the prompt "Determine whether this dialogue history constitutes harassment" to have the generating AI model analyze the dialogue history. The generating AI model uses natural language processing techniques to analyze the content of the dialogue and determine whether or not harassment is present. The output of this step is the result of the harassment determination.

[0535] Step 3:

[0536] The server receives the judgment result from the generating AI model, and if it determines that harassment has occurred, it generates visual information using a visual information generation means. Specifically, based on the judgment result, it creates visual information that includes educational content. Video editing software may be used to generate the visual information. The output of this step is the generated visual information.

[0537] Step 4:

[0538] The server presents the generated visual information to the user. The user can view the visual information through their device and learn about harassment issues and countermeasures. The output of this step is the visual information presented to the user.

[0539] (Application Example 2)

[0540] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0541] Workplace harassment is a serious problem that harms employees' mental health and leads to decreased productivity. However, detecting and preventing harassment is difficult, and prompt and appropriate measures are necessary, especially in situations where real-time response is required. Traditional methods often only identify harassment incidents after they have occurred, making immediate response difficult.

[0542] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0543] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the determination means; and a presentation means that presents the video generated by the generation means to the user. This makes it possible to detect harassment in the workplace environment in real time and immediately generate and present educational videos.

[0544] "Dialogue history" refers to data that records the content of a conversation and is saved in audio or text format.

[0545] A "determination means" is a device or program that has the function of analyzing input data and making a judgment based on specific conditions.

[0546] "Generation means" refers to a device or program that has the function of automatically creating new content based on input information.

[0547] "Presentation means" refers to a device or program that has the function of providing generated information or content to the user visually or audibly.

[0548] A "generative AI model" is an algorithm or program that uses artificial intelligence technology to generate new information or content from data.

[0549] A "prompt" is an instruction or question given to a generative AI model to obtain a specific output.

[0550] The system for implementing this invention mainly consists of a server and a user terminal. The server receives dialogue history as input and converts the audio data into text using a speech recognition API (e.g., Google Cloud Speech-to-Text) to determine whether the dialogue history constitutes harassment. Next, it analyzes the converted data using a natural language processing library (e.g., spaCy) to determine the possibility of harassment.

[0551] Based on the determined results, the server uses a generation AI model (e.g., OpenAI's GPT-4) to generate a video that points out the harassment issues and proposes solutions. In this process, the generation AI model receives prompt text to obtain specific output. An example of a prompt text would be, "This conversation has been determined to be power harassment. Please generate a video that points out the power harassment issues and proposes solutions."

[0552] The generated video is displayed on the user's device. The user's device is a smartphone or smart glasses, and its role is to visually provide the generated video to the user. This allows the user to understand the harassment issue in real time and take appropriate action.

[0553] For example, if a workplace conversation includes a statement like, "You're always so slow, you need to work harder," the server will identify this as harassment and input a prompt into a generation AI model to create an educational video. This video is then immediately presented to the relevant parties via the user's terminal.

[0554] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0555] Step 1:

[0556] The server receives audio data sent from the user's terminal. This audio data is a recording of a workplace conversation. The server converts this audio data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is audio data, and the output is text data.

[0557] Step 2:

[0558] The server analyzes text data using a natural language processing library (e.g., spaCy). The purpose of the analysis is to determine whether there are signs of harassment in the text. The input is text data, and the output is a judgment result indicating the possibility of harassment. Specifically, it analyzes keywords and context within the text and scores the likelihood of harassment.

[0559] Step 3:

[0560] Based on the judgment result, the server inputs a prompt message into the generating AI model (e.g., OpenAI's GPT-4). The prompt message is: "This conversation has been determined to be harassment. Please generate a video that points out the problems with harassment and proposes solutions." The input is the judgment result and the prompt message, and the output is the content of the video generated by the generating AI model. Specifically, the generating AI model analyzes the prompt message and generates appropriate video content.

[0561] Step 4:

[0562] The server uses a video generation tool to create the actual video based on the generated video content. The input is the video content generated by the AI ​​model, and the output is the completed video file. Specifically, the video generation tool converts text-based content into a visual video.

[0563] Step 5:

[0564] The server sends the completed video file to the user's terminal. The user's terminal then displays this video to the user. The input is a video file, and the output is the display of the video to the user. Specifically, the user's terminal plays the video and provides the user with visual information.

[0565] (Example 3)

[0566] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0567] In today's workplace, harassment remains a serious problem, and there is a need to prevent it and to implement concrete corrective measures when it occurs. However, traditional methods have the challenge of making it difficult to recognize harassment and obtain specific feedback for improvement.

[0568] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[0569] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a visual information generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the information processing means; and a presentation means that presents the generated visual information to the user. This enables the user to recognize the possibility that their words and actions constitute harassment and to take concrete corrective measures.

[0570] "Dialogue history" refers to a record of conversations a user has had in the past, and is data entered in text format.

[0571] "Information processing means" refers to technical means for analyzing input dialogue history and determining whether or not harassment has occurred.

[0572] "Visual information generation means" refers to a technical means that automatically generates visual information reflecting the content of harassment based on a conversation history that has been determined to be harassment.

[0573] A "presentation means" is a technical means that provides generated visual information to a user, enabling the user to visually confirm it.

[0574] "Harassment" refers to inappropriate words or actions towards others in the workplace or in society, and particularly includes inappropriate behavior and sexual harassment in the workplace.

[0575] A description of embodiments for carrying out this invention will be given.

[0576] The user first inputs their conversation history into the system via their terminal. The conversation history is in text format and includes past conversations. The user copies and pastes the conversation history into the input form on their terminal and clicks the "Start Analysis" button.

[0577] The server receives the conversation history sent by the user. The received data is preprocessed for natural language processing. Specifically, text cleaning and tokenization are performed. The server inputs the preprocessed text data into a natural language processing engine, which is an AI model for generating conversations. This engine analyzes the conversation content and determines whether or not harassment has occurred. The result of the determination is output as a score indicating whether or not harassment is likely.

[0578] Based on the assessment results, the server extracts specific details of any harassment that may have occurred. Using this information, it creates a scenario for generating visual information. The server then uses video editing software to generate visual information that reflects the extracted harassment. This visual information is designed to allow users to visually confirm their own behavior. The generated visual information is sent to the user's device for viewing.

[0579] As a concrete example, if a user wants to "check whether what they said in yesterday's meeting was appropriate," they input the conversation history into the system. An example of a prompt would be, "Please analyze what I said in yesterday's meeting and assess the possibility of harassment." Based on this prompt, the server performs the analysis, generates visual information as needed, and provides it to the user. The flow of specific processing in Example 3 will be explained using Figure 15.

[0580] Step 1:

[0581] The user enters their conversation history using a terminal. The input is in text format; the user copies and pastes the conversation history into the input form on the terminal and clicks the "Start Analysis" button. The entered conversation history is sent to the server.

[0582] Step 2:

[0583] The server preprocesses the dialogue history received from the user. Specifically, it performs text cleaning (removing unnecessary spaces and special characters) and tokenization (dividing words and phrases into parsable units). This preprocessing prepares the data in a format that can be easily analyzed by the generative AI model.

[0584] Step 3:

[0585] The server inputs pre-processed text data into a generative AI model. The generative AI model uses natural language processing techniques to analyze the dialogue and determine whether or not harassment has occurred. The result is output as a score indicating the likelihood of harassment. This score quantifies the risk of harassment.

[0586] Step 4:

[0587] Based on the assessment results, the server extracts specific details if harassment is detected. Using this extracted information, it creates a scenario for generating visual information. This scenario serves as a blueprint for determining what kind of visual information to generate.

[0588] Step 5:

[0589] The server uses video editing software to generate visual information that reflects the extracted harassment content. The generated visual information is designed to allow users to visually review their own words and actions. The visual information is sent to the user's device.

[0590] Step 6:

[0591] Users view visual information sent to their devices. This visual information includes specific examples of harassment and suggestions for improvement, allowing users to review their own behavior based on it.

[0592] (Application Example 3)

[0593] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0594] Harassment remains a serious problem in modern workplaces and educational institutions. Traditional methods are often insufficient in preventing harassment from occurring in the first place, and responses after it happens are frequently inadequate. In particular, the lack of real-time warnings and concrete improvement measures makes effective improvement difficult.

[0595] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[0596] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the determination means; and a presentation means that issues a warning to the user in real time and presents visual information including specific advice for improvement. This enables the prevention of harassment and a rapid response after it occurs.

[0597] "Dialogue history" refers to a record of conversations between users, which is saved in audio or text format.

[0598] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not its content constitutes harassment.

[0599] "Generation means" refers to a device or program that has the function of automatically creating visual information based on the dialogue history that has been determined to be harassment by the determination means.

[0600] "Presentation means" refers to a device or program that has the function of displaying generated visual information to the user and issuing warnings in real time.

[0601] "Visual information" refers to visual content such as videos and images presented to users, including details of harassment and measures to improve it.

[0602] "Real-time" refers to the instantaneous processing and response that occurs at the very moment the user's interaction is taking place.

[0603] "Advice" refers to information that suggests specific actions or behavioral changes that users should take to improve their behavior and address harassment.

[0604] The system for implementing this invention mainly consists of a server and a user terminal. The server receives dialogue history as input and is equipped with a determination means for determining whether or not the dialogue history constitutes harassment. This determination means uses a generative AI model to analyze the input text data and evaluate the likelihood of harassment.

[0605] The user's device uses a speech recognition API (e.g., Google Speech-to-Text) to convert audio data into text data. This text data is sent to a server and analyzed by a judgment mechanism. If harassment is detected as a result of the judgment, the server uses a generation mechanism to generate visual information that reflects the content of the harassment. This visual information is created using video generation software (e.g., Adobe Premiere Pro).

[0606] The generated visual information is sent to the user's device and displayed to the user in real time via a presentation tool. This allows the user to recognize if their words or actions may constitute harassment and receive specific advice for improvement.

[0607] For example, workplace conversations may be recorded, and a warning message such as "That way of speaking is inappropriate" may appear. In this case, the video may also include advice such as, "This way of speaking may hurt the other person. Next time, try rephrasing it like this."

[0608] An example of a prompt for a generative AI model is: "Analyze the following conversation and determine if it may be harassment. Conversation: 'Your way of doing things is terrible. Do it properly.'"

[0609] The flow of the specific processing in Application Example 3 will be explained using Figure 16.

[0610] Step 1:

[0611] The user's device accepts voice input. When the user speaks, the device's microphone records the audio. The recorded audio data is converted into text data using a speech recognition API (e.g., Google Speech-to-Text). This converted text data becomes the input for the next process.

[0612] Step 2:

[0613] The server receives text data sent from the user's terminal. The server analyzes this text data using a generative AI model to determine the possibility of harassment. In this analysis, prompt sentences are input into the generative AI model to determine whether or not harassment is present. The output of this step is the result of the harassment determination.

[0614] Step 3:

[0615] Based on the assessment results, the server generates visual information using a generation mechanism if harassment is detected. This visual information is created using video generation software (e.g., Adobe Premiere Pro). The generated visual information includes specific improvement measures and advice. This visual information serves as input for the next process.

[0616] Step 4:

[0617] The server transmits the generated visual information to the user's terminal. The user's terminal displays this visual information to the user in real time using a presentation mechanism. Through the presented visual information, the user can recognize if their words or actions may constitute harassment and receive specific advice for improvement.

[0618] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0619] "Example of form 1"

[0620] One embodiment of the present invention provides a system incorporating an emotion engine. This system receives a user's dialogue history as input and determines whether or not that dialogue history constitutes harassment. This determination is made by an emotion engine that recognizes the user's emotions. Specifically, it extracts emotions from the user's dialogue history and determines whether those emotions are related to harassment. For example, if a user expresses anger or dissatisfaction, it is determined that this may constitute harassment.

[0621] "Example of form 2"

[0622] Furthermore, the system provides a mechanism to adjust harassment detection based on emotions recognized by the emotion engine. Specifically, it strengthens harassment detection when a user's emotions are strongly expressed or when certain emotions are repeatedly expressed. For example, if a user repeatedly expresses anger, it may be determined to be a strong form of harassment, and this result is provided as feedback to the user.

[0623] "Example of form 3"

[0624] Furthermore, the system provides a means to generate videos that reflect the user's emotions based on the history of conversations that have been identified as harassment. Specifically, it automatically generates videos that reflect both the content of the harassment and the user's emotions. For example, if a user expresses anger, a video reflecting that anger is generated and presented to the user. This allows the user to visually understand the relationship between their emotions and the harassment.

[0625] The following describes the processing flow for each example of the form.

[0626] "Example of form 1"

[0627] Step 1: The system receives the user's interaction history.

[0628] Step 2: The emotion engine extracts the user's emotions from the conversation history.

[0629] Step 3: Determine whether the emotions extracted by the emotion engine are related to harassment.

[0630] "Example of form 2"

[0631] Step 1: Receive the user's conversation history and the emotion recognition results from the emotion engine.

[0632] Step 2: Strengthen the criteria for determining harassment when a user's emotions are strongly expressed or when certain emotions are repeatedly expressed.

[0633] Step 3: Provide feedback to the user regarding the enhanced harassment assessment results.

[0634] "Example of form 3"

[0635] Step 1: Receive the conversation history and user sentiment that were identified as harassment.

[0636] Step 2: Combine the content of the harassment with the user's emotions and automatically generate a video that reflects both.

[0637] Step 3: Present the generated video to the user.

[0638] (Example 1)

[0639] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0640] In today's workplace, harassment is a serious problem, and early detection and countermeasures are essential. However, traditional methods often rely on subjective and inaccurate assessments of harassment. Furthermore, victims may face psychological resistance to reporting their experiences. This leads to challenges such as harassment being overlooked or responses being delayed.

[0641] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0642] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; an analysis means that analyzes the dialogue history using natural language processing technology; and an emotion recognition means that extracts emotions and determines whether or not the emotions are related to harassment. This makes it possible to objectively and quickly determine whether or not harassment has occurred and to respond promptly.

[0643] "Dialogue history" is a record of linguistic information exchanged between users during communication with others.

[0644] "Information processing means" refers to a device or program that has the function of analyzing input data and making a decision based on specific conditions.

[0645] "Generation means" refers to a device or program that has the function of creating and outputting new visual information based on input information.

[0646] "Natural language processing technology" is a technology that enables computers to understand and analyze human language.

[0647] "Analysis means" refers to a device or program that has the function of examining input data in detail and understanding its structure and meaning.

[0648] An "emotion recognition means" is a device or program that has the function of extracting emotions from input data and determining whether those emotions are related to specific conditions.

[0649] "Visual information" refers to information that can be recognized visually, such as images and videos.

[0650] As an embodiment for carrying out this invention, the automated harassment analysis system is configured as follows.

[0651] The user enters the conversation history using a terminal. This conversation history is a record of the linguistic information exchanged between the user and others. The entered conversation history is transmitted to the server via the internet.

[0652] The server analyzes the received dialogue history using information processing tools. These tools include Python's natural language processing libraries, NLTK and spaCy. This allows the server to analyze the dialogue history text, extract keywords, and understand the context.

[0653] Furthermore, the server uses emotion recognition to extract emotions from the conversation history. This emotion recognition uses emotion analysis tools such as Google Cloud's Natural Language API. The server then determines whether the extracted emotions are related to harassment.

[0654] For example, if a user enters a conversation history such as "My boss insults me every day," the server analyzes this history and extracts the keyword "insult." Next, the emotion engine identifies the emotion of "anger," and based on this information, it determines that there is a high probability of harassment.

[0655] An example of a prompt to input into a generative AI model would be: "Please determine if the following conversation history constitutes harassment: 'My boss insults me every day.'" Using this prompt, the generative AI model analyzes the content of the conversation history and determines whether or not harassment occurred.

[0656] The flow of the specific processing in Example 1 will be explained using Figure 17.

[0657] Step 1:

[0658] The user enters the conversation history using a terminal. The entered conversation history is in text format and includes linguistic information exchanged between the user and others. This data is transmitted to the server via the internet.

[0659] Step 2:

[0660] The server passes the received dialogue history to an information processing system. Here, the server uses Python's natural language processing libraries, NLTK and spaCy, to parse the dialogue history text. Specifically, the server extracts keywords from the text and performs data processing to understand the context. As a result of this analysis, important keywords and phrases contained in the dialogue history are output.

[0661] Step 3:

[0662] The server passes the analyzed data to the emotion recognition system. Here, the server uses emotion analysis tools such as Google Cloud's Natural Language API to extract emotions from the conversation history. Specifically, the server identifies emotions from words and phrases in the text and determines whether those emotions are related to harassment. As a result of this process, the extracted emotions and their relevance are output.

[0663] Step 4:

[0664] The server integrates the results obtained from the information processing and emotion recognition means to ultimately determine whether the dialogue history constitutes harassment. Specifically, the server evaluates the possibility of harassment based on the extracted keywords and emotions. As a result of this determination, it outputs whether or not the dialogue history constitutes harassment.

[0665] Step 5:

[0666] The server returns the judgment result to the user. The user's terminal is notified of the judgment result and displayed as a detailed report. Specifically, the user can check whether the conversation history constitutes harassment and take appropriate action if necessary.

[0667] (Application Example 1)

[0668] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0669] Workplace harassment is a problem that negatively impacts employees' mental health and workplace productivity. However, early detection of signs of harassment and implementation of appropriate countermeasures are difficult. In particular, there is a need for a system that can detect and warn about harassment in real time.

[0670] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0671] In this invention, the server includes a determination means that receives dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs information automatically generated based on the dialogue history determined to be harassment by the determination means; and a processing means that converts the user's voice data into text and analyzes it using natural language processing technology. This enables early detection of harassment in the workplace and prompt warnings.

[0672] "Dialogue history" refers to a record of conversations between users, which is saved as audio or text data.

[0673] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not its content constitutes harassment.

[0674] "Generation means" refers to a device or program that has the function of automatically creating and outputting relevant information and warnings based on the dialogue history that has been determined to be harassment by the determination means.

[0675] "Processing means" refers to a device or program that has the function of converting the user's voice data into text and analyzing that text using natural language processing technology.

[0676] "Analysis tool" refers to a device or program that has the function of extracting emotions from text data and evaluating the possibility of harassment.

[0677] "Notification means" refers to a device or program that has the function of issuing a warning to a user in the event of potential harassment.

[0678] A "portable information terminal" is an electronic device that is portable and capable of processing information, such as a smartphone or tablet.

[0679] The system that realizes this invention mainly consists of a server and a user's mobile information terminal. The server receives dialogue history as input and is equipped with a determination means for determining whether or not the dialogue history constitutes harassment. The determination means analyzes the dialogue history using natural language processing technology and evaluates the possibility of harassment. Specifically, it converts the user's voice data into text and performs analysis using a natural language processing library (e.g., spaCy, NLTK).

[0680] Furthermore, the server is equipped with analytical tools to extract emotions from text data and determine the possibility of harassment. Using emotion analysis libraries (e.g., TextBlob, VADER), it evaluates the user's emotions, and if anger or dissatisfaction is detected, it determines that there is a possibility of harassment.

[0681] The user's mobile device is equipped with a notification system to receive notifications from the server. If harassment is suspected, a warning will be displayed on the device to alert the user. This enables early detection and rapid response to harassment in the workplace.

[0682] As a concrete example, if a supervisor makes an inappropriate remark to a subordinate during a workplace conversation, the server analyzes the remark and, if it determines that it may constitute harassment, displays a warning on the subordinate's mobile device. An example of a prompt to be input to the generating AI model would be, "Analyze the content of this conversation and determine if it may constitute harassment."

[0683] The flow of a specific process in Application Example 1 will be explained using Figure 18.

[0684] Step 1:

[0685] The server receives audio data from the user's mobile device. This audio data is a recording of a conversation at work. The server uses speech recognition software to convert this audio data into text data.

[0686] Step 2:

[0687] The server analyzes the converted text data using natural language processing libraries (e.g., spaCy, NLTK). This analysis extracts grammatical structures and keywords from the text, generating foundational data for understanding the dialogue.

[0688] Step 3:

[0689] The server extracts emotions from the analyzed text data using a sentiment analysis library (e.g., TextBlob, VADER). In this step, it evaluates the emotional expressions contained in the text and determines whether negative emotions such as anger or frustration are present.

[0690] Step 4:

[0691] The server determines the possibility of harassment based on the results of sentiment analysis. Specifically, if negative emotions exceed a certain threshold, it determines that harassment is possible. This determination can be further improved using a generative AI model.

[0692] Step 5:

[0693] If the server determines that harassment may be occurring, it sends a warning to the user's mobile device. The device then notifies the user of this warning and prompts them to take action. This allows the user to recognize the risk of harassment in real time.

[0694] (Example 2)

[0695] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0696] In today's workplace, harassment is a serious problem, and it is essential to properly detect it and take countermeasures. However, traditional methods have challenges in making accurate judgments because the determination of harassment is subjective and does not take into account changes in emotions.

[0697] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0698] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the information processing means; and an adjustment means that analyzes emotions and adjusts the harassment determination based on the analysis results. This makes it possible to make an objective and emotionally conscious determination of harassment.

[0699] "Dialogue history" refers to a record of conversations that users have had, and is data saved in audio or text format.

[0700] "Information processing means" refers to technical means for analyzing dialogue history and determining whether or not harassment occurred.

[0701] "Generation means" refers to a technical means for automatically creating and outputting visual information based on dialogue history that has been determined to be harassment.

[0702] "Adjustment measures" refer to technical means for analyzing user emotions and modifying or strengthening harassment judgments based on the results.

[0703] "Visual information" refers to visually presented information such as videos and images that reflect the content of the harassment.

[0704] "Emotion" refers to the user's psychological state and includes changes in feelings and emotions inferred from voice and text.

[0705] In an embodiment of this invention, a server, a terminal, and a user cooperate to operate the system. First, the user uses a terminal to record everyday conversations. The terminal is equipped with speech recognition software, which converts the speech data into text data. The converted text data is sent to the server as a conversation history.

[0706] The server analyzes the received dialogue history using information processing tools. This analysis employs natural language processing technology to detect specific keywords and phrases, thereby determining whether harassment has occurred. Furthermore, the server uses an emotion engine to analyze the user's emotions and adjusts the harassment determination based on the results. For example, if a user repeatedly expresses "anger," the server takes that emotion into account and strengthens its determination.

[0707] If harassment is detected, the server automatically generates visual information using a generation mechanism. This visual information is a video created using a generation AI model to represent a scenario corresponding to the content of the conversation. The generated video is sent to the user's device and presented to the user. By watching this video, the user can deepen their understanding of the nature of the harassment and its impact.

[0708] As a concrete example, let's look at a prompt message to input into a generation AI model: "If a user repeatedly expresses anger in conversations with their boss, please generate a video that points out the possibility of workplace harassment based on that conversation history." Using this prompt message, the server can generate an appropriate video and provide it to the user.

[0709] The flow of the specific processing in Example 2 will be explained using Figure 19.

[0710] Step 1:

[0711] The user records a conversation using a device. The device uses speech recognition software to convert the audio data into text data. The input for this step is audio data, and the output is text data. Specifically, the user records a conversation at work and then transcribes that recording into text.

[0712] Step 2:

[0713] The server receives text data sent from the terminal. The server uses information processing tools to analyze the text data and determine whether or not harassment has occurred. The input for this step is text data, and the output is the result of the harassment determination. Specifically, the server utilizes natural language processing technology to detect certain keywords and phrases.

[0714] Step 3:

[0715] The server uses an emotion engine to analyze the user's emotions from text data. Based on this analysis, it adjusts the harassment judgment. The input for this step is text data, and the output is the adjusted judgment result. Specifically, the server infers emotions from voice tone and text content, and strengthens or modifies the judgment.

[0716] Step 4:

[0717] If the server determines that harassment has occurred, it automatically generates visual information using a generation mechanism. The input for this step is the adjusted judgment result, and the output is visual information (video). Specifically, the server utilizes a generation AI model to create a video of a scenario that corresponds to the content of the conversation.

[0718] Step 5:

[0719] The server sends the generated visual information to the user's device and presents it to the user. The input in this step is visual information, and the output is the presentation to the user. Specifically, the user watches a video on their device to deepen their understanding of the content and impact of harassment.

[0720] (Application Example 2)

[0721] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0722] In modern workplaces and educational institutions, harassment is a serious problem, and early detection and countermeasures are essential. However, traditional methods can lead to delays in detecting harassment or failure to implement appropriate measures. In particular, there is a challenge in determining harassment while considering emotional changes.

[0723] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0724] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs a video that is automatically generated based on the dialogue history determined to be harassment by the determination means; and an adjustment means that recognizes the user's emotions and adjusts the harassment determination based on those emotions. This enables early detection of harassment and the implementation of appropriate countermeasures.

[0725] "Dialogue history" refers to a record of conversations between users, which is saved as text data.

[0726] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not it constitutes harassment.

[0727] "Generation means" refers to a device or program that has the function of automatically creating and outputting a video that points out the problematic aspects based on the history of a conversation that has been determined to be harassment.

[0728] "Adjustment measures" refer to devices or programs that recognize the user's emotions and have the function of strengthening or mitigating the determination of harassment based on those emotions.

[0729] "Emotions" refer to the psychological states and reactions that users exhibit during a conversation, including states such as anger, sadness, and joy.

[0730] The system for carrying out this invention includes a server and a user terminal. The server is equipped with a determination means for receiving and analyzing dialogue history. The determination means uses natural language processing technology to determine whether the dialogue history constitutes harassment. Specifically, it uses the Google Cloud Natural Language API to analyze text data and extract emotions and intentions.

[0731] The user's device, such as a smartphone or computer, is responsible for sending conversation history to the server. It also receives generated video from the server and presents it to the user. The generation method uses the Adobe Premiere Pro API to automatically generate video based on conversation history identified as harassment. This video includes content that points out the problem and prompts the user to take action.

[0732] Furthermore, the server is equipped with adjustment mechanisms to recognize user emotions in real time. The emotion engine enhances harassment detection when user emotions are strongly expressed or when certain emotions are repeatedly expressed. This enables more accurate detection.

[0733] For example, if workplace conversations are recorded and the user repeatedly expresses "anger," the server analyzes the conversation and determines that it may constitute workplace harassment. It then generates a video highlighting the harassment issues and notifies the user.

[0734] Examples of prompts for a generative AI model include the following:

[0735] "Analyze the following conversation history and determine if harassment is likely. Strengthen the assessment if the user's emotions are strongly expressed. If harassment is detected, generate a video highlighting the problem."

[0736] The flow of a specific process in Application Example 2 will be explained using Figure 20.

[0737] Step 1:

[0738] The user's terminal sends the conversation history as text data to the server. The input is the user's conversation history, and the output is the data transmission to the server. In this step, the user's terminal collects the conversation history and sends it to the server over the network.

[0739] Step 2:

[0740] The server analyzes the received dialogue history using the Google Cloud Natural Language API. The input is the text data of the dialogue history, and the output is the result of the sentiment and intent analysis. In this step, the server uses natural language processing techniques to extract sentiment and intent from the dialogue history.

[0741] Step 3:

[0742] The server determines the possibility of harassment based on the analysis results. The input is the analysis results of emotions and intentions, and the output is the harassment determination result. In this step, the server uses a determination means to evaluate the analysis results and determine whether or not harassment occurred.

[0743] Step 4:

[0744] If harassment is detected, the server generates a video highlighting the problem using the Adobe Premiere Pro API. The input is the harassment detection result and the dialogue history, and the output is the generated video. In this step, the server uses a generation mechanism to automatically create a video based on the dialogue history.

[0745] Step 5:

[0746] The server sends the generated video to the user's terminal. The input is the generated video, and the output is the transmission of the video to the user's terminal. In this step, the server transmits the video to the user's terminal over the network.

[0747] Step 6:

[0748] The user's device displays the received video to the user. The input is the video transmitted from the server, and the output is the video displayed to the user. In this step, the user's device plays the video and provides information to the user visually.

[0749] (Example 3)

[0750] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0751] In modern workplaces and society, harassment remains a serious problem, and preventing its occurrence is essential. However, it is difficult to determine what kind of words and actions constitute harassment in individual conversations and communications, and when harassment does occur, concrete corrective measures are required. Furthermore, there is a lack of visual means to understand the emotions associated with harassment, so these challenges need to be addressed.

[0752] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[0753] In this invention, the server includes: information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; data generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the information processing means; and data generation means that analyzes the emotions contained in the dialogue history and generates a video that reflects those emotions. This makes it possible for users to recognize the possibility that their words and actions constitute harassment and to visually understand concrete measures to improve their behavior.

[0754] "Dialogue history" refers to a record of conversations and communications between users, and is data saved in text format.

[0755] "Information processing means" refers to a device or program that has the function of analyzing input dialogue history and determining whether its content constitutes harassment.

[0756] "Data generation means" refers to a device or program that has the function of automatically creating images and other visual content based on the determination results of information processing means.

[0757] "Emotional analysis" is the process of identifying the user's emotions contained in the conversation history and evaluating the type and intensity of those emotions.

[0758] "Video" refers to content that includes visual information and is presented to users in the form of videos, animations, and other visual media.

[0759] A description of embodiments for carrying out this invention will be given.

[0760] The server receives the conversation history entered by the user. This conversation history is a record of the conversations and communications the user has had, and is provided in text format. The server analyzes this conversation history using natural language processing technology. Specifically, a common cloud-based natural language processing API can be used as the natural language processing software. This allows the server to understand the content of the conversation and detect signs of harassment.

[0761] Next, the server uses information processing means to determine whether the conversation history constitutes harassment. Based on this determination, the server uses data generation means to generate a video that reflects the content of the harassment. Video editing software can be used for this video generation.

[0762] Furthermore, the server uses sentiment analysis tools to analyze the user's emotions contained in the conversation history. Based on the results of the sentiment analysis, the server generates a video that reflects the user's emotions. This video is intended to help the user visually understand the relationship between their emotions and harassment.

[0763] As a concrete example, consider a scenario where a user enters a conversation history stating, "I used harsh words towards my subordinate during yesterday's meeting." The server analyzes this history and determines whether the harsh language constitutes harassment. If it does, the server generates a video based on the content that reflects the user's anger and presents it to the user.

[0764] An example of a prompt to be input to the generation AI model is, "Based on this dialogue history, determine whether harassment occurred and, if necessary, generate a video that reflects the emotions." The flow of the specific processing in Example 3 will be explained using Figure 21.

[0765] Step 1:

[0766] The user enters their conversation history. The user enters their conversation history into the system interface. The entered data is sent to the server in text format.

[0767] Step 2:

[0768] The server receives the dialogue history and performs natural language processing. The server analyzes the received text data using natural language processing software. This analysis helps understand the content and context of the dialogue and prepares it to detect signs of harassment. The input is text data, and the output is the analysis result.

[0769] Step 3:

[0770] The server determines whether harassment has occurred. Based on the results of natural language processing, the server determines whether harassment is present in the dialogue history. Specifically, it evaluates whether aggressive words or inappropriate expressions are included. The input is the analysis result, and the output is the determination result regarding the presence or absence of harassment.

[0771] Step 4:

[0772] If the server detects harassment, it prepares data for video generation. When harassment is detected, the server creates a script or template for video generation based on the details. The input is the detection result, and the output is the data for video generation.

[0773] Step 5:

[0774] The server analyzes the user's emotions. The server uses an emotion analysis tool to identify the user's emotions as they appear in the conversation history. The input is the conversation history, and the output is the result of the emotion analysis.

[0775] Step 6:

[0776] The server generates a video that reflects emotions. The server generates a video that reflects the content of the harassment and the user's emotions. Specifically, it uses video editing software to create visually represented content. The input is data for video generation and the results of emotion analysis, and the output is the generated video.

[0777] Step 7:

[0778] The server generates a video and presents it to the user. The server sends the generated video to the user, making it viewable. The input is the generated video, and the output is its presentation to the user.

[0779] (Application Example 3)

[0780] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0781] Harassment in the workplace and online communication is a problem that negatively impacts individual mental health and workplace productivity. However, there are insufficient systems to detect signs of harassment early and provide appropriate feedback. Furthermore, there is a lack of means for users to visually understand the connection between their own feelings and harassment and to take concrete actions toward improvement.

[0782] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[0783] In this invention, the server includes means for receiving a conversation history as input and determining whether the conversation history constitutes harassment; means for outputting a video automatically generated based on the conversation history determined to be harassment by the harassment determination means; and means for analyzing the user's emotions and generating a video that reflects those emotions. This allows the user to receive specific feedback to improve their communication style and prevent harassment from occurring.

[0784] "Dialogue history" refers to a record of conversations between users, which is saved in text format.

[0785] A "harassment determination tool" is a device that analyzes the input conversation history and determines whether or not its content constitutes harassment.

[0786] A "video generation method" is a system that automatically creates videos to provide visual feedback based on conversation history and user emotions that have been identified as harassment.

[0787] "User emotions" refers to the user's psychological state and emotional expression in their dialogue history, and this is analyzed and reflected in the video generation process.

[0788] A "generative AI model" refers to an algorithm or system that uses artificial intelligence technology to generate new content from input data.

[0789] "Feedback" refers to information and advice provided to users that helps improve communication styles.

[0790] The system for implementing this invention consists of a server and a user terminal. The server receives the conversation history transmitted from the user and analyzes its content using a harassment detection means. Specifically, it processes the text data using the natural language processing library spaCy and evaluates the possibility of harassment. Based on the evaluation results, it generates a video that reflects the content of the harassment using OpenAI's GPT generative AI model.

[0791] The video generation process includes analyzing the user's emotions. These emotions are extracted from the conversation history and visually represented by the video generation tool. Video editing libraries such as MoviePy are used for video generation. The generated video is sent to the user's device and presented to them. This allows the user to receive specific feedback to improve their communication style.

[0792] As a concrete example, consider email exchanges in the workplace. If a user sends an email saying, "This project isn't progressing at all, it's all your fault," the server analyzes this text and detects the possibility of harassment. An example of a prompt would be, "Evaluate whether the following text constitutes harassment: 'This project isn't progressing at all, it's all your fault'," which would be input into the GPT model. The generated video provides visual feedback to help the user reflect on their own behavior and make improvements.

[0793] The flow of the specific processing in Application Example 3 will be explained using Figure 22.

[0794] Step 1:

[0795] The user sends the conversation history to the server using their device. The input is the user's conversation history, sent to the server in text format. The server receives this data and prepares for the next processing step.

[0796] Step 2:

[0797] The server analyzes the received dialogue history using the natural language processing library spaCy. The input is the text data of the dialogue history, and the output is the analysis results indicating the possibility of harassment. The server detects signs of harassment by tokenizing the text and analyzing its grammatical structure.

[0798] Step 3:

[0799] Based on the analysis results, the server inputs prompt text into the OpenAI GPT, a generative AI model. The input is prompt text that reflects the analysis results, and the output is an evaluation of whether or not harassment occurred. Specifically, the server generates prompt text and sends it to the GPT model.

[0800] Step 4:

[0801] The server receives evaluation results from the GPT model and, if it determines that harassment has occurred, generates a video using a video generation method. The input is the evaluation results and user sentiment data, and the output is the generated video. The server uses a video editing library such as MoviePy to create a video that reflects the user's sentiment.

[0802] Step 5:

[0803] The server sends the generated video to the user's terminal and presents it to the user. The input is the generated video, and the output is the video playback on the user's terminal. The user watches this video and receives feedback to improve their communication style.

[0804] (Other examples)

[0805] Since this is the same as the specific processing described in the other embodiments of the first embodiment above, the explanation will be omitted.

[0806] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0807] The data generation model 58 is a form of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0808] Other examples of generative AI include Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) are some examples.

[0809] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0810] [Third Embodiment]

[0811] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0812] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0813] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0814] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0815] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0816] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0817] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0818] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0819] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0820] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0821] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0822] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.

[0823] "Example of form 1"

[0824] One embodiment of the present invention provides an automated harassment analysis system. This system receives a conversation history from a user as input and has a harassment determination means that determines whether or not the conversation history constitutes harassment. The harassment determination means analyzes the conversation history using natural language processing technology and determines whether or not harassment such as power harassment or sexual harassment has occurred.

[0825] "Example of form 2"

[0826] Furthermore, the system of the present invention includes a video generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the harassment determination means. The video generation means generates a video that reflects the content of the dialogue history determined to be harassment and presents it to the user. For example, if the dialogue history determined to be harassment is related to power harassment, it generates a video that points out the problems of power harassment.

[0827] "Example of form 3"

[0828] The system of the present invention can prevent harassment from occurring and provide concrete material for reflection and improvement in cases of harassment that have occurred. Specifically, by having the user input their conversation history into the system, the presence or absence of harassment is determined, and if harassment is found, a video reflecting the content is generated and presented. As a result, the user can recognize that their words and actions may constitute harassment and take concrete actions to improve them.

[0829] The following describes the processing flow for each example of the form.

[0830] "Example of form 1"

[0831] Step 1: The system receives the user's interaction history.

[0832] Step 2: The system's harassment detection mechanism analyzes the conversation history. This analysis uses natural language processing technology to determine whether or not harassment, such as power harassment or sexual harassment, has occurred.

[0833] "Example of form 2"

[0834] Step 1: The system receives the conversation history that has been determined to be harassment by the harassment detection method.

[0835] Step 2: The system's video generation mechanism automatically generates a video based on the conversation history that was identified as harassment. The video content reflects the content of the harassment.

[0836] Step 3: Present the generated video to the user.

[0837] "Example of form 3"

[0838] Step 1: The user enters their conversation history into the system.

[0839] Step 2: The system determines whether harassment occurred, and if so, generates a video reflecting the details of the harassment.

[0840] Step 3: Present the generated video to the user, allowing them to recognize that their words and actions may constitute harassment.

[0841] Step 4: Based on the presented video, users take specific actions to improve the harassment situation.

[0842] (Example 1)

[0843] Next, we will describe Embodiment 1 of Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0844] In today's workplace, harassment is a serious problem, and its detection and response are crucial. However, traditional methods have made it difficult to efficiently analyze dialogue history and accurately determine the presence or absence of harassment. Furthermore, there has been a lack of visual representations of harassment, making it difficult to understand and share the problem.

[0845] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0846] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; an analysis means that analyzes the dialogue history using natural language processing technology; and an evaluation means that evaluates the analysis results using a generative AI model and determines whether or not harassment has occurred. This enables efficient analysis of dialogue history and accurate determination of harassment.

[0847] "Dialogue history" refers to a record of conversations and messages exchanged between users, and is data saved in text format.

[0848] "Information processing means" refers to a device or program that has the function of analyzing input data and making decisions based on specific conditions.

[0849] "Generation means" refers to a device or program for automatically creating new visual information or content based on input information.

[0850] "Natural language processing technology" refers to technologies for understanding and analyzing human language using computers, and includes text tokenization and grammatical analysis.

[0851] "Analysis means" refers to a device or program for analyzing input data in detail and understanding its structure and meaning.

[0852] A "generative AI model" is a model that uses artificial intelligence technology to learn from data and is trained to perform a specific task.

[0853] "Evaluation means" refers to a device or program that makes a judgment based on analyzed data according to specific criteria.

[0854] To implement this invention, the user must first input their conversation history. The user uses their device to input the content of workplace conversations and messages in text format. The inputted conversation history is then sent from the device to the server.

[0855] The server uses Python's natural language processing libraries, NLTK and spaCy, to analyze the received dialogue history. Using this software, the server tokenizes the text data, tags it with parts of speech, and analyzes its grammatical structure.

[0856] Next, the server evaluates the analysis results using a generative AI model. This model is built using deep learning frameworks such as TensorFlow and PyTorch and is pre-trained on a dataset related to harassment. The model detects elements of power harassment and sexual harassment in the dialogue history and determines whether or not they exist.

[0857] For example, if a user enters a conversation history such as "My boss forces me to work unreasonable overtime every day," the server analyzes this data and uses a generative AI model to determine that "it contains elements of power harassment." This result is sent back to the terminal and displayed to the user.

[0858] An example of a prompt message is, "Please determine if this conversation history contains elements of harassment." By using this prompt message, the server can efficiently analyze the conversation history and accurately determine whether or not harassment has occurred.

[0859] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0860] Step 1:

[0861] Users input conversation history in text format using a terminal. This input data includes the content of workplace conversations and messages. The entered conversation history is sent from the terminal to the server.

[0862] Step 2:

[0863] The server uses natural language processing libraries such as NLTK and spaCy to analyze the dialogue history received from the terminal. The server tokenizes the input text data, tags it with parts of speech, and analyzes its grammatical structure. This analysis extracts structural information from the text.

[0864] Step 3:

[0865] The server inputs the analyzed data into a generating AI model. This model is built using TensorFlow and PyTorch and has been pre-trained on a dataset related to harassment. Based on the input data, the model detects elements of power harassment and sexual harassment and determines whether or not they exist. This determination result is then output.

[0866] Step 4:

[0867] The server returns the judgment result obtained from the generated AI model to the terminal. The judgment result indicates whether or not the conversation history contains elements of harassment.

[0868] Step 5:

[0869] The terminal displays the judgment results received from the server to the user. For example, a message such as "This conversation contains elements of power harassment" might be displayed. This allows the user to understand the content of the conversation history and take necessary measures.

[0870] (Application Example 1)

[0871] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0872] Workplace harassment is a serious problem that negatively impacts employees' mental health and workplace productivity. However, early detection and appropriate response to signs of harassment are difficult. Traditional methods often rely on victims reporting the issue themselves, and problems may go undetected until they become severe. An effective system is needed to improve this situation and prevent harassment proactively.

[0873] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0874] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a conversion means that converts speech into text; and an analysis means that analyzes the text converted by the conversion means and detects signs of harassment. This makes it possible to detect signs of harassment in the workplace environment in real time and notify administrators.

[0875] "Dialogue history" refers to a record of conversations between users, which is saved as audio or text data.

[0876] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not it constitutes harassment.

[0877] "Generation means" refers to a device or program that has the function of automatically creating relevant information and content based on the dialogue history that has been determined to be harassment by the determination means.

[0878] "Conversion means" refers to a device or program that has the function of converting audio data into text data.

[0879] "Analysis means" refers to a device or program that has the function of analyzing text data and detecting signs of harassment.

[0880] A "notification means" is a device or program that has the function of sending warnings or information to administrators or relevant parties based on signs of harassment detected by an analysis means.

[0881] The system for implementing this invention mainly consists of a server and a user terminal. The server receives dialogue history as input and is equipped with a determination means for determining whether or not the dialogue history constitutes harassment. This determination means analyzes the dialogue history using natural language processing technology to determine whether or not harassment has occurred. Specifically, it uses the Google Cloud Natural Language API to analyze text data.

[0882] The user terminal is equipped with a conversion mechanism for converting speech to text. This conversion mechanism uses the Google Cloud Speech-to-Text API to convert speech data into text data. The converted text data is sent to a server and analyzed by a determination mechanism.

[0883] If the server detects signs of harassment as a result of its analysis, it will send a warning to the administrator using a notification system. This notification system uses Firebase Cloud Messaging to provide real-time notifications.

[0884] For example, if the statement "You're always useless" is made during a conversation at work, this statement may be judged as potentially being harassment, and a notification will be sent to the manager. An example of a prompt to the generating AI model would be, "Please determine whether the following conversation constitutes harassment: 'You're always useless.'"

[0885] In this way, the system can detect signs of harassment in the workplace in real time, enabling a rapid response.

[0886] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0887] Step 1:

[0888] The user terminal captures workplace conversations as audio data. This audio data is collected in real time through the user terminal's microphone.

[0889] Step 2:

[0890] The user's device converts the acquired audio data into text data using the Google Cloud Speech-to-Text API. This conversion process analyzes the audio signal and generates the corresponding text. The input is audio data, and the output is text data.

[0891] Step 3:

[0892] The user's terminal sends the converted text data to the server. The server analyzes the received text data using the Google Cloud Natural Language API. The analysis evaluates linguistic features within the text and detects signs of harassment. The input is text data, and the output is a determination of whether or not harassment occurred.

[0893] Step 4:

[0894] If the server detects signs of harassment based on its analysis, it will send a notification to the administrator using Firebase Cloud Messaging. The notification will include detailed information such as the nature and time of the harassment. The input is the analysis result, and the output is the notification to the administrator.

[0895] Step 5:

[0896] The administrator will take necessary actions based on the received notification. This includes confirming with relevant parties and conducting further investigations. The administrator's actions will be based on the system output.

[0897] (Example 2)

[0898] Next, we will describe Example 2 of the morphological example. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0899] In today's workplace, harassment is a serious problem, and early detection and appropriate response are essential. However, traditional methods have the drawback of subjective harassment assessments, making it difficult to implement appropriate education and countermeasures. Therefore, a system is needed that objectively assesses harassment and provides educational visual information based on that assessment.

[0900] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0901] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a visual information generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the information processing means; and a presentation means that presents the visual information generated by the visual information generation means to the user. This makes it possible to objectively determine harassment and provide educational visual information based on its content.

[0902] "Dialogue history" refers to a record of conversations between users, which is saved in text format.

[0903] "Information processing means" refers to a device or program that has the function of analyzing input data and making decisions based on specific conditions.

[0904] "Visual information generation means" refers to a device or program that has the function of automatically creating visual content based on analysis results.

[0905] "Presentation means" refers to a device or program for displaying or providing generated visual information to a user.

[0906] "Harassment" refers to inappropriate words or actions towards others in the workplace or other environments, and is an act that causes mental or physical distress.

[0907] A description of embodiments for carrying out this invention will be given.

[0908] The server receives the conversation history provided by the user. The user inputs the conversation history in text format via their terminal and sends it to the server. This conversation history is used as data for harassment assessment.

[0909] The server analyzes the dialogue history using a generative AI model. Specifically, it uses natural language processing techniques to determine whether the content of the dialogue constitutes harassment. In this process, a generative AI model with excellent natural language processing capabilities is used as the general-purpose model. The server inputs the prompt "Please determine whether this dialogue history constitutes harassment" into the generative AI model.

[0910] If harassment is detected, the server automatically generates visual information based on the detection result using a visual information generation system. This visual information includes educational content and aims to inform users about the problems and countermeasures against harassment. General video editing software may be used to generate the visual information.

[0911] For example, if a user enters a conversation history stating that "a superior excessively reprimanded a subordinate," the server will determine this to be harassment. Next, the server generates and presents visual information on the theme of "the impact of inappropriate behavior in the workplace and countermeasures." This visual information includes specific examples and countermeasures, allowing the user to deepen their understanding.

[0912] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0913] Step 1:

[0914] The user uses a terminal to input the conversation history in text format and sends it to the server. The input data is a record of the conversation between users and serves as the basis for determining whether harassment occurred. The server receives this conversation history and prepares for the next processing step.

[0915] Step 2:

[0916] The server inputs the received dialogue history into the generating AI model. Specifically, it uses the prompt "Determine whether this dialogue history constitutes harassment" to have the generating AI model analyze the dialogue history. The generating AI model uses natural language processing techniques to analyze the content of the dialogue and determine whether or not harassment is present. The output of this step is the result of the harassment determination.

[0917] Step 3:

[0918] The server receives the judgment result from the generating AI model, and if it determines that harassment has occurred, it generates visual information using a visual information generation means. Specifically, based on the judgment result, it creates visual information that includes educational content. Video editing software may be used to generate the visual information. The output of this step is the generated visual information.

[0919] Step 4:

[0920] The server presents the generated visual information to the user. The user can view the visual information through their device and learn about harassment issues and countermeasures. The output of this step is the visual information presented to the user.

[0921] (Application Example 2)

[0922] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0923] Workplace harassment is a serious problem that harms employees' mental health and leads to decreased productivity. However, detecting and preventing harassment is difficult, and prompt and appropriate measures are necessary, especially in situations where real-time response is required. Traditional methods often only identify harassment incidents after they have occurred, making immediate response difficult.

[0924] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0925] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the determination means; and a presentation means that presents the video generated by the generation means to the user. This makes it possible to detect harassment in the workplace environment in real time and immediately generate and present educational videos.

[0926] "Dialogue history" refers to data that records the content of a conversation and is saved in audio or text format.

[0927] A "determination means" is a device or program that has the function of analyzing input data and making a judgment based on specific conditions.

[0928] "Generation means" refers to a device or program that has the function of automatically creating new content based on input information.

[0929] "Presentation means" refers to a device or program that has the function of providing generated information or content to the user visually or audibly.

[0930] A "generative AI model" is an algorithm or program that uses artificial intelligence technology to generate new information or content from data.

[0931] A "prompt" is an instruction or question given to a generative AI model to obtain a specific output.

[0932] The system for implementing this invention mainly consists of a server and a user terminal. The server receives dialogue history as input and converts the audio data into text using a speech recognition API (e.g., Google Cloud Speech-to-Text) to determine whether the dialogue history constitutes harassment. Next, it analyzes the converted data using a natural language processing library (e.g., spaCy) to determine the possibility of harassment.

[0933] Based on the determined results, the server uses a generation AI model (e.g., OpenAI's GPT-4) to generate a video that points out the harassment issues and proposes solutions. In this process, the generation AI model receives prompt text to obtain specific output. An example of a prompt text would be, "This conversation has been determined to be power harassment. Please generate a video that points out the power harassment issues and proposes solutions."

[0934] The generated video is displayed on the user's device. The user's device is a smartphone or smart glasses, and its role is to visually provide the generated video to the user. This allows the user to understand the harassment issue in real time and take appropriate action.

[0935] For example, if a workplace conversation includes a statement like, "You're always so slow, you need to work harder," the server will identify this as harassment and input a prompt into a generation AI model to create an educational video. This video is then immediately presented to the relevant parties via the user's terminal.

[0936] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0937] Step 1:

[0938] The server receives audio data sent from the user's terminal. This audio data is a recording of a workplace conversation. The server converts this audio data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is audio data, and the output is text data.

[0939] Step 2:

[0940] The server analyzes text data using a natural language processing library (e.g., spaCy). The purpose of the analysis is to determine whether there are signs of harassment in the text. The input is text data, and the output is a judgment result indicating the possibility of harassment. Specifically, it analyzes keywords and context within the text and scores the likelihood of harassment.

[0941] Step 3:

[0942] Based on the judgment result, the server inputs a prompt message into the generating AI model (e.g., OpenAI's GPT-4). The prompt message is: "This conversation has been determined to be harassment. Please generate a video that points out the problems with harassment and proposes solutions." The input is the judgment result and the prompt message, and the output is the content of the video generated by the generating AI model. Specifically, the generating AI model analyzes the prompt message and generates appropriate video content.

[0943] Step 4:

[0944] The server uses a video generation tool to create the actual video based on the generated video content. The input is the video content generated by the AI ​​model, and the output is the completed video file. Specifically, the video generation tool converts text-based content into a visual video.

[0945] Step 5:

[0946] The server sends the completed video file to the user's terminal. The user's terminal then displays this video to the user. The input is a video file, and the output is the display of the video to the user. Specifically, the user's terminal plays the video and provides the user with visual information.

[0947] (Example 3)

[0948] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0949] In today's workplace, harassment remains a serious problem, and there is a need to prevent it and to implement concrete corrective measures when it occurs. However, traditional methods have the challenge of making it difficult to recognize harassment and obtain specific feedback for improvement.

[0950] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[0951] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a visual information generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the information processing means; and a presentation means that presents the generated visual information to the user. This enables the user to recognize the possibility that their words and actions constitute harassment and to take concrete corrective measures.

[0952] "Dialogue history" refers to a record of conversations a user has had in the past, and is data entered in text format.

[0953] "Information processing means" refers to technical means for analyzing input dialogue history and determining whether or not harassment has occurred.

[0954] "Visual information generation means" refers to a technical means that automatically generates visual information reflecting the content of harassment based on a conversation history that has been determined to be harassment.

[0955] A "presentation means" is a technical means that provides generated visual information to a user, enabling the user to visually confirm it.

[0956] "Harassment" refers to inappropriate words or actions towards others in the workplace or in society, and particularly includes inappropriate behavior and sexual harassment in the workplace.

[0957] A description of embodiments for carrying out this invention will be given.

[0958] The user first inputs their conversation history into the system via their terminal. The conversation history is in text format and includes past conversations. The user copies and pastes the conversation history into the input form on their terminal and clicks the "Start Analysis" button.

[0959] The server receives the conversation history sent by the user. The received data is preprocessed for natural language processing. Specifically, text cleaning and tokenization are performed. The server inputs the preprocessed text data into a natural language processing engine, which is an AI model for generating conversations. This engine analyzes the conversation content and determines whether or not harassment has occurred. The result of the determination is output as a score indicating whether or not harassment is likely.

[0960] Based on the assessment results, the server extracts specific details of any harassment that may have occurred. Using this information, it creates a scenario for generating visual information. The server then uses video editing software to generate visual information that reflects the extracted harassment. This visual information is designed to allow users to visually confirm their own behavior. The generated visual information is sent to the user's device for viewing.

[0961] As a concrete example, if a user wants to "check whether what they said in yesterday's meeting was appropriate," they input the conversation history into the system. An example of a prompt would be, "Please analyze what I said in yesterday's meeting and assess the possibility of harassment." Based on this prompt, the server performs the analysis, generates visual information as needed, and provides it to the user. The flow of specific processing in Example 3 will be explained using Figure 15.

[0962] Step 1:

[0963] The user enters their conversation history using a terminal. The input is in text format; the user copies and pastes the conversation history into the input form on the terminal and clicks the "Start Analysis" button. The entered conversation history is sent to the server.

[0964] Step 2:

[0965] The server preprocesses the dialogue history received from the user. Specifically, it performs text cleaning (removing unnecessary spaces and special characters) and tokenization (dividing words and phrases into parsable units). This preprocessing prepares the data in a format that can be easily analyzed by the generative AI model.

[0966] Step 3:

[0967] The server inputs pre-processed text data into a generative AI model. The generative AI model uses natural language processing techniques to analyze the dialogue and determine whether or not harassment has occurred. The result is output as a score indicating the likelihood of harassment. This score quantifies the risk of harassment.

[0968] Step 4:

[0969] Based on the assessment results, the server extracts specific details if harassment is detected. Using this extracted information, it creates a scenario for generating visual information. This scenario serves as a blueprint for determining what kind of visual information to generate.

[0970] Step 5:

[0971] The server uses video editing software to generate visual information that reflects the extracted harassment content. The generated visual information is designed to allow users to visually review their own words and actions. The visual information is sent to the user's device.

[0972] Step 6:

[0973] Users view visual information sent to their devices. This visual information includes specific examples of harassment and suggestions for improvement, allowing users to review their own behavior based on it.

[0974] (Application Example 3)

[0975] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0976] Harassment remains a serious problem in modern workplaces and educational institutions. Traditional methods are often insufficient in preventing harassment from occurring in the first place, and responses after it happens are frequently inadequate. In particular, the lack of real-time warnings and concrete improvement measures makes effective improvement difficult.

[0977] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[0978] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the determination means; and a presentation means that issues a warning to the user in real time and presents visual information including specific advice for improvement. This enables the prevention of harassment and a rapid response after it occurs.

[0979] "Dialogue history" refers to a record of conversations between users, which is saved in audio or text format.

[0980] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not its content constitutes harassment.

[0981] "Generation means" refers to a device or program that has the function of automatically creating visual information based on the dialogue history that has been determined to be harassment by the determination means.

[0982] "Presentation means" refers to a device or program that has the function of displaying generated visual information to the user and issuing warnings in real time.

[0983] "Visual information" refers to visual content such as videos and images presented to users, including details of harassment and measures to improve it.

[0984] "Real-time" refers to the instantaneous processing and response that occurs at the very moment the user's interaction is taking place.

[0985] "Advice" refers to information that suggests specific actions or behavioral changes that users should take to improve their behavior and address harassment.

[0986] The system for implementing this invention mainly consists of a server and a user terminal. The server receives dialogue history as input and is equipped with a determination means for determining whether or not the dialogue history constitutes harassment. This determination means uses a generative AI model to analyze the input text data and evaluate the likelihood of harassment.

[0987] The user's device uses a speech recognition API (e.g., Google Speech-to-Text) to convert audio data into text data. This text data is sent to a server and analyzed by a judgment mechanism. If harassment is detected as a result of the judgment, the server uses a generation mechanism to generate visual information that reflects the content of the harassment. This visual information is created using video generation software (e.g., Adobe Premiere Pro).

[0988] The generated visual information is sent to the user's device and displayed to the user in real time via a presentation tool. This allows the user to recognize if their words or actions may constitute harassment and receive specific advice for improvement.

[0989] For example, workplace conversations may be recorded, and a warning message such as "That way of speaking is inappropriate" may appear. In this case, the video may also include advice such as, "This way of speaking may hurt the other person. Next time, try rephrasing it like this."

[0990] An example of a prompt for a generative AI model is: "Analyze the following conversation and determine if it may be harassment. Conversation: 'Your way of doing things is terrible. Do it properly.'"

[0991] The flow of the specific processing in Application Example 3 will be explained using Figure 16.

[0992] Step 1:

[0993] The user's device accepts voice input. When the user speaks, the device's microphone records the audio. The recorded audio data is converted into text data using a speech recognition API (e.g., Google Speech-to-Text). This converted text data becomes the input for the next process.

[0994] Step 2:

[0995] The server receives text data sent from the user's terminal. The server analyzes this text data using a generative AI model to determine the possibility of harassment. In this analysis, prompt sentences are input into the generative AI model to determine whether or not harassment is present. The output of this step is the result of the harassment determination.

[0996] Step 3:

[0997] Based on the assessment results, the server generates visual information using a generation mechanism if harassment is detected. This visual information is created using video generation software (e.g., Adobe Premiere Pro). The generated visual information includes specific improvement measures and advice. This visual information serves as input for the next process.

[0998] Step 4:

[0999] The server transmits the generated visual information to the user's terminal. The user's terminal displays this visual information to the user in real time using a presentation mechanism. Through the presented visual information, the user can recognize if their words or actions may constitute harassment and receive specific advice for improvement.

[1000] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1001] "Example of form 1"

[1002] One embodiment of the present invention provides a system incorporating an emotion engine. This system receives a user's dialogue history as input and determines whether or not that dialogue history constitutes harassment. This determination is made by an emotion engine that recognizes the user's emotions. Specifically, it extracts emotions from the user's dialogue history and determines whether those emotions are related to harassment. For example, if a user expresses anger or dissatisfaction, it is determined that this may constitute harassment.

[1003] "Example of form 2"

[1004] Furthermore, the system provides a mechanism to adjust harassment detection based on emotions recognized by the emotion engine. Specifically, it strengthens harassment detection when a user's emotions are strongly expressed or when certain emotions are repeatedly expressed. For example, if a user repeatedly expresses anger, it may be determined to be a strong form of harassment, and this result is provided as feedback to the user.

[1005] "Example of form 3"

[1006] Furthermore, the system provides a means to generate videos that reflect the user's emotions based on the history of conversations that have been identified as harassment. Specifically, it automatically generates videos that reflect both the content of the harassment and the user's emotions. For example, if a user expresses anger, a video reflecting that anger is generated and presented to the user. This allows the user to visually understand the relationship between their emotions and the harassment.

[1007] The following describes the processing flow for each example of the form.

[1008] "Example of form 1"

[1009] Step 1: The system receives the user's interaction history.

[1010] Step 2: The emotion engine extracts the user's emotions from the conversation history.

[1011] Step 3: Determine whether the emotions extracted by the emotion engine are related to harassment.

[1012] "Example of form 2"

[1013] Step 1: Receive the user's conversation history and the emotion recognition results from the emotion engine.

[1014] Step 2: Strengthen the criteria for determining harassment when a user's emotions are strongly expressed or when certain emotions are repeatedly expressed.

[1015] Step 3: Provide feedback to the user regarding the enhanced harassment assessment results.

[1016] "Example of form 3"

[1017] Step 1: Receive the conversation history and user sentiment that were identified as harassment.

[1018] Step 2: Combine the content of the harassment with the user's emotions and automatically generate a video that reflects both.

[1019] Step 3: Present the generated video to the user.

[1020] (Example 1)

[1021] Next, we will describe Embodiment 1 of Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1022] In today's workplace, harassment is a serious problem, and early detection and countermeasures are essential. However, traditional methods often rely on subjective and inaccurate assessments of harassment. Furthermore, victims may face psychological resistance to reporting their experiences. This leads to challenges such as harassment being overlooked or responses being delayed.

[1023] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1024] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; an analysis means that analyzes the dialogue history using natural language processing technology; and an emotion recognition means that extracts emotions and determines whether or not the emotions are related to harassment. This makes it possible to objectively and quickly determine whether or not harassment has occurred and to respond promptly.

[1025] "Dialogue history" is a record of linguistic information exchanged between users during communication with others.

[1026] "Information processing means" refers to a device or program that has the function of analyzing input data and making a decision based on specific conditions.

[1027] "Generation means" refers to a device or program that has the function of creating and outputting new visual information based on input information.

[1028] "Natural language processing technology" is a technology that enables computers to understand and analyze human language.

[1029] "Analysis means" refers to a device or program that has the function of examining input data in detail and understanding its structure and meaning.

[1030] An "emotion recognition means" is a device or program that has the function of extracting emotions from input data and determining whether those emotions are related to specific conditions.

[1031] "Visual information" refers to information that can be recognized visually, such as images and videos.

[1032] As an embodiment for carrying out this invention, the automated harassment analysis system is configured as follows.

[1033] The user enters the conversation history using a terminal. This conversation history is a record of the linguistic information exchanged between the user and others. The entered conversation history is transmitted to the server via the internet.

[1034] The server analyzes the received dialogue history using information processing tools. These tools include Python's natural language processing libraries, NLTK and spaCy. This allows the server to analyze the dialogue history text, extract keywords, and understand the context.

[1035] Furthermore, the server uses emotion recognition to extract emotions from the conversation history. This emotion recognition uses emotion analysis tools such as Google Cloud's Natural Language API. The server then determines whether the extracted emotions are related to harassment.

[1036] For example, if a user enters a conversation history such as "My boss insults me every day," the server analyzes this history and extracts the keyword "insult." Next, the emotion engine identifies the emotion of "anger," and based on this information, it determines that there is a high probability of harassment.

[1037] An example of a prompt to input into a generative AI model would be: "Please determine if the following conversation history constitutes harassment: 'My boss insults me every day.'" Using this prompt, the generative AI model analyzes the content of the conversation history and determines whether or not harassment occurred.

[1038] The flow of the specific processing in Example 1 will be explained using Figure 17.

[1039] Step 1:

[1040] The user enters the conversation history using a terminal. The entered conversation history is in text format and includes linguistic information exchanged between the user and others. This data is transmitted to the server via the internet.

[1041] Step 2:

[1042] The server passes the received dialogue history to an information processing system. Here, the server uses Python's natural language processing libraries, NLTK and spaCy, to parse the dialogue history text. Specifically, the server extracts keywords from the text and performs data processing to understand the context. As a result of this analysis, important keywords and phrases contained in the dialogue history are output.

[1043] Step 3:

[1044] The server passes the analyzed data to the emotion recognition system. Here, the server uses emotion analysis tools such as Google Cloud's Natural Language API to extract emotions from the conversation history. Specifically, the server identifies emotions from words and phrases in the text and determines whether those emotions are related to harassment. As a result of this process, the extracted emotions and their relevance are output.

[1045] Step 4:

[1046] The server integrates the results obtained from the information processing and emotion recognition means to ultimately determine whether the dialogue history constitutes harassment. Specifically, the server evaluates the possibility of harassment based on the extracted keywords and emotions. As a result of this determination, it outputs whether or not the dialogue history constitutes harassment.

[1047] Step 5:

[1048] The server returns the judgment result to the user. The user's terminal is notified of the judgment result and displayed as a detailed report. Specifically, the user can check whether the conversation history constitutes harassment and take appropriate action if necessary.

[1049] (Application Example 1)

[1050] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1051] Workplace harassment is a problem that negatively impacts employees' mental health and workplace productivity. However, early detection of signs of harassment and implementation of appropriate countermeasures are difficult. In particular, there is a need for a system that can detect and warn about harassment in real time.

[1052] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1053] In this invention, the server includes a determination means that receives dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs information automatically generated based on the dialogue history determined to be harassment by the determination means; and a processing means that converts the user's voice data into text and analyzes it using natural language processing technology. This enables early detection of harassment in the workplace and prompt warnings.

[1054] "Dialogue history" refers to a record of conversations between users, which is saved as audio or text data.

[1055] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not its content constitutes harassment.

[1056] "Generation means" refers to a device or program that has the function of automatically creating and outputting relevant information and warnings based on the dialogue history that has been determined to be harassment by the determination means.

[1057] "Processing means" refers to a device or program that has the function of converting the user's voice data into text and analyzing that text using natural language processing technology.

[1058] "Analysis tool" refers to a device or program that has the function of extracting emotions from text data and evaluating the possibility of harassment.

[1059] "Notification means" refers to a device or program that has the function of issuing a warning to a user in the event of potential harassment.

[1060] A "portable information terminal" is an electronic device that is portable and capable of processing information, such as a smartphone or tablet.

[1061] The system that realizes this invention mainly consists of a server and a user's mobile information terminal. The server receives dialogue history as input and is equipped with a determination means for determining whether or not the dialogue history constitutes harassment. The determination means analyzes the dialogue history using natural language processing technology and evaluates the possibility of harassment. Specifically, it converts the user's voice data into text and performs analysis using a natural language processing library (e.g., spaCy, NLTK).

[1062] Furthermore, the server is equipped with analytical tools to extract emotions from text data and determine the possibility of harassment. Using emotion analysis libraries (e.g., TextBlob, VADER), it evaluates the user's emotions, and if anger or dissatisfaction is detected, it determines that there is a possibility of harassment.

[1063] The user's mobile device is equipped with a notification system to receive notifications from the server. If harassment is suspected, a warning will be displayed on the device to alert the user. This enables early detection and rapid response to harassment in the workplace.

[1064] As a concrete example, if a supervisor makes an inappropriate remark to a subordinate during a workplace conversation, the server analyzes the remark and, if it determines that it may constitute harassment, displays a warning on the subordinate's mobile device. An example of a prompt to be input to the generating AI model would be, "Analyze the content of this conversation and determine if it may constitute harassment."

[1065] The flow of a specific process in Application Example 1 will be explained using Figure 18.

[1066] Step 1:

[1067] The server receives audio data from the user's mobile device. This audio data is a recording of a conversation at work. The server uses speech recognition software to convert this audio data into text data.

[1068] Step 2:

[1069] The server analyzes the converted text data using natural language processing libraries (e.g., spaCy, NLTK). This analysis extracts grammatical structures and keywords from the text, generating foundational data for understanding the dialogue.

[1070] Step 3:

[1071] The server extracts emotions from the analyzed text data using a sentiment analysis library (e.g., TextBlob, VADER). In this step, it evaluates the emotional expressions contained in the text and determines whether negative emotions such as anger or frustration are present.

[1072] Step 4:

[1073] The server determines the possibility of harassment based on the results of sentiment analysis. Specifically, if negative emotions exceed a certain threshold, it determines that harassment is possible. This determination can be further improved using a generative AI model.

[1074] Step 5:

[1075] If the server determines that harassment may be occurring, it sends a warning to the user's mobile device. The device then notifies the user of this warning and prompts them to take action. This allows the user to recognize the risk of harassment in real time.

[1076] (Example 2)

[1077] Next, we will describe Example 2 of the morphological example. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1078] In today's workplace, harassment is a serious problem, and it is essential to properly detect it and take countermeasures. However, traditional methods have challenges in making accurate judgments because the determination of harassment is subjective and does not take into account changes in emotions.

[1079] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1080] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the information processing means; and an adjustment means that analyzes emotions and adjusts the harassment determination based on the analysis results. This makes it possible to make an objective and emotionally conscious determination of harassment.

[1081] "Dialogue history" refers to a record of conversations that users have had, and is data saved in audio or text format.

[1082] "Information processing means" refers to technical means for analyzing dialogue history and determining whether or not harassment occurred.

[1083] "Generation means" refers to a technical means for automatically creating and outputting visual information based on dialogue history that has been determined to be harassment.

[1084] "Adjustment measures" refer to technical means for analyzing user emotions and modifying or strengthening harassment judgments based on the results.

[1085] "Visual information" refers to visually presented information such as videos and images that reflect the content of the harassment.

[1086] "Emotion" refers to the user's psychological state and includes changes in feelings and emotions inferred from voice and text.

[1087] In an embodiment of this invention, a server, a terminal, and a user cooperate to operate the system. First, the user uses a terminal to record everyday conversations. The terminal is equipped with speech recognition software, which converts the speech data into text data. The converted text data is sent to the server as a conversation history.

[1088] The server analyzes the received dialogue history using information processing tools. This analysis employs natural language processing technology to detect specific keywords and phrases, thereby determining whether harassment has occurred. Furthermore, the server uses an emotion engine to analyze the user's emotions and adjusts the harassment determination based on the results. For example, if a user repeatedly expresses "anger," the server takes that emotion into account and strengthens its determination.

[1089] If harassment is detected, the server automatically generates visual information using a generation mechanism. This visual information is a video created using a generation AI model to represent a scenario corresponding to the content of the conversation. The generated video is sent to the user's device and presented to the user. By watching this video, the user can deepen their understanding of the nature of the harassment and its impact.

[1090] As a concrete example, let's look at a prompt message to input into a generation AI model: "If a user repeatedly expresses anger in conversations with their boss, please generate a video that points out the possibility of workplace harassment based on that conversation history." Using this prompt message, the server can generate an appropriate video and provide it to the user.

[1091] The flow of the specific processing in Example 2 will be explained using Figure 19.

[1092] Step 1:

[1093] The user records a conversation using a device. The device uses speech recognition software to convert the audio data into text data. The input for this step is audio data, and the output is text data. Specifically, the user records a conversation at work and then transcribes that recording into text.

[1094] Step 2:

[1095] The server receives text data sent from the terminal. The server uses information processing tools to analyze the text data and determine whether or not harassment has occurred. The input for this step is text data, and the output is the result of the harassment determination. Specifically, the server utilizes natural language processing technology to detect certain keywords and phrases.

[1096] Step 3:

[1097] The server uses an emotion engine to analyze the user's emotions from text data. Based on this analysis, it adjusts the harassment judgment. The input for this step is text data, and the output is the adjusted judgment result. Specifically, the server infers emotions from voice tone and text content, and strengthens or modifies the judgment.

[1098] Step 4:

[1099] If the server determines that harassment has occurred, it automatically generates visual information using a generation mechanism. The input for this step is the adjusted judgment result, and the output is visual information (video). Specifically, the server utilizes a generation AI model to create a video of a scenario that corresponds to the content of the conversation.

[1100] Step 5:

[1101] The server sends the generated visual information to the user's device and presents it to the user. The input in this step is visual information, and the output is the presentation to the user. Specifically, the user watches a video on their device to deepen their understanding of the content and impact of harassment.

[1102] (Application Example 2)

[1103] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1104] In modern workplaces and educational institutions, harassment is a serious problem, and early detection and countermeasures are essential. However, traditional methods can lead to delays in detecting harassment or failure to implement appropriate measures. In particular, there is a challenge in determining harassment while considering emotional changes.

[1105] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1106] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs a video that is automatically generated based on the dialogue history determined to be harassment by the determination means; and an adjustment means that recognizes the user's emotions and adjusts the harassment determination based on those emotions. This enables early detection of harassment and the implementation of appropriate countermeasures.

[1107] "Dialogue history" refers to a record of conversations between users, which is saved as text data.

[1108] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not it constitutes harassment.

[1109] "Generation means" refers to a device or program that has the function of automatically creating and outputting a video that points out the problematic aspects based on the history of a conversation that has been determined to be harassment.

[1110] "Adjustment measures" refer to devices or programs that recognize the user's emotions and have the function of strengthening or mitigating the determination of harassment based on those emotions.

[1111] "Emotions" refer to the psychological states and reactions that users exhibit during a conversation, including states such as anger, sadness, and joy.

[1112] The system for carrying out this invention includes a server and a user terminal. The server is equipped with a determination means for receiving and analyzing dialogue history. The determination means uses natural language processing technology to determine whether the dialogue history constitutes harassment. Specifically, it uses the Google Cloud Natural Language API to analyze text data and extract emotions and intentions.

[1113] The user's device, such as a smartphone or computer, is responsible for sending conversation history to the server. It also receives generated video from the server and presents it to the user. The generation method uses the Adobe Premiere Pro API to automatically generate video based on conversation history identified as harassment. This video includes content that points out the problem and prompts the user to take action.

[1114] Furthermore, the server is equipped with adjustment mechanisms to recognize user emotions in real time. The emotion engine enhances harassment detection when user emotions are strongly expressed or when certain emotions are repeatedly expressed. This enables more accurate detection.

[1115] For example, if workplace conversations are recorded and the user repeatedly expresses "anger," the server analyzes the conversation and determines that it may constitute workplace harassment. It then generates a video highlighting the harassment issues and notifies the user.

[1116] Examples of prompts for a generative AI model include the following:

[1117] "Analyze the following conversation history and determine if harassment is likely. Strengthen the assessment if the user's emotions are strongly expressed. If harassment is detected, generate a video highlighting the problem."

[1118] The flow of a specific process in Application Example 2 will be explained using Figure 20.

[1119] Step 1:

[1120] The user's terminal sends the conversation history as text data to the server. The input is the user's conversation history, and the output is the data transmission to the server. In this step, the user's terminal collects the conversation history and sends it to the server over the network.

[1121] Step 2:

[1122] The server analyzes the received dialogue history using the Google Cloud Natural Language API. The input is the text data of the dialogue history, and the output is the result of the sentiment and intent analysis. In this step, the server uses natural language processing techniques to extract sentiment and intent from the dialogue history.

[1123] Step 3:

[1124] The server determines the possibility of harassment based on the analysis results. The input is the analysis results of emotions and intentions, and the output is the harassment determination result. In this step, the server uses a determination means to evaluate the analysis results and determine whether or not harassment occurred.

[1125] Step 4:

[1126] If harassment is detected, the server generates a video highlighting the problem using the Adobe Premiere Pro API. The input is the harassment detection result and the dialogue history, and the output is the generated video. In this step, the server uses a generation mechanism to automatically create a video based on the dialogue history.

[1127] Step 5:

[1128] The server sends the generated video to the user's terminal. The input is the generated video, and the output is the transmission of the video to the user's terminal. In this step, the server transmits the video to the user's terminal over the network.

[1129] Step 6:

[1130] The user's device displays the received video to the user. The input is the video transmitted from the server, and the output is the video displayed to the user. In this step, the user's device plays the video and provides information to the user visually.

[1131] (Example 3)

[1132] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1133] In modern workplaces and society, harassment remains a serious problem, and preventing its occurrence is essential. However, it is difficult to determine what kind of words and actions constitute harassment in individual conversations and communications, and when harassment does occur, concrete corrective measures are required. Furthermore, there is a lack of visual means to understand the emotions associated with harassment, so these challenges need to be addressed.

[1134] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[1135] In this invention, the server includes: information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; data generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the information processing means; and data generation means that analyzes the emotions contained in the dialogue history and generates a video that reflects those emotions. This makes it possible for users to recognize the possibility that their words and actions constitute harassment and to visually understand concrete measures to improve their behavior.

[1136] "Dialogue history" refers to a record of conversations and communications between users, and is data saved in text format.

[1137] "Information processing means" refers to a device or program that has the function of analyzing input dialogue history and determining whether its content constitutes harassment.

[1138] "Data generation means" refers to a device or program that has the function of automatically creating images and other visual content based on the determination results of information processing means.

[1139] "Emotional analysis" is the process of identifying the user's emotions contained in the conversation history and evaluating the type and intensity of those emotions.

[1140] "Video" refers to content that includes visual information and is presented to users in the form of videos, animations, and other visual media.

[1141] A description of embodiments for carrying out this invention will be given.

[1142] The server receives the conversation history entered by the user. This conversation history is a record of the conversations and communications the user has had, and is provided in text format. The server analyzes this conversation history using natural language processing technology. Specifically, a common cloud-based natural language processing API can be used as the natural language processing software. This allows the server to understand the content of the conversation and detect signs of harassment.

[1143] Next, the server uses information processing means to determine whether the conversation history constitutes harassment. Based on this determination, the server uses data generation means to generate a video that reflects the content of the harassment. Video editing software can be used for this video generation.

[1144] Furthermore, the server uses sentiment analysis tools to analyze the user's emotions contained in the conversation history. Based on the results of the sentiment analysis, the server generates a video that reflects the user's emotions. This video is intended to help the user visually understand the relationship between their emotions and harassment.

[1145] As a concrete example, consider a scenario where a user enters a conversation history stating, "I used harsh words towards my subordinate during yesterday's meeting." The server analyzes this history and determines whether the harsh language constitutes harassment. If it does, the server generates a video based on the content that reflects the user's anger and presents it to the user.

[1146] An example of a prompt to be input to the generation AI model is, "Based on this dialogue history, determine whether harassment occurred and, if necessary, generate a video that reflects the emotions." The flow of the specific processing in Example 3 will be explained using Figure 21.

[1147] Step 1:

[1148] The user enters their conversation history. The user enters their conversation history into the system interface. The entered data is sent to the server in text format.

[1149] Step 2:

[1150] The server receives the dialogue history and performs natural language processing. The server analyzes the received text data using natural language processing software. This analysis helps understand the content and context of the dialogue and prepares it to detect signs of harassment. The input is text data, and the output is the analysis result.

[1151] Step 3:

[1152] The server determines whether harassment has occurred. Based on the results of natural language processing, the server determines whether harassment is present in the dialogue history. Specifically, it evaluates whether aggressive words or inappropriate expressions are included. The input is the analysis result, and the output is the determination result regarding the presence or absence of harassment.

[1153] Step 4:

[1154] If the server detects harassment, it prepares data for video generation. When harassment is detected, the server creates a script or template for video generation based on the details. The input is the detection result, and the output is the data for video generation.

[1155] Step 5:

[1156] The server analyzes the user's emotions. The server uses an emotion analysis tool to identify the user's emotions as they appear in the conversation history. The input is the conversation history, and the output is the result of the emotion analysis.

[1157] Step 6:

[1158] The server generates a video that reflects emotions. The server generates a video that reflects the content of the harassment and the user's emotions. Specifically, it uses video editing software to create visually represented content. The input is data for video generation and the results of emotion analysis, and the output is the generated video.

[1159] Step 7:

[1160] The server generates a video and presents it to the user. The server sends the generated video to the user, making it viewable. The input is the generated video, and the output is its presentation to the user.

[1161] (Application Example 3)

[1162] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1163] Harassment in the workplace and online communication is a problem that negatively impacts individual mental health and workplace productivity. However, there are insufficient systems to detect signs of harassment early and provide appropriate feedback. Furthermore, there is a lack of means for users to visually understand the connection between their own feelings and harassment and to take concrete actions toward improvement.

[1164] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[1165] In this invention, the server includes means for receiving a conversation history as input and determining whether the conversation history constitutes harassment; means for outputting a video automatically generated based on the conversation history determined to be harassment by the harassment determination means; and means for analyzing the user's emotions and generating a video that reflects those emotions. This allows the user to receive specific feedback to improve their communication style and prevent harassment from occurring.

[1166] "Dialogue history" refers to a record of conversations between users, which is saved in text format.

[1167] A "harassment determination tool" is a device that analyzes the input conversation history and determines whether or not its content constitutes harassment.

[1168] A "video generation method" is a system that automatically creates videos to provide visual feedback based on conversation history and user emotions that have been identified as harassment.

[1169] "User emotions" refers to the user's psychological state and emotional expression in their dialogue history, and this is analyzed and reflected in the video generation process.

[1170] A "generative AI model" refers to an algorithm or system that uses artificial intelligence technology to generate new content from input data.

[1171] "Feedback" refers to information and advice provided to users that helps improve communication styles.

[1172] The system for implementing this invention consists of a server and a user terminal. The server receives the conversation history transmitted from the user and analyzes its content using a harassment detection means. Specifically, it processes the text data using the natural language processing library spaCy and evaluates the possibility of harassment. Based on the evaluation results, it generates a video that reflects the content of the harassment using OpenAI's GPT generative AI model.

[1173] The video generation process includes analyzing the user's emotions. These emotions are extracted from the conversation history and visually represented by the video generation tool. Video editing libraries such as MoviePy are used for video generation. The generated video is sent to the user's device and presented to them. This allows the user to receive specific feedback to improve their communication style.

[1174] As a concrete example, consider email exchanges in the workplace. If a user sends an email saying, "This project isn't progressing at all, it's all your fault," the server analyzes this text and detects the possibility of harassment. An example of a prompt would be, "Evaluate whether the following text constitutes harassment: 'This project isn't progressing at all, it's all your fault'," which would be input into the GPT model. The generated video provides visual feedback to help the user reflect on their own behavior and make improvements.

[1175] The flow of the specific processing in Application Example 3 will be explained using Figure 22.

[1176] Step 1:

[1177] The user sends the conversation history to the server using their device. The input is the user's conversation history, sent to the server in text format. The server receives this data and prepares for the next processing step.

[1178] Step 2:

[1179] The server analyzes the received dialogue history using the natural language processing library spaCy. The input is the text data of the dialogue history, and the output is the analysis results indicating the possibility of harassment. The server detects signs of harassment by tokenizing the text and analyzing its grammatical structure.

[1180] Step 3:

[1181] Based on the analysis results, the server inputs prompt text into the OpenAI GPT, a generative AI model. The input is prompt text that reflects the analysis results, and the output is an evaluation of whether or not harassment occurred. Specifically, the server generates prompt text and sends it to the GPT model.

[1182] Step 4:

[1183] The server receives evaluation results from the GPT model and, if it determines that harassment has occurred, generates a video using a video generation method. The input is the evaluation results and user sentiment data, and the output is the generated video. The server uses a video editing library such as MoviePy to create a video that reflects the user's sentiment.

[1184] Step 5:

[1185] The server sends the generated video to the user's terminal and presents it to the user. The input is the generated video, and the output is the video playback on the user's terminal. The user watches this video and receives feedback to improve their communication style.

[1186] (Other examples)

[1187] Since this is the same as the specific processing described in the other embodiments of the first embodiment above, the explanation will be omitted.

[1188] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1189] The data generation model 58 is a form of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1190] Other examples of generative AI include Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) are some examples.

[1191] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1192] [Fourth Embodiment]

[1193] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1194] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1195] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1196] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1197] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1198] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1199] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1200] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1201] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1202] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1203] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1204] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1205] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.

[1206] "Example of form 1"

[1207] One embodiment of the present invention provides an automated harassment analysis system. This system receives a conversation history from a user as input and has a harassment determination means that determines whether or not the conversation history constitutes harassment. The harassment determination means analyzes the conversation history using natural language processing technology and determines whether or not harassment such as power harassment or sexual harassment has occurred.

[1208] "Example of form 2"

[1209] Furthermore, the system of the present invention includes a video generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the harassment determination means. The video generation means generates a video that reflects the content of the dialogue history determined to be harassment and presents it to the user. For example, if the dialogue history determined to be harassment is related to power harassment, it generates a video that points out the problems of power harassment.

[1210] "Example of form 3"

[1211] The system of the present invention can prevent harassment from occurring and provide concrete material for reflection and improvement in cases of harassment that have occurred. Specifically, by having the user input their conversation history into the system, the presence or absence of harassment is determined, and if harassment is found, a video reflecting the content is generated and presented. As a result, the user can recognize that their words and actions may constitute harassment and take concrete actions to improve them.

[1212] The following describes the processing flow for each example of the form.

[1213] "Example of form 1"

[1214] Step 1: The system receives the user's interaction history.

[1215] Step 2: The system's harassment detection mechanism analyzes the conversation history. This analysis uses natural language processing technology to determine whether or not harassment, such as power harassment or sexual harassment, has occurred.

[1216] "Example of form 2"

[1217] Step 1: The system receives the conversation history that has been determined to be harassment by the harassment detection method.

[1218] Step 2: The system's video generation mechanism automatically generates a video based on the conversation history that was identified as harassment. The video content reflects the content of the harassment.

[1219] Step 3: Present the generated video to the user.

[1220] "Example of form 3"

[1221] Step 1: The user enters their conversation history into the system.

[1222] Step 2: The system determines whether harassment occurred, and if so, generates a video reflecting the details of the harassment.

[1223] Step 3: Present the generated video to the user, allowing them to recognize that their words and actions may constitute harassment.

[1224] Step 4: Based on the presented video, users take specific actions to improve the harassment situation.

[1225] (Example 1)

[1226] Next, we will describe Embodiment 1 of Example Form 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1227] In today's workplace, harassment is a serious problem, and its detection and response are crucial. However, traditional methods have made it difficult to efficiently analyze dialogue history and accurately determine the presence or absence of harassment. Furthermore, there has been a lack of visual representations of harassment, making it difficult to understand and share the problem.

[1228] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1229] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; an analysis means that analyzes the dialogue history using natural language processing technology; and an evaluation means that evaluates the analysis results using a generative AI model and determines whether or not harassment has occurred. This enables efficient analysis of dialogue history and accurate determination of harassment.

[1230] "Dialogue history" refers to a record of conversations and messages exchanged between users, and is data saved in text format.

[1231] "Information processing means" refers to a device or program that has the function of analyzing input data and making decisions based on specific conditions.

[1232] "Generation means" refers to a device or program for automatically creating new visual information or content based on input information.

[1233] "Natural language processing technology" refers to technologies for understanding and analyzing human language using computers, and includes text tokenization and grammatical analysis.

[1234] "Analysis means" refers to a device or program for analyzing input data in detail and understanding its structure and meaning.

[1235] A "generative AI model" is a model that uses artificial intelligence technology to learn from data and is trained to perform a specific task.

[1236] "Evaluation means" refers to a device or program that makes a judgment based on analyzed data according to specific criteria.

[1237] To implement this invention, the user must first input their conversation history. The user uses their device to input the content of workplace conversations and messages in text format. The inputted conversation history is then sent from the device to the server.

[1238] The server uses Python's natural language processing libraries, NLTK and spaCy, to analyze the received dialogue history. Using this software, the server tokenizes the text data, tags it with parts of speech, and analyzes its grammatical structure.

[1239] Next, the server evaluates the analysis results using a generative AI model. This model is built using deep learning frameworks such as TensorFlow and PyTorch and is pre-trained on a dataset related to harassment. The model detects elements of power harassment and sexual harassment in the dialogue history and determines whether or not they exist.

[1240] For example, if a user enters a conversation history such as "My boss forces me to work unreasonable overtime every day," the server analyzes this data and uses a generative AI model to determine that "it contains elements of power harassment." This result is sent back to the terminal and displayed to the user.

[1241] An example of a prompt message is, "Please determine if this conversation history contains elements of harassment." By using this prompt message, the server can efficiently analyze the conversation history and accurately determine whether or not harassment has occurred.

[1242] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1243] Step 1:

[1244] Users input conversation history in text format using a terminal. This input data includes the content of workplace conversations and messages. The entered conversation history is sent from the terminal to the server.

[1245] Step 2:

[1246] The server uses natural language processing libraries such as NLTK and spaCy to analyze the dialogue history received from the terminal. The server tokenizes the input text data, tags it with parts of speech, and analyzes its grammatical structure. This analysis extracts structural information from the text.

[1247] Step 3:

[1248] The server inputs the analyzed data into a generating AI model. This model is built using TensorFlow and PyTorch and has been pre-trained on a dataset related to harassment. Based on the input data, the model detects elements of power harassment and sexual harassment and determines whether or not they exist. This determination result is then output.

[1249] Step 4:

[1250] The server returns the judgment result obtained from the generated AI model to the terminal. The judgment result indicates whether or not the conversation history contains elements of harassment.

[1251] Step 5:

[1252] The terminal displays the judgment results received from the server to the user. For example, a message such as "This conversation contains elements of power harassment" might be displayed. This allows the user to understand the content of the conversation history and take necessary measures.

[1253] (Application Example 1)

[1254] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1255] Workplace harassment is a serious problem that negatively impacts employees' mental health and workplace productivity. However, early detection and appropriate response to signs of harassment are difficult. Traditional methods often rely on victims reporting the issue themselves, and problems may go undetected until they become severe. An effective system is needed to improve this situation and prevent harassment proactively.

[1256] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1257] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a conversion means that converts speech into text; and an analysis means that analyzes the text converted by the conversion means and detects signs of harassment. This makes it possible to detect signs of harassment in the workplace environment in real time and notify administrators.

[1258] "Dialogue history" refers to a record of conversations between users, which is saved as audio or text data.

[1259] "Determination means" refers to a device or program that analyzes the input dialogue history and determines whether or not it constitutes harassment.

[1260] "Generation means" refers to a device or program that has the function of automatically creating relevant information and content based on the dialogue history that has been determined to be harassment by the determination means.

[1261] "Conversion means" refers to a device or program that has the function of converting audio data into text data.

[1262] "Analysis means" refers to a device or program that has the function of analyzing text data and detecting signs of harassment.

[1263] A "notification means" is a device or program that has the function of sending warnings or information to administrators or relevant parties based on signs of harassment detected by an analysis means.

[1264] The system for implementing this invention mainly consists of a server and a user terminal. The server receives dialogue history as input and is equipped with a determination means for determining whether or not the dialogue history constitutes harassment. This determination means analyzes the dialogue history using natural language processing technology to determine whether or not harassment has occurred. Specifically, it uses the Google Cloud Natural Language API to analyze text data.

[1265] The user terminal is equipped with a conversion mechanism for converting speech to text. This conversion mechanism uses the Google Cloud Speech-to-Text API to convert speech data into text data. The converted text data is sent to a server and analyzed by a determination mechanism.

[1266] If the server detects signs of harassment as a result of its analysis, it will send a warning to the administrator using a notification system. This notification system uses Firebase Cloud Messaging to provide real-time notifications.

[1267] For example, if the statement "You're always useless" is made during a conversation at work, this statement may be judged as potentially being harassment, and a notification will be sent to the manager. An example of a prompt to the generating AI model would be, "Please determine whether the following conversation constitutes harassment: 'You're always useless.'"

[1268] In this way, the system can detect signs of harassment in the workplace in real time, enabling a rapid response.

[1269] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1270] Step 1:

[1271] The user terminal captures workplace conversations as audio data. This audio data is collected in real time through the user terminal's microphone.

[1272] Step 2:

[1273] The user's device converts the acquired audio data into text data using the Google Cloud Speech-to-Text API. This conversion process analyzes the audio signal and generates the corresponding text. The input is audio data, and the output is text data.

[1274] Step 3:

[1275] The user's terminal sends the converted text data to the server. The server analyzes the received text data using the Google Cloud Natural Language API. The analysis evaluates linguistic features within the text and detects signs of harassment. The input is text data, and the output is a determination of whether or not harassment occurred.

[1276] Step 4:

[1277] If the server detects signs of harassment based on its analysis, it will send a notification to the administrator using Firebase Cloud Messaging. The notification will include detailed information such as the nature and time of the harassment. The input is the analysis result, and the output is the notification to the administrator.

[1278] Step 5:

[1279] The administrator will take necessary actions based on the received notification. This includes confirming with relevant parties and conducting further investigations. The administrator's actions will be based on the system output.

[1280] (Example 2)

[1281] Next, we will describe Example 2 of the morphological example. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1282] In today's workplace, harassment is a serious problem, and early detection and appropriate response are essential. However, traditional methods have the drawback of subjective harassment assessments, making it difficult to implement appropriate education and countermeasures. Therefore, a system is needed that objectively assesses harassment and provides educational visual information based on that assessment.

[1283] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1284] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a visual information generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the information processing means; and a presentation means that presents the visual information generated by the visual information generation means to the user. This makes it possible to objectively determine harassment and provide educational visual information based on its content.

[1285] "Dialogue history" refers to a record of conversations between users, which is saved in text format.

[1286] "Information processing means" refers to a device or program that has the function of analyzing input data and making decisions based on specific conditions.

[1287] "Visual information generation means" refers to a device or program that has the function of automatically creating visual content based on analysis results.

[1288] "Presentation means" refers to a device or program for displaying or providing generated visual information to a user.

[1289] "Harassment" refers to inappropriate words or actions towards others in the workplace or other environments, and is an act that causes mental or physical distress.

[1290] A description of embodiments for carrying out this invention will be given.

[1291] The server receives the conversation history provided by the user. The user inputs the conversation history in text format via their terminal and sends it to the server. This conversation history is used as data for harassment assessment.

[1292] The server analyzes the dialogue history using a generative AI model. Specifically, it uses natural language processing techniques to determine whether the content of the dialogue constitutes harassment. In this process, a generative AI model with excellent natural language processing capabilities is used as the general-purpose model. The server inputs the prompt "Please determine whether this dialogue history constitutes harassment" into the generative AI model.

[1293] If harassment is detected, the server automatically generates visual information based on the detection result using a visual information generation system. This visual information includes educational content and aims to inform users about the problems and countermeasures against harassment. General video editing software may be used to generate the visual information.

[1294] For example, if a user enters a conversation history stating that "a superior excessively reprimanded a subordinate," the server will determine this to be harassment. Next, the server generates and presents visual information on the theme of "the impact of inappropriate behavior in the workplace and countermeasures." This visual information includes specific examples and countermeasures, allowing the user to deepen their understanding.

[1295] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1296] Step 1:

[1297] The user uses a terminal to input the conversation history in text format and sends it to the server. The input data is a record of the conversation between users and serves as the basis for determining whether harassment occurred. The server receives this conversation history and prepares for the next processing step.

[1298] Step 2:

[1299] The server inputs the received dialogue history into the generating AI model. Specifically, it uses the prompt "Determine whether this dialogue history constitutes harassment" to have the generating AI model analyze the dialogue history. The generating AI model uses natural language processing techniques to analyze the content of the dialogue and determine whether or not harassment is present. The output of this step is the result of the harassment determination.

[1300] Step 3:

[1301] The server receives the judgment result from the generating AI model, and if it determines that harassment has occurred, it generates visual information using a visual information generation means. Specifically, based on the judgment result, it creates visual information that includes educational content. Video editing software may be used to generate the visual information. The output of this step is the generated visual information.

[1302] Step 4:

[1303] The server presents the generated visual information to the user. The user can view the visual information through their device and learn about harassment issues and countermeasures. The output of this step is the visual information presented to the user.

[1304] (Application Example 2)

[1305] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1306] Workplace harassment is a serious problem that harms employees' mental health and leads to decreased productivity. However, detecting and preventing harassment is difficult, and prompt and appropriate measures are necessary, especially in situations where real-time response is required. Traditional methods often only identify harassment incidents after they have occurred, making immediate response difficult.

[1307] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1308] In this invention, the server includes a determination means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a generation means that outputs a video automatically generated based on the dialogue history determined to be harassment by the determination means; and a presentation means that presents the video generated by the generation means to the user. This makes it possible to detect harassment in the workplace environment in real time and immediately generate and present educational videos.

[1309] "Dialogue history" refers to data that records the content of a conversation and is saved in audio or text format.

[1310] A "determination means" is a device or program that has the function of analyzing input data and making a judgment based on specific conditions.

[1311] "Generation means" refers to a device or program that has the function of automatically creating new content based on input information.

[1312] "Presentation means" refers to a device or program that has the function of providing generated information or content to the user visually or audibly.

[1313] A "generative AI model" is an algorithm or program that uses artificial intelligence technology to generate new information or content from data.

[1314] A "prompt" is an instruction or question given to a generative AI model to obtain a specific output.

[1315] The system for implementing this invention mainly consists of a server and a user terminal. The server receives dialogue history as input and converts the audio data into text using a speech recognition API (e.g., Google Cloud Speech-to-Text) to determine whether the dialogue history constitutes harassment. Next, it analyzes the converted data using a natural language processing library (e.g., spaCy) to determine the possibility of harassment.

[1316] Based on the determined results, the server uses a generation AI model (e.g., OpenAI's GPT-4) to generate a video that points out the harassment issues and proposes solutions. In this process, the generation AI model receives prompt text to obtain specific output. An example of a prompt text would be, "This conversation has been determined to be power harassment. Please generate a video that points out the power harassment issues and proposes solutions."

[1317] The generated video is displayed on the user's device. The user's device is a smartphone or smart glasses, and its role is to visually provide the generated video to the user. This allows the user to understand the harassment issue in real time and take appropriate action.

[1318] For example, if a workplace conversation includes a statement like, "You're always so slow, you need to work harder," the server will identify this as harassment and input a prompt into a generation AI model to create an educational video. This video is then immediately presented to the relevant parties via the user's terminal.

[1319] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1320] Step 1:

[1321] The server receives audio data sent from the user's terminal. This audio data is a recording of a workplace conversation. The server converts this audio data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is audio data, and the output is text data.

[1322] Step 2:

[1323] The server analyzes text data using a natural language processing library (e.g., spaCy). The purpose of the analysis is to determine whether there are signs of harassment in the text. The input is text data, and the output is a judgment result indicating the possibility of harassment. Specifically, it analyzes keywords and context within the text and scores the likelihood of harassment.

[1324] Step 3:

[1325] Based on the judgment result, the server inputs a prompt message into the generating AI model (e.g., OpenAI's GPT-4). The prompt message is: "This conversation has been determined to be harassment. Please generate a video that points out the problems with harassment and proposes solutions." The input is the judgment result and the prompt message, and the output is the content of the video generated by the generating AI model. Specifically, the generating AI model analyzes the prompt message and generates appropriate video content.

[1326] Step 4:

[1327] The server uses a video generation tool to create the actual video based on the generated video content. The input is the video content generated by the AI ​​model, and the output is the completed video file. Specifically, the video generation tool converts text-based content into a visual video.

[1328] Step 5:

[1329] The server sends the completed video file to the user's terminal. The user's terminal then displays this video to the user. The input is a video file, and the output is the display of the video to the user. Specifically, the user's terminal plays the video and provides the user with visual information.

[1330] (Example 3)

[1331] Next, we will describe Embodiment 3 of Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1332] In today's workplace, harassment remains a serious problem, and there is a need to prevent it and to implement concrete corrective measures when it occurs. However, traditional methods have the challenge of making it difficult to recognize harassment and obtain specific feedback for improvement.

[1333] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[1334] In this invention, the server includes an information processing means that receives a dialogue history as input and determines whether or not the dialogue history constitutes harassment; a visual information generation means that outputs visual information automatically generated based on the dialogue history determined to be harassment by the information processing means; and a presentation means that presents the generated visual information to the user. This enables the user to recognize the possibility that their words and actions constitute harassment and to take concrete corrective measures.

[1335] "Dialogue history" refers to a record of conversations a user has had in the past, and is data entered in text format.

[1336] "Information processing means" refers to technical means for analyzing input dialogue history and determining whether or not harassment has occurred.

[1337] "Visual information generation means" refers to a technical means that automatically generates visual information reflecting the content of harassment based on a conversation history that has been determined to be harassment.

[1338] A "presentation means" is a technical means that provides generated visual information to a user, enabling the user to visually confirm it.

[1339] "Harassment" refers to inappropriate words or actions towards others in the workplace or in society, and particularly includes inappropriate behavior and sexual harassment in the workplace.

[1340] A description of embodiments for carrying out this invention will be given.

[1341] The user first inputs their conversation history into the system via their terminal. The conversation history is in text format and includes past conversations. The user copies and pastes the conversation history into the input form on their terminal and clicks the "Start Analysis" button.

[1342] The server receives the conversation history sent by the user. The received data is preprocessed for natural language processing. Specifically, text cleaning and tokenization are performed. The server inputs the preprocessed text data into a natural language processing engine, which is an AI model for generating conversations. This engine analyzes the conversation content and determines whether or not harassment has occurred. The result of the determination is output as a score indicating whether or not harassment is likely.

[1343] Based on the assessment results, the...

Claims

[Claim 1] Equipped with a processor, The aforementioned processor, The system receives the dialogue history as input, analyzes the keywords and context within the text of the dialogue history, and scores the likelihood of power harassment to determine whether or not the dialogue history constitutes power harassment. If the aforementioned dialogue history is determined to be power harassment, a prompt message is generated to instruct the system to create a video that includes a statement indicating that the dialogue history has been determined to be power harassment, and that points out the problems of power harassment and proposes solutions for improvement. By inputting the aforementioned prompt text into a generating AI model that obtains output by inputting the aforementioned prompt text, a video is automatically generated that points out the problems of power harassment and proposes solutions to improve those problems. system.

Citation Information

Patent Citations

  • Harmful act detection system and method

    JP2020123204A

  • Persona chatbot control method and system

    JP2022180282A

  • Action control system

    JP2025022813A