System

A system with data collection, voice recognition, and generative AI analysis provides real-time detection and prevention of sexual harassment, enhancing workplace safety by improving detection accuracy through continuous learning.

JP2026030670APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024133654
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Current systems lack the ability to detect signs of sexual harassment in real time and provide appropriate feedback, leading to discomfort for victims and damage to a company's credibility.

Method used

A system that includes data collection, voice or text data capture, voice recognition, analysis using a generative AI model to detect sexual harassment remarks, alert generation, and notification to users, with continuous learning to improve detection accuracy.

Benefits of technology

Enables real-time detection and prevention of sexual harassment, maintaining a healthy communication environment by providing immediate warnings and improving detection accuracy over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026030670000001_ABST
    Figure 2026030670000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: a data collection means; a voice or text data capture means; a voice recognition means for converting the captured voice to text data; an analyzing means for analyzing the text data and using a generative AI model to detect possible sexual harassment; an alert generating means for generating a warning message if sexual harassment is detected; and a notifying means for notifying a user of the warning message.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, sexual harassment has become a major issue in companies and organizations, but many people still commit sexual harassment without even realizing it. This not only causes discomfort for the victim, but can also damage the company's credibility. While effective preventative measures and countermeasures are needed to address this issue, there are currently no adequate solutions. In particular, there is a lack of systems that can detect signs of sexual harassment in real time and provide appropriate feedback. This means that in order to maintain a safe and healthy communication environment, a system that can detect sexually harassing remarks in real time and issue appropriate warnings is needed. [Means for solving the problem]

[0005] To solve the above-mentioned problems, the present invention provides the following means. It proposes a system including a data collection means, a voice or text data capture means, a voice recognition means for converting the captured voice data into text data, an analysis means using a generative AI model to analyze the text data and detect possible sexual harassment remarks, an alert generation means for generating a warning message when sexual harassment remarks are detected, and a notification means for notifying the user of the warning message. This system can detect sexual harassment remarks in real time and immediately issue a warning, thereby preventing unconscious sexual harassment acts by users. Furthermore, by continuously learning new sexual harassment case data using a database, the system's detection accuracy can be improved, making it possible to provide more advanced countermeasures over time.

[0006] "Data collection means" refers to the part of the system that has the function of importing past sexual harassment case data and storing, updating, and managing it in a database.

[0007] "Voice or text data capture means" refers to a part of a system that has the ability to collect a user's voice conversations or text messages in real time using a microphone or chat app, etc.

[0008] "Speech recognition means" refers to a technique for analyzing captured voice data and converting it into text data, and generally uses a voice recognition engine.

[0009] The "analysis means" is part of a system that uses a generative AI model to analyze text data and determine whether a comment may be sexually harassing.

[0010] A "generative AI model" refers to an artificial intelligence model that learns from data on past sexual harassment cases and analyzes new text data to evaluate the possibility of sexual harassment remarks.

[0011] The "alert generation means" is a part of the system that creates and provides a warning message to the user when a sexually harassing remark is detected.

[0012] A "notification means" is a part of the system that has the function of visually and audibly communicating generated warning messages to the user. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] Specific Embodiments of the Invention

[0035] The present invention relates to a system that monitors conversations and chats within a company, detects possible sexual harassment comments, and issues a warning in real time. Specific embodiments for implementing this system will be described below.

[0036] Program processing

[0037] Data collection

[0038] The server imports data on past sexual harassment cases and stores it in a centrally managed database, allowing the system to learn in a way that responds to the diversity of situations and word usage.

[0039] Voice and text data capture

[0040] The device (smartphone or PC) captures voice or text data from a microphone or chat app. In the case of voice data, the next step is to convert it into text data.

[0041] Voice Recognition

[0042] The device uses voice recognition technology to convert the captured voice data into text, putting the data in a format that can be analyzed by the generative AI model.

[0043] Text data analysis

[0044] The device analyzes the text data using a generative AI model that has been trained on data from past sexual harassment cases and has the ability to score whether a comment is likely to be sexual harassment.

[0045] Detecting sexual harassment remarks

[0046] The device determines whether the remark is likely to be sexually harassing based on the scores and results obtained from the generative AI model. If the score exceeds a certain threshold during this process, the device proceeds to the next step.

[0047] Generate alerts

[0048] If the device determines that a remark is likely to be sexually harassing, it will generate a warning message, which will be customized depending on the situation and will alert the person making the remark.

[0049] User Notification

[0050] The terminal notifies the user of the generated warning message both visually (e.g., on the screen) and audibly, allowing the user to immediately become aware of their own remarks and make any necessary corrections.

[0051] Specific examples

[0052] Scenario 1: Remote conference operation

[0053] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[0054] 2. The device captures this speech and converts it into text using speech recognition technology.

[0055] 3. The device uses a generative AI model to analyze the comment as potentially sexual harassment.

[0056] 4. The device generates a warning saying, "Caution: potential sexual harassment."

[0057] 5. The device displays a warning message on User A's screen.

[0058] 6. User A realizes what he or she said and corrects himself or herself by saying, "I'm sorry, I made an inappropriate comment."

[0059] Scenario 2: Chat operation

[0060] 1. User B sends a message in a chat room saying, "New employee XX is so beautiful."

[0061] 2. The device captures this text message.

[0062] 3. The device uses a generative AI model to analyze the message and assess its potential for sexual harassment.

[0063] 4. The device generates an alert saying, "Be careful, there is a possibility of sexual harassment."

[0064] 5. The device displays an alert on User B's chat screen.

[0065] 6. User B immediately cancels the message and sends a correction message saying, "I apologize for my inappropriate comment."

[0066] In this way, by introducing this system, it is possible to prevent sexually harassing remarks in everyday communication, thereby maintaining and improving a healthy workplace environment.

[0067] The processing flow will be explained below.

[0068] Step 1:

[0069] The server imports past sexual harassment case data and stores it in a database, which is used to train generative AI models.

[0070] Step 2:

[0071] The device captures voice or text data from a microphone or chat app. In the case of voice data, it needs to be converted into text format for further processing.

[0072] Step 3:

[0073] When the device captures voice data, it uses a speech recognition engine to convert the speech into text data, which can then be analyzed by a generative AI model.

[0074] Step 4:

[0075] The device then inputs the converted text data into a generative AI model that analyzes potential sexual harassment comments. This AI model was trained based on the data collected in step 1.

[0076] Step 5:

[0077] The device evaluates the likelihood of a remark being sexually harassing based on the output from the generated AI model, and if it is determined to be sexually harassing with a high probability, it activates an alert generation method.

[0078] Step 6:

[0079] The device generates a warning message. Specifically, it creates a message indicating the possibility of sexual harassment according to a set format.

[0080] Step 7:

[0081] The device notifies the user of the generated warning message. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification.

[0082] Step 8:

[0083] The user receives a warning message and is made aware that their comments or messages may constitute sexual harassment. The user can correct their comments or take appropriate action.

[0084] Step 9:

[0085] The server continuously updates the generative AI model by collecting newly detected sexual harassment case data and adding it to the database, thereby improving the model's accuracy and effectiveness over time.

[0086] Example 1

[0087] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0088] Preventing sexual harassment in the modern workplace is a difficult task. Traditional methods require time for sexual harassment comments to be reported externally, so a prompt response is required to maintain a healthy workplace. Real-time monitoring and warnings are necessary, especially with the increase in remote work and online chat, but current systems tend to be slow to respond.

[0089] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0090] In this invention, the server includes a means for collecting data, a means for capturing voice or text data, a speech recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, a means for inputting a prompt sentence into the generative AI model, an alert generation means for generating a warning message when sexual harassment remarks are detected, and a notification means for notifying the user of the warning message. This enables real-time collection and analysis of data, and rapid generation and notification of warning messages.

[0091] "Means for collecting data" refers to systems or devices for importing and centrally managing data on past sexual harassment cases and related information.

[0092] "Audio or text data capture means" means a system or device for capturing and storing conversations or text messages in real time.

[0093] "Speech recognition means" refers to the technology or software used to convert captured voice data into text.

[0094] "Analytical means using generative AI models" refers to systems or devices that use AI models trained based on past cases of sexual harassment to analyze text data and evaluate the likelihood of sexual harassment remarks.

[0095] A "means for inputting prompts to a generative AI model" is an input interface or software for providing specific instructions or questions to an AI model.

[0096] An "alert generation means" is a system or device that generates a warning message when a sexually harassing remark is detected.

[0097] The "notification means" is a system or device for visually and audibly notifying the user of the generated warning message.

[0098] MODE FOR CARRYING OUT THE INVENTION

[0099] The present invention relates to a system that monitors conversations and chats within a company, detects possible sexual harassment comments, and issues a warning in real time. Specific embodiments for implementing this system will be described below.

[0100] The server is primarily used to collect data on past sexual harassment cases and store it in a centralized database. This system can use MySQL or PostgreSQL as the database. The collected data is used as training data for the generative AI model, and data preprocessing is performed to improve the model's accuracy. Preprocessing includes filling in missing values ​​and normalizing the data.

[0101] The device (smartphone or personal computer) is used to capture voice or text data in real time from remote meetings or chat apps. Voice can be recorded through a microphone and saved in WAV or MP3 format. Text data can also be obtained using the chat app's API.

[0102] For speech recognition, we use voice recognition technologies such as Google Cloud Speech-to-Text and Amazon Transcribe. This converts the audio data into text, which is then formatted for easy analysis by the generative AI model. The converted text data is then stripped of unnecessary spaces and special characters to prepare it for analysis.

[0103] Text data is analyzed in real time using an analytical method that uses a generative AI model (e.g., OpenAI GPT-4). During this process, a prompt is input into the generative AI model, which evaluates whether the remark constitutes sexual harassment. Including specific examples in the prompt allows for a more accurate judgment.

[0104] Some common prompts are:

[0105] Example prompt 1:

[0106] "Please determine whether the following statements constitute sexual harassment:

[0107] "You're so cute, how about we have dinner sometime?"

[0108] Example prompt 2:

[0109] "Please determine whether the following statements constitute sexual harassment:

[0110] "New employee XX is really beautiful."

[0111] If the analysis results of the generative AI model exceed a threshold, a warning message is generated as an alert generation means. This message includes specific instructions and warns the speaker that the comment is inappropriate.

[0112] The warning message is notified to the user through a notification means, which is visual (for example, a screen display) and audio, so that the user can immediately pay attention to the message.

[0113] As a result, the system can detect sexually harassing remarks in real time and take prompt action to maintain a healthy work environment.

[0114] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0115] Program processing flow

[0116] Step 1:

[0117] The server collects data on past sexual harassment cases and stores it in a centrally managed database.

[0118] Input: Historical sexual harassment case data from a company database.

[0119] Data processing: Preprocessing of collected data (filling in missing values ​​and normalizing data).

[0120] Output: Data stored as a training dataset for a generative AI model.

[0121] How it works: The server uses SQL queries to extract sexual harassment cases from the company's internal database over the past five years, imports the data into a MySQL database, and performs preprocessing such as missing value imputation and normalization to prepare a training dataset.

[0122] Step 2:

[0123] The device captures voice or text data in real time from conversations and chat apps.

[0124] Input: Audio data from a remote meeting or text messages from a chat app.

[0125] Data processing: Audio data is saved in WAV or MP3 format. Text data is saved in a temporary folder.

[0126] Output: The captured audio or text data.

[0127] Specific operation: During remote meetings, the device records audio data in real time in WAV format through the microphone, and uses the chat app's API to capture all messages sent in text format and save them in a temporary folder.

[0128] Step 3:

[0129] The terminal uses voice recognition technology to convert the captured voice data into text.

[0130] Input: Captured audio data (WAV or MP3 format).

[0131] Data processing: Convert speech to text using Google Cloud Speech-to-Text or Amazon Transcribe. Remove unnecessary spaces and special characters.

[0132] Output: The converted text data.

[0133] Specific operation: The device calls the Google Cloud Speech-to-Text API to convert the WAV audio file into a string, then removes unnecessary spaces and special characters from the converted text data in preparation for analysis.

[0134] Step 4:

[0135] The device uses a generative AI model to analyze the text data.

[0136] Input: The converted text data.

[0137] Data processing: A prompt was generated and input into a generative AI model (OpenAI GPT-4). This prompt included instructions to evaluate whether the remark constituted sexual harassment.

[0138] Output: Scoring and analysis results from the generative AI model.

[0139] Specific operation: The device generates the following prompt sentence and inputs it into the AI ​​model:

[0140] "Please determine whether the following statements constitute sexual harassment:

[0141] "You're so cute, how about we have dinner sometime?"

[0142] Obtain the scoring and analysis results returned by the AI ​​model and format the data.

[0143] Step 5:

[0144] The device determines whether a comment is likely to be sexual harassment based on the analysis results obtained from the generative AI model.

[0145] Input: Scoring and analysis results from the AI ​​model.

[0146] Data processing: Determine whether the score exceeds a predefined threshold.

[0147] Output: Triggers an alert generator when a threshold is exceeded.

[0148] Specific operation: The device checks whether the score returned by the AI ​​model is 80 or above (pre-set threshold), and if so, proceeds to the next alert generation step.

[0149] Step 6:

[0150] The terminal generates a warning message if a sexually harassing remark is detected.

[0151] Input: Speech data with scores above a threshold.

[0152] Data processing: Generation of customized warning messages.

[0153] Output: The warning message generated.

[0154] Specific behavior: If the device determines that a comment is sexually harassing, it will generate a warning message stating, "This comment may be inappropriate. Please correct it."

[0155] Step 7:

[0156] The terminal notifies the user of the generated warning message.

[0157] Input: The generated warning message.

[0158] Data processing: Preparation of visual and audio notifications.

[0159] Output: Notification to the user.

[0160] Specific operation: The device will display a warning message on the user's screen and simultaneously issue an audio alert to draw attention to the statement, allowing the user to immediately pay attention to the statement and make corrections.

[0161] (Application example 1)

[0162] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0163] In conventional workplaces, there was no system that could detect sexual harassment remarks in real time and issue a warning. As a result, when a victim received unpleasant remarks, it was difficult to take immediate action, making it difficult to maintain a healthy work environment. There was also a risk that employees who might make sexually harassing remarks would cause problems without realizing that they were making such remarks.

[0164] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0165] In this invention, the server includes a data collection means, a voice or text data capture means, a voice recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, an alert generation means for generating a warning message when sexual harassment remarks are detected, a notification means for notifying a user of the warning message, and a real-time notification means for capturing data in real time using a smart device and immediately notifying a user of the warning message based on the results of the data analysis. This makes it possible to monitor conversations and chats in the office in real time, immediately detect sexual harassment remarks, issue a warning to employees, and encourage them to take prompt action.

[0166] "Data collection means" refers to a means for collecting information such as voice data and text data and saving it in a format required for subsequent processing.

[0167] "Audio or text data capture means" means a device or software for capturing audio or text data in real time.

[0168] "Speech recognition means" is a technology that converts captured voice data into text data.

[0169] The "analysis means" is a means for analyzing collected text data and evaluating and classifying it based on specific conditions.

[0170] A "generative AI model" is an artificial intelligence model that learns from past data and makes predictions and classifications for new data.

[0171] An "alert generator" is a device or software for generating a warning message when a particular condition is met.

[0172] The "notification means" is a means for transmitting the generated warning message to the user.

[0173] A "smart device" is a mobile terminal equipped with advanced computing power and communication functions.

[0174] "Real-time data capture" is the act of collecting data in synchronization with real time.

[0175] "Real-time notification means" refers to a means for transmitting a warning message to the user in a timely manner according to the analysis results.

[0176] This invention relates to a system for monitoring conversations and chats in a workplace and detecting and issuing warnings about sexual harassment comments in real time. Specific embodiments for implementing this system will be described below.

[0177] The system includes a data collection means, a voice or text data capture means, a voice recognition means, an analysis means, an alert generation means, a notification means, real-time data capture using a smart device, and a real-time notification means.

[0178] Server processing

[0179] The server processes data using the following hardware and software:

[0180] Hardware: Server

[0181] Software: Database (MySQL), Generative AI model (OpenAI GPT-3)

[0182] As a means of collecting data, past sexual harassment cases are stored in a database, and the generative AI model learns from this data. New sexual harassment cases are continually added to the database, and it is updated with the latest data.

[0183] What the device is doing

[0184] The terminal uses the following hardware and software:

[0185] Hardware: Smartphone, microphone

[0186] Software: Speech recognition API (Google Speech-to-Text)

[0187] The smartphone's microphone is used to capture voice or text data, and office conversations are captured in real time. The voice data is converted into text data using a speech recognition API, and the converted text data is sent to the server.

[0188] Server text data analysis

[0189] The server uses a generative AI model to analyze text data and score it for potential sexual harassment. This analysis is performed in real time based on a historical database, and a warning message is generated if the score exceeds a certain threshold.

[0190] Alerting and Notifications

[0191] As an alert generation means, the server generates a warning message and immediately notifies the user through a notification means, which is visually and audibly displayed on the smartphone in real time.

[0192] Specific examples

[0193] For example, if User A says, "You're so cute," the smartphone's microphone captures this remark. The speech data is converted into text data using a speech recognition API (Google Speech-to-Text). This text data is sent to a server and analyzed by a generative AI model (OpenAI GPT-3). If the remark is determined to be highly likely to be sexual harassment, a warning message is generated and a message stating, "Please be careful as this may be sexual harassment," is immediately displayed on User A's smartphone.

[0194] Prompt Sentence Examples

[0195] For example, the input prompt for a generative AI model might look like this:

[0196] "New employee XX is so beautiful."

[0197] Based on this prompt, the generative AI model evaluates the statement and issues a warning if necessary.

[0198] As described above, the present invention provides a system for preventing sexual harassment remarks in a workplace environment and maintaining healthy communication.

[0199] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0200] Step 1:

[0201] The device performs audio capture. To obtain audio data, the device (smartphone or PC) constantly monitors the ambient sound using a microphone. The input is the user's voice, and the output is audio data.

[0202] Step 2:

[0203] The device performs speech recognition and converts the speech data into text data. The acquired speech data is converted into text using a speech recognition API (Google Speech-to-Text). The input is speech data, and the output is text data.

[0204] Step 3:

[0205] The terminal sends the converted text data to the server. The text data is sent to the server via the HTTPS protocol. The input is text data, and the output is data sent to the server.

[0206] Step 4:

[0207] The server analyzes the text data using a generative AI model. Based on the received text data, the server uses a generative AI model (OpenAI GPT-3) to score the likelihood of sexual harassment remarks. At this time, a prompt sentence may be used as input. The input is text data, and the output is a score of the likelihood of sexual harassment remarks.

[0208] Step 5:

[0209] The server generates a warning message based on the analysis results. If the scoring result by the generative AI model exceeds a certain threshold, the server generates a customized warning message. The input is the likelihood score of sexual harassment remarks, and the output is a warning message.

[0210] Step 6:

[0211] The server sends a warning message to the terminal. The generated warning message is sent to the terminal via HTTPS protocol and notified in real time. The input is the warning message and the output is the data sent to the terminal.

[0212] Step 7:

[0213] The terminal displays a warning message to the user using a notification means. The terminal notifies the user of the received warning message visually or audibly. Specifically, the warning message is displayed on the screen and an audio notification is played. The input is the sent warning message, and the output is a notification message to the user.

[0214] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0215] Specific Embodiments of the Invention

[0216] This invention relates to a system that monitors conversations and chats, detects possible sexual harassment comments, and issues a warning in real time, and further includes a function to recognize the user's emotions and provide appropriate feedback. Specific embodiments for implementing this system are described below.

[0217] Program processing

[0218] Data collection

[0219] The server imports past sexual harassment case data and stores it in a database, which is used to train the generative AI model and improve the accuracy of the emotion engine.

[0220] Voice and text data capture

[0221] The device (smartphone or PC) captures voice or text data in real time from a microphone or chat app. In the case of voice data, this is then converted into text format.

[0222] Voice Recognition

[0223] The device uses voice recognition technology to convert the captured voice data into text data, which is then used for subsequent analysis by generative AI models and emotion engines.

[0224] Text data analysis

[0225] The device analyzes the text data and assesses the likelihood of sexual harassment using a generative AI model, which is trained on data from past sexual harassment cases.

[0226] emotion recognition

[0227] The device uses an emotion engine to recognize emotions from the user's facial expressions, tone of voice, etc. Emotion data is combined with the results of text data analysis to generate the final warning message.

[0228] Detecting sexual harassment remarks

[0229] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If the remark is judged to be sexually harassing with a high probability, it will activate an alert generation method.

[0230] Generate alerts

[0231] If the device determines that a remark is likely to be sexually harassing, it generates a warning message that is customized to the situation by incorporating emotional data provided by the emotion engine.

[0232] User Notification

[0233] The device notifies the user of the generated warning message. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification. The user is immediately aware of what they said and can make any necessary corrections.

[0234] Specific examples

[0235] Scenario 1: Remote conference operation

[0236] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[0237] 2. The device captures this speech and converts it into text using speech recognition technology.

[0238] 3. The device uses a generative AI model to analyze the comment and determine that it may be sexual harassment.

[0239] 4. The device uses the emotion engine to recognize User A's emotions, for example, determining whether User A is speaking in a joking or serious tone.

[0240] 5. The device generates a warning saying, "Be careful, this may be sexual harassment," and customizes the tone and content based on data from the emotion engine.

[0241] 6. The device displays a warning message on User A's screen.

[0242] 7. User A recognizes what he said and corrects himself by saying, "I'm sorry, I made an inappropriate comment."

[0243] Scenario 2: Chat operation

[0244] 1. User B sends a message in a chat room saying, "New employee XX is so beautiful."

[0245] 2. The device captures this text message.

[0246] 3. The device uses a generative AI model to analyze the message and determine that it may be sexually harassing.

[0247] 4. The device uses an emotion engine to recognize User B's emotions. For example, it determines whether User B is speaking in a friendly manner or teasing the user.

[0248] 5. The device generates a warning saying, "Be careful, this may be sexual harassment," and customizes the tone and content based on data from the emotion engine.

[0249] 6. The device displays an alert on User B's chat screen.

[0250] 7. User B immediately cancels the message and sends a correction message saying, "I apologize for my inappropriate comment."

[0251] This system enables the detection and feedback of sexually harassing remarks taking into account the user's emotions, which is expected to lead to more accurate and effective communication.

[0252] The processing flow will be explained below.

[0253] Step 1:

[0254] The server imports historical sexual harassment case data and stores it in a database, which is used to train generative AI models and emotion engines.

[0255] Step 2:

[0256] The device captures voice or text data from the microphone or chat app, allowing users to collect their conversations and messages in real time.

[0257] Step 3:

[0258] When the device receives voice data, it uses a speech recognition engine to convert the voice into text data, which prepares the data in a format that can be analyzed.

[0259] Step 4:

[0260] The device inputs the text data into a generative AI model that analyzes the possibility of sexual harassment remarks. The generative AI model is trained from data on past sexual harassment cases.

[0261] Step 5:

[0262] The device captures the user's facial expressions and tone of voice, and uses an emotion engine to recognize the user's emotions, allowing it to understand the user's mental state and emotions.

[0263] Step 6:

[0264] The device combines the analysis results from the generative AI model with the emotional data from the emotion engine to assess the likelihood of sexual harassment. By combining the two sets of data, a more accurate judgment can be made.

[0265] Step 7:

[0266] If the device determines that a comment is likely to be sexually harassing, it will generate a warning message, dynamically adjusting the content and tone of the message based on the emotional data.

[0267] Step 8:

[0268] The terminal will notify the user of the generated warning message visually and audibly. The user will be alerted by a pop-up display and / or an audio notification.

[0269] Step 9:

[0270] Users will receive a warning message and be made aware that their comments or actions may constitute sexual harassment. Users can correct or retract their comments as necessary.

[0271] Step 10:

[0272] The server collects newly detected sexual harassment cases and emotion data and adds them to the database, which continuously updates the generative AI model and emotion engine, improving the accuracy of the entire system.

[0273] Example 2

[0274] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0275] Sexually harassing remarks have become a problem in modern remote conference and chat environments. However, there are few systems that can detect them in real time and take immediate countermeasures. It has also been pointed out that there is a lack of systems that not only analyze the content of remarks but also recognize the speaker's emotions and provide appropriate feedback. The purpose of this invention is to solve these problems and provide a system that effectively detects and warns about sexually harassing remarks.

[0276] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0277] In this invention, the server includes a data collection means, a voice or text data capture means, a voice recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment expressions, an emotion recognition means for recognizing a user's emotions and integrating them with the analysis results, an alert generation means for generating a warning message when possible sexual harassment expressions are detected, and a notification means for notifying the user of the warning message. This enables effective detection and warning of sexual harassment remarks in real time and appropriate feedback that takes the user's emotions into consideration.

[0278] "Data Collection Measures" refers to the devices and processes used to import and store historical sexual harassment case data in a database.

[0279] "Voice or text data capture means" refers to devices and processes that obtain voice or text data through a microphone or chat app.

[0280] "Speech recognition means" refers to the technology and processes that convert captured voice data into text data.

[0281] A "generative AI model" refers to an artificial intelligence model that is trained based on data from past sexual harassment cases and analyzes text data to detect possible sexual harassment remarks.

[0282] "Analysis means" refers to the devices and processes that use generative AI models to analyze text data and assess potential sexual harassment comments.

[0283] "Emotion recognition means" refers to technology and processes for recognizing emotions from a user's facial expressions and tone of voice and generating emotion data.

[0284] "Alert generation means" refers to a device and process that generates a warning message when it determines that a sexually harassing remark is likely.

[0285] "Notification means" refers to devices and processes for notifying a user of generated warning messages.

[0286] MODE FOR CARRYING OUT THE INVENTION

[0287] The present invention is a system that monitors conversations and chats in real time, detects possible sexual harassment comments, and issues a warning. It also includes a function to recognize the user's emotions and provide appropriate feedback. A specific embodiment of this system will be described below.

[0288] Data collection methods

[0289] The server imports past sexual harassment case data and stores it in a database. This data contains a wide variety of sexual harassment cases and is used to train the generative AI model and emotion recognition engine. The server implements data collection and storage processes using Python and SQL, which significantly improves the accuracy and effectiveness of the system.

[0290] A means of capturing voice or text data

[0291] The device (smartphone or PC) captures voice or text data in real time through a microphone or chat app. Voice data is converted into text using speech recognition technology such as the Google Cloud Speech-to-Text API. This allows the spoken content to be obtained as text and used for subsequent processing.

[0292] Voice recognition means

[0293] The device uses voice recognition technology to convert the captured audio data into text data, using services such as Google Cloud Speech-to-Text, which is then immediately stored internally and passed on to the next analysis step.

[0294] Analytical tools using generative AI models

[0295] The device analyzes text data using a generative AI model to assess the likelihood of sexual harassment. This AI model is trained based on past data on sexual harassment cases and can evaluate the content of statements with high accuracy. For example, it uses advanced generative AI models such as BERT and GPT-3.

[0296] emotion recognition means

[0297] The device uses an emotion recognition engine to recognize emotions from the user's facial expressions and tone of voice, utilizing the Microsoft Azure Emotion API, among other tools. By acquiring emotion data, the device can more accurately understand the intention and nuance of what is being said.

[0298] Alert generation method

[0299] The device evaluates the likelihood of sexual harassment based on the analysis results of the generative AI model and data from the emotion recognition engine. If there is a high probability that the remark is sexual harassment, it automatically generates a warning message. This message is customized to an appropriate tone and content based on the emotion data.

[0300] Notification means

[0301] The terminal generates a warning message and notifies the user. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification. In particular, if the message is urgent, it can immediately draw the user's attention.

[0302] Specific operation example

[0303] Scenario 1: Remote conference operation

[0304] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[0305] 2. The device captures this speech and converts it into text using speech recognition technology (powered by the Google Cloud Speech-to-Text API).

[0306] 3. The device analyzes the comment using a generative AI model (BERT, GPT-3, etc.) and determines that it may be sexual harassment.

[0307] 4. The device uses an emotion engine (such as Azure Emotion API) to recognize User A's emotions.

[0308] 5. The device generates a warning saying, "Beware of potential sexual harassment," and customizes the tone and content accordingly.

[0309] 6. The device pops up a warning message on User A's screen.

[0310] 7. User A corrects himself by saying, "I'm sorry, I made an inappropriate comment."

[0311] Examples of prompt statements

[0312] During a remote meeting, User A says, "You're so cute. How about we go out for dinner sometime?" Capture this remark and analyze it using a sexual harassment remark detection system. Explain the process of generating an appropriate warning message based on the analysis results of the generative AI model and emotion recognition results, and notifying User A.

[0313] This system can detect sexual harassment remarks in real time and provide appropriate warning messages that take emotions into account. This configuration can improve effective communication.

[0314] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0315] Step 1:

[0316] The server imports data on past sexual harassment cases and stores it in a database. The input is data on past sexual harassment cases, and the output is the data stored in the database. A data collection program using Python and SQL reads and saves the initial data.

[0317] What happens: The server periodically checks the data source (e.g., CSV file or external API) and imports any new data it finds. The imported data is then stored in the database.

[0318] Step 2:

[0319] The device captures voice or text data in real time through a microphone or chat app. The input is the user's speech data, and the output is the captured voice file or text data. The data is acquired in real time using JavaScript or Python libraries.

[0320] What it does: When a conversation starts, an app installed on the device activates the microphone to capture voice data in real time, and also captures text messages directly from chat apps.

[0321] Step 3:

[0322] The device converts the voice data into text data using the Google Cloud Speech-to-Text API or similar. The input is the captured voice data, and the output is the converted text data. The device calls a speech recognition API to convert the voice to text.

[0323] Specific operation: The device sends voice data to the API, receives the voice recognition results, stores the results in internal memory, and passes them on to the next process.

[0324] Step 4:

[0325] The device analyzes text data using a generative AI model to assess the likelihood of sexually harassing comments. The input is text data, and the output is a probability score of sexually harassing comments. AI models such as BERT and GPT-3 are used.

[0326] How it works: The device inputs text data into the generative AI model, which then outputs a score indicating the likelihood of the comment being sexually harassing. This score is then used in the next step.

[0327] Step 5:

[0328] The device uses an emotion recognition engine to recognize the user's emotions. The input is the user's facial expression and tone of voice, and the output is the recognized emotion data. It uses the Microsoft Azure Emotion API, etc.

[0329] Specific operation: The device uses the camera and microphone connected to the device to capture the user's facial expressions and tone of voice, and sends them to the emotion recognition API. The obtained emotion data is then stored in the device's internal memory.

[0330] Step 6:

[0331] The device evaluates the probability of sexual harassment remarks based on the analysis results of the generated AI model and emotion recognition data. The input is the probability score of sexual harassment remarks and emotion data, and the output is an alert generation flag. A threshold is set, and if it is exceeded, an alert generation flag is set.

[0332] Specific operation: The analysis results of the text data (probability score) are integrated with the emotion data, and if a certain threshold is exceeded, an alert generation flag is set and the next process is carried out.

[0333] Step 7:

[0334] Generates a warning message when the device has an alert generation flag. The input is the alert generation flag and emotion data, and the output is the warning message. Generates the warning message using a Python script.

[0335] Specific operation: Prepare a warning message template and generate a customized message based on the emotion data. This message is notified to the user in the next step.

[0336] Step 8:

[0337] The terminal notifies the user of the generated warning message. The input is the warning message and the output is the notification to the user. The notification methods are visual (screen popup) and audible (voice notification).

[0338] Specific behavior: The device will display the generated warning message as a pop-up on the screen and, if necessary, provide an audio notification. This is implemented using JavaScript or the Android and iOS notification APIs.

[0339] (Application example 2)

[0340] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0341] In recent years, sexual harassment has become a problem in the workplace and online communities. However, systems that can detect sexual harassment in real time and issue appropriate warnings are not yet widely available. Furthermore, there is a need for technology that provides more accurate and effective feedback by taking user emotions into account. The present invention aims to provide a sexual harassment detection system that solves these problems.

[0342] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0343] In this invention, the server includes a data collection means, a means for capturing voice or text data, a means for converting the captured voice data into text data, a means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, a means for generating a warning message when a sexual harassment remark is detected, a means for notifying the user of the warning message, and an emotion recognition means for recognizing the user's emotion. This makes it possible to detect sexual harassment remarks in real time and issue an appropriate warning. Furthermore, by recognizing the user's emotion, it is possible to customize the warning message according to the situation and provide more effective feedback.

[0344] "Data collection means" is a function that imports data on past sexual harassment cases and stores it in a database.

[0345] A "voice or text data capture means" is a device that acquires voice or text data in real time from a microphone or chat app.

[0346] A "voice recognition means" is a technique or device that converts captured voice data into text data.

[0347] "Analysis tools" are technologies or devices that use generative AI models to analyze text data and assess the likelihood of sexual harassment statements.

[0348] A "generative AI model" is an artificial intelligence model trained based on data on past sexual harassment cases.

[0349] The "alert generation means" is a function that generates a warning message when a sexual harassment remark is detected.

[0350] The "notification means" is a function that notifies the user of the generated warning message visually or audibly.

[0351] The "emotion recognition means" is a function that recognizes the user's emotions from facial expressions, tone of voice, etc.

[0352] A specific embodiment for implementing this invention will be described below. This system monitors comments in real time when users communicate via voice or chat, and issues an immediate warning if possible sexual harassment comments are detected. Furthermore, it recognizes the user's emotions and appropriately customizes the warning message, thereby achieving effective feedback.

[0353] Data collection

[0354] The server imports past sexual harassment case data and stores it in a database. This data is used to train the generative AI model and also helps improve the accuracy of the emotion engine. For example, a dataset could include sexual harassment cases collected from various industries.

[0355] Voice and text data capture

[0356] The device captures voice or text data in real time from a microphone or chat app. In the case of voice data, it needs to be converted into text. For example, a smartphone microphone can be used to capture voice and a chat app can accept input as a text message.

[0357] Voice Recognition

[0358] The device uses voice recognition technology to convert the captured voice data into text data, using Google's voice recognition service as an example. This converted text data is then used for subsequent analysis by generative AI models and sentiment engines.

[0359] Text data analysis

[0360] The device uses a generative AI model to analyze the captured text data and evaluate the possibility of sexual harassment. The generative AI model is trained based on data from past sexual harassment cases and evaluates the content of comments with high accuracy. For example, a comment such as "You're so cute. How about going out for dinner sometime?" is judged to be potentially sexual harassment.

[0361] emotion recognition

[0362] The device uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice. The emotion data is combined with the results of text data analysis to generate the final warning message. For example, it can determine whether the user is speaking in a joking or serious tone.

[0363] Detecting sexual harassment remarks

[0364] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If the remark is judged to be sexually harassing with a high probability, it will activate an alert generation method.

[0365] Alert generation and notification

[0366] If the device determines that a remark is likely to be sexual harassment, it generates a warning message. The generated warning message incorporates emotional data provided by the emotion engine and is customized according to the situation. The device then notifies the user of the warning message visually or audibly using a notification means. For example, a warning such as "Be careful as this may be sexual harassment" may be displayed on the user's screen.

[0367] Specific examples

[0368] 1. Scenario 1: Remote Meeting Operation

[0369] If a user says, "You're so cute. Let's have dinner together sometime," during a remote meeting, the remark is captured and converted into text. A generative AI model then detects potential sexual harassment, and an emotion recognition engine recognizes the user's playful tone. As a result, a warning message appears on the user's screen saying, "Be careful, this may be sexual harassment."

[0370] 2. Scenario 2: Chat operation

[0371] If a user sends a message in a chat room saying, "New employee XX is so beautiful," this text message is captured and the generative AI model detects potential sexual harassment. The emotion recognition engine determines whether the user is speaking in a friendly or teasing manner, and a warning message is displayed on the user's chat screen saying, "Be careful, this may be sexual harassment."

[0372] Prompt Sentence Examples

[0373] Example: "You're so cute, how about we have dinner sometime?"

[0374] Example: "The new employee, Ms. XX, is very beautiful."

[0375] In this way, this system enables the detection of sexually harassing remarks and feedback in real time, taking into account the user's emotions, which is expected to lead to more accurate and effective communication.

[0376] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0377] Step 1: Data collection

[0378] The server imports past sexual harassment case data and stores it in a database. This is primarily used to improve the accuracy of the generative AI model and emotion recognition engine. Specifically, it analyzes past sexual harassment remark data collected from within the company or from public datasets and registers it in the database. The input is past sexual harassment case data, and the output is the data stored in the database.

[0379] Step 2: Capture audio and text data

[0380] The device captures voice and text data in real time from a microphone or chat app. In the case of voice data, it must be converted into text format. Specifically, the device picks up what the user says with a microphone and captures the voice data, or receives text data directly from a chat app. The input is the user's real-time voice or chat message, and the output is the captured voice data or text data.

[0381] Step 3: Voice Recognition

[0382] The device uses speech recognition technology to convert the captured voice data into text data. Here, Google's speech recognition API is used for speech-to-text conversion. The input is the captured voice data, and the output is text data.

[0383] Step 4: Analyze the text data

[0384] The device analyzes text data using a generative AI model to assess the likelihood of sexual harassment. The model is a trained generative AI model that evaluates the content of statements with high accuracy based on past data on sexual harassment cases. The input is text data, and the output is an assessment result indicating the likelihood of sexual harassment.

[0385] Step 5: Emotion Recognition

[0386] The device uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice. For example, it determines whether the user is speaking in a joking or serious tone. The input is the user's facial expression data and tone of voice data, and the output is the recognized emotion data.

[0387] Step 6: Detect sexual harassment comments

[0388] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If it determines with a high probability that the remark was sexually harassing, it activates an alert generation mechanism. The inputs are the analysis results and emotional data, and the output is the evaluation result of the remark being sexually harassing.

[0389] Step 7: Generate an alert

[0390] The device generates a warning message if it determines that a remark is likely to be sexual harassment. The generated warning message is customized by incorporating emotional data provided by an emotion recognition engine. The input is the evaluation result of the sexual harassment remark and the emotional data, and the output is a customized warning message.

[0391] Step 8: Notify users

[0392] The terminal notifies the user of the generated warning message visually or audibly. For example, it notifies the user of the warning by a screen pop-up or a voice notification. The input is the customized warning message, and the output is the warning message notified to the user.

[0393] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0394] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0395] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0396] [Second embodiment]

[0397] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0398] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0399] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0400] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0401] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0402] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0403] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0404] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0405] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0406] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0407] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0408] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0409] Specific Embodiments of the Invention

[0410] The present invention relates to a system that monitors conversations and chats within a company, detects possible sexual harassment comments, and issues a warning in real time. Specific embodiments for implementing this system will be described below.

[0411] Program processing

[0412] Data collection

[0413] The server imports data on past sexual harassment cases and stores it in a centrally managed database, allowing the system to learn in a way that responds to the diversity of situations and word usage.

[0414] Voice and text data capture

[0415] The device (smartphone or PC) captures voice or text data from a microphone or chat app. In the case of voice data, the next step is to convert it into text data.

[0416] Voice Recognition

[0417] The device uses voice recognition technology to convert the captured voice data into text, putting the data in a format that can be analyzed by the generative AI model.

[0418] Text data analysis

[0419] The device analyzes the text data using a generative AI model that has been trained on data from past sexual harassment cases and has the ability to score whether a comment is likely to be sexual harassment.

[0420] Detecting sexual harassment remarks

[0421] The device determines whether the remark is likely to be sexually harassing based on the scores and results obtained from the generative AI model. If the score exceeds a certain threshold during this process, the device proceeds to the next step.

[0422] Generate alerts

[0423] If the device determines that a remark is likely to be sexually harassing, it will generate a warning message, which will be customized depending on the situation and will alert the person making the remark.

[0424] User Notification

[0425] The terminal notifies the user of the generated warning message both visually (e.g., on the screen) and audibly, allowing the user to immediately become aware of their own remarks and make any necessary corrections.

[0426] Specific examples

[0427] Scenario 1: Remote conference operation

[0428] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[0429] 2. The device captures this speech and converts it into text using speech recognition technology.

[0430] 3. The device uses a generative AI model to analyze the comment as potentially sexual harassment.

[0431] 4. The device generates a warning saying, "Caution: potential sexual harassment."

[0432] 5. The device displays a warning message on User A's screen.

[0433] 6. User A realizes what he or she said and corrects himself or herself by saying, "I'm sorry, I made an inappropriate comment."

[0434] Scenario 2: Chat operation

[0435] 1. User B sends a message in a chat room saying, "New employee XX is so beautiful."

[0436] 2. The device captures this text message.

[0437] 3. The device uses a generative AI model to analyze the message and assess its potential for sexual harassment.

[0438] 4. The device generates an alert saying, "Be careful, there is a possibility of sexual harassment."

[0439] 5. The device displays an alert on User B's chat screen.

[0440] 6. User B immediately cancels the message and sends a correction message saying, "I apologize for my inappropriate comment."

[0441] In this way, by introducing this system, it is possible to prevent sexually harassing remarks in everyday communication, thereby maintaining and improving a healthy workplace environment.

[0442] The processing flow will be explained below.

[0443] Step 1:

[0444] The server imports past sexual harassment case data and stores it in a database, which is used to train generative AI models.

[0445] Step 2:

[0446] The device captures voice or text data from a microphone or chat app. In the case of voice data, it needs to be converted into text format for further processing.

[0447] Step 3:

[0448] When the device captures voice data, it uses a speech recognition engine to convert the speech into text data, which can then be analyzed by a generative AI model.

[0449] Step 4:

[0450] The device then inputs the converted text data into a generative AI model that analyzes potential sexual harassment comments. This AI model was trained based on the data collected in step 1.

[0451] Step 5:

[0452] The device evaluates the likelihood of a remark being sexually harassing based on the output from the generated AI model, and if it is determined to be sexually harassing with a high probability, it activates an alert generation method.

[0453] Step 6:

[0454] The device generates a warning message. Specifically, it creates a message indicating the possibility of sexual harassment according to a set format.

[0455] Step 7:

[0456] The device notifies the user of the generated warning message. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification.

[0457] Step 8:

[0458] The user receives a warning message and is made aware that their comments or messages may constitute sexual harassment. The user can correct their comments or take appropriate action.

[0459] Step 9:

[0460] The server continuously updates the generative AI model by collecting newly detected sexual harassment case data and adding it to the database, thereby improving the model's accuracy and effectiveness over time.

[0461] Example 1

[0462] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0463] Preventing sexual harassment in the modern workplace is a difficult task. Traditional methods require time for sexual harassment comments to be reported externally, so a prompt response is required to maintain a healthy workplace. Real-time monitoring and warnings are necessary, especially with the increase in remote work and online chat, but current systems tend to be slow to respond.

[0464] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0465] In this invention, the server includes a means for collecting data, a means for capturing voice or text data, a speech recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, a means for inputting a prompt sentence into the generative AI model, an alert generation means for generating a warning message when sexual harassment remarks are detected, and a notification means for notifying the user of the warning message. This enables real-time collection and analysis of data, and rapid generation and notification of warning messages.

[0466] "Means for collecting data" refers to systems or devices for importing and centrally managing data on past sexual harassment cases and related information.

[0467] "Audio or text data capture means" means a system or device for capturing and storing conversations or text messages in real time.

[0468] "Speech recognition means" refers to the technology or software used to convert captured voice data into text.

[0469] "Analytical means using generative AI models" refers to systems or devices that use AI models trained based on past cases of sexual harassment to analyze text data and evaluate the likelihood of sexual harassment remarks.

[0470] A "means for inputting prompts to a generative AI model" is an input interface or software for providing specific instructions or questions to an AI model.

[0471] An "alert generation means" is a system or device that generates a warning message when a sexually harassing remark is detected.

[0472] The "notification means" is a system or device for visually and audibly notifying the user of the generated warning message.

[0473] MODE FOR CARRYING OUT THE INVENTION

[0474] The present invention relates to a system that monitors conversations and chats within a company, detects possible sexual harassment comments, and issues a warning in real time. Specific embodiments for implementing this system will be described below.

[0475] The server is primarily used to collect data on past sexual harassment cases and store it in a centralized database. This system can use MySQL or PostgreSQL as the database. The collected data is used as training data for the generative AI model, and data preprocessing is performed to improve the model's accuracy. Preprocessing includes filling in missing values ​​and normalizing the data.

[0476] The device (smartphone or personal computer) is used to capture voice or text data in real time from remote meetings or chat apps. Voice can be recorded through a microphone and saved in WAV or MP3 format. Text data can also be obtained using the chat app's API.

[0477] For speech recognition, we use voice recognition technologies such as Google Cloud Speech-to-Text and Amazon Transcribe. This converts the audio data into text, which is then formatted for easy analysis by the generative AI model. The converted text data is then stripped of unnecessary spaces and special characters to prepare it for analysis.

[0478] Text data is analyzed in real time using an analytical method that uses a generative AI model (e.g., OpenAI GPT-4). During this process, a prompt is input into the generative AI model, which evaluates whether the remark constitutes sexual harassment. Including specific examples in the prompt allows for a more accurate judgment.

[0479] Some common prompts are:

[0480] Example prompt 1:

[0481] "Please determine whether the following statements constitute sexual harassment:

[0482] "You're so cute, how about we have dinner sometime?"

[0483] Example prompt 2:

[0484] "Please determine whether the following statements constitute sexual harassment:

[0485] "New employee XX is really beautiful."

[0486] If the analysis results of the generative AI model exceed a threshold, a warning message is generated as an alert generation means. This message includes specific instructions and warns the speaker that the comment is inappropriate.

[0487] The warning message is notified to the user through a notification means, which is visual (for example, a screen display) and audio, so that the user can immediately pay attention to the message.

[0488] As a result, the system can detect sexually harassing remarks in real time and take prompt action to maintain a healthy work environment.

[0489] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0490] Program processing flow

[0491] Step 1:

[0492] The server collects data on past sexual harassment cases and stores it in a centrally managed database.

[0493] Input: Historical sexual harassment case data from a company database.

[0494] Data processing: Preprocessing of collected data (filling in missing values ​​and normalizing data).

[0495] Output: Data stored as a training dataset for a generative AI model.

[0496] How it works: The server uses SQL queries to extract sexual harassment cases from the company's internal database over the past five years, imports the data into a MySQL database, and performs preprocessing such as missing value imputation and normalization to prepare a training dataset.

[0497] Step 2:

[0498] The device captures voice or text data in real time from conversations and chat apps.

[0499] Input: Audio data from a remote meeting or text messages from a chat app.

[0500] Data processing: Audio data is saved in WAV or MP3 format. Text data is saved in a temporary folder.

[0501] Output: The captured audio or text data.

[0502] Specific operation: During remote meetings, the device records audio data in real time in WAV format through the microphone, and uses the chat app's API to capture all messages sent in text format and save them in a temporary folder.

[0503] Step 3:

[0504] The terminal uses voice recognition technology to convert the captured voice data into text.

[0505] Input: Captured audio data (WAV or MP3 format).

[0506] Data processing: Convert speech to text using Google Cloud Speech-to-Text or Amazon Transcribe. Remove unnecessary spaces and special characters.

[0507] Output: The converted text data.

[0508] Specific operation: The device calls the Google Cloud Speech-to-Text API to convert the WAV audio file into a string, then removes unnecessary spaces and special characters from the converted text data in preparation for analysis.

[0509] Step 4:

[0510] The device uses a generative AI model to analyze the text data.

[0511] Input: The converted text data.

[0512] Data processing: A prompt was generated and input into a generative AI model (OpenAI GPT-4). This prompt included instructions to evaluate whether the remark constituted sexual harassment.

[0513] Output: Scoring and analysis results from the generative AI model.

[0514] Specific operation: The device generates the following prompt sentence and inputs it into the AI ​​model:

[0515] "Please determine whether the following statements constitute sexual harassment:

[0516] "You're so cute, how about we have dinner sometime?"

[0517] Obtain the scoring and analysis results returned by the AI ​​model and format the data.

[0518] Step 5:

[0519] The device determines whether a comment is likely to be sexual harassment based on the analysis results obtained from the generative AI model.

[0520] Input: Scoring and analysis results from the AI ​​model.

[0521] Data processing: Determine whether the score exceeds a predefined threshold.

[0522] Output: Triggers an alert generator when a threshold is exceeded.

[0523] Specific operation: The device checks whether the score returned by the AI ​​model is 80 or above (pre-set threshold), and if so, proceeds to the next alert generation step.

[0524] Step 6:

[0525] The terminal generates a warning message if a sexually harassing remark is detected.

[0526] Input: Speech data with scores above a threshold.

[0527] Data processing: Generation of customized warning messages.

[0528] Output: The warning message generated.

[0529] Specific behavior: If the device determines that a comment is sexually harassing, it will generate a warning message stating, "This comment may be inappropriate. Please correct it."

[0530] Step 7:

[0531] The terminal notifies the user of the generated warning message.

[0532] Input: The generated warning message.

[0533] Data processing: Preparation of visual and audio notifications.

[0534] Output: Notification to the user.

[0535] Specific operation: The device will display a warning message on the user's screen and simultaneously issue an audio alert to draw attention to the statement, allowing the user to immediately pay attention to the statement and make corrections.

[0536] (Application example 1)

[0537] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0538] In conventional workplaces, there was no system that could detect sexual harassment remarks in real time and issue a warning. As a result, when a victim received unpleasant remarks, it was difficult to take immediate action, making it difficult to maintain a healthy work environment. There was also a risk that employees who might make sexually harassing remarks would cause problems without realizing that they were making such remarks.

[0539] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0540] In this invention, the server includes a data collection means, a voice or text data capture means, a voice recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, an alert generation means for generating a warning message when sexual harassment remarks are detected, a notification means for notifying a user of the warning message, and a real-time notification means for capturing data in real time using a smart device and immediately notifying a user of the warning message based on the results of the data analysis. This makes it possible to monitor conversations and chats in the office in real time, immediately detect sexual harassment remarks, issue a warning to employees, and encourage them to take prompt action.

[0541] "Data collection means" refers to a means for collecting information such as voice data and text data and saving it in a format required for subsequent processing.

[0542] "Audio or text data capture means" means a device or software for capturing audio or text data in real time.

[0543] "Speech recognition means" is a technology that converts captured voice data into text data.

[0544] The "analysis means" is a means for analyzing collected text data and evaluating and classifying it based on specific conditions.

[0545] A "generative AI model" is an artificial intelligence model that learns from past data and makes predictions and classifications for new data.

[0546] An "alert generator" is a device or software for generating a warning message when a particular condition is met.

[0547] The "notification means" is a means for transmitting the generated warning message to the user.

[0548] A "smart device" is a mobile terminal equipped with advanced computing power and communication functions.

[0549] "Real-time data capture" is the act of collecting data in synchronization with real time.

[0550] "Real-time notification means" refers to a means for transmitting a warning message to the user in a timely manner according to the analysis results.

[0551] This invention relates to a system for monitoring conversations and chats in a workplace and detecting and issuing warnings about sexual harassment comments in real time. Specific embodiments for implementing this system will be described below.

[0552] The system includes a data collection means, a voice or text data capture means, a voice recognition means, an analysis means, an alert generation means, a notification means, real-time data capture using a smart device, and a real-time notification means.

[0553] Server processing

[0554] The server processes data using the following hardware and software:

[0555] Hardware: Server

[0556] Software: Database (MySQL), Generative AI model (OpenAI GPT-3)

[0557] As a means of collecting data, past sexual harassment cases are stored in a database, and the generative AI model learns from this data. New sexual harassment cases are continually added to the database, and it is updated with the latest data.

[0558] What the device is doing

[0559] The terminal uses the following hardware and software:

[0560] Hardware: Smartphone, microphone

[0561] Software: Speech recognition API (Google Speech-to-Text)

[0562] The smartphone's microphone is used to capture voice or text data, and office conversations are captured in real time. The voice data is converted into text data using a speech recognition API, and the converted text data is sent to the server.

[0563] Server text data analysis

[0564] The server uses a generative AI model to analyze text data and score it for potential sexual harassment. This analysis is performed in real time based on a historical database, and a warning message is generated if the score exceeds a certain threshold.

[0565] Alerting and Notifications

[0566] As an alert generation means, the server generates a warning message and immediately notifies the user through a notification means, which is visually and audibly displayed on the smartphone in real time.

[0567] Specific examples

[0568] For example, if User A says, "You're so cute," the smartphone's microphone captures this remark. The speech data is converted into text data using a speech recognition API (Google Speech-to-Text). This text data is sent to a server and analyzed by a generative AI model (OpenAI GPT-3). If the remark is determined to be highly likely to be sexual harassment, a warning message is generated and a message stating, "Please be careful as this may be sexual harassment," is immediately displayed on User A's smartphone.

[0569] Prompt Sentence Examples

[0570] For example, the input prompt for a generative AI model might look like this:

[0571] "New employee XX is so beautiful."

[0572] Based on this prompt, the generative AI model evaluates the statement and issues a warning if necessary.

[0573] As described above, the present invention provides a system for preventing sexual harassment remarks in a workplace environment and maintaining healthy communication.

[0574] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0575] Step 1:

[0576] The device performs audio capture. To obtain audio data, the device (smartphone or PC) constantly monitors the ambient sound using a microphone. The input is the user's voice, and the output is audio data.

[0577] Step 2:

[0578] The device performs speech recognition and converts the speech data into text data. The acquired speech data is converted into text using a speech recognition API (Google Speech-to-Text). The input is speech data, and the output is text data.

[0579] Step 3:

[0580] The terminal sends the converted text data to the server. The text data is sent to the server via the HTTPS protocol. The input is text data, and the output is data sent to the server.

[0581] Step 4:

[0582] The server analyzes the text data using a generative AI model. Based on the received text data, the server uses a generative AI model (OpenAI GPT-3) to score the likelihood of sexual harassment remarks. At this time, a prompt sentence may be used as input. The input is text data, and the output is a score of the likelihood of sexual harassment remarks.

[0583] Step 5:

[0584] The server generates a warning message based on the analysis results. If the scoring result by the generative AI model exceeds a certain threshold, the server generates a customized warning message. The input is the likelihood score of sexual harassment remarks, and the output is a warning message.

[0585] Step 6:

[0586] The server sends a warning message to the terminal. The generated warning message is sent to the terminal via HTTPS protocol and notified in real time. The input is the warning message and the output is the data sent to the terminal.

[0587] Step 7:

[0588] The terminal displays a warning message to the user using a notification means. The terminal notifies the user of the received warning message visually or audibly. Specifically, the warning message is displayed on the screen and an audio notification is played. The input is the sent warning message, and the output is a notification message to the user.

[0589] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0590] Specific Embodiments of the Invention

[0591] This invention relates to a system that monitors conversations and chats, detects possible sexual harassment comments, and issues a warning in real time, and further includes a function to recognize the user's emotions and provide appropriate feedback. Specific embodiments for implementing this system are described below.

[0592] Program processing

[0593] Data collection

[0594] The server imports past sexual harassment case data and stores it in a database, which is used to train the generative AI model and improve the accuracy of the emotion engine.

[0595] Voice and text data capture

[0596] The device (smartphone or PC) captures voice or text data in real time from a microphone or chat app. In the case of voice data, this is then converted into text format.

[0597] Voice Recognition

[0598] The device uses voice recognition technology to convert the captured voice data into text data, which is then used for subsequent analysis by generative AI models and emotion engines.

[0599] Text data analysis

[0600] The device analyzes the text data and assesses the likelihood of sexual harassment using a generative AI model, which is trained on data from past sexual harassment cases.

[0601] emotion recognition

[0602] The device uses an emotion engine to recognize emotions from the user's facial expressions, tone of voice, etc. Emotion data is combined with the results of text data analysis to generate the final warning message.

[0603] Detecting sexual harassment remarks

[0604] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If the remark is judged to be sexually harassing with a high probability, it will activate an alert generation method.

[0605] Generate alerts

[0606] If the device determines that a remark is likely to be sexually harassing, it generates a warning message that is customized to the situation by incorporating emotional data provided by the emotion engine.

[0607] User Notification

[0608] The device notifies the user of the generated warning message. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification. The user is immediately aware of what they said and can make any necessary corrections.

[0609] Specific examples

[0610] Scenario 1: Remote conference operation

[0611] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[0612] 2. The device captures this speech and converts it into text using speech recognition technology.

[0613] 3. The device uses a generative AI model to analyze the comment and determine that it may be sexual harassment.

[0614] 4. The device uses the emotion engine to recognize User A's emotions, for example, determining whether User A is speaking in a joking or serious tone.

[0615] 5. The device generates a warning saying, "Be careful, this may be sexual harassment," and customizes the tone and content based on data from the emotion engine.

[0616] 6. The device displays a warning message on User A's screen.

[0617] 7. User A recognizes what he said and corrects himself by saying, "I'm sorry, I made an inappropriate comment."

[0618] Scenario 2: Chat operation

[0619] 1. User B sends a message in a chat room saying, "New employee XX is so beautiful."

[0620] 2. The device captures this text message.

[0621] 3. The device uses a generative AI model to analyze the message and determine that it may be sexually harassing.

[0622] 4. The device uses an emotion engine to recognize User B's emotions. For example, it determines whether User B is speaking in a friendly manner or teasing the user.

[0623] 5. The device generates a warning saying, "Be careful, this may be sexual harassment," and customizes the tone and content based on data from the emotion engine.

[0624] 6. The device displays an alert on User B's chat screen.

[0625] 7. User B immediately cancels the message and sends a correction message saying, "I apologize for my inappropriate comment."

[0626] This system enables the detection and feedback of sexually harassing remarks taking into account the user's emotions, which is expected to lead to more accurate and effective communication.

[0627] The processing flow will be explained below.

[0628] Step 1:

[0629] The server imports historical sexual harassment case data and stores it in a database, which is used to train generative AI models and emotion engines.

[0630] Step 2:

[0631] The device captures voice or text data from the microphone or chat app, allowing users to collect their conversations and messages in real time.

[0632] Step 3:

[0633] When the device receives voice data, it uses a speech recognition engine to convert the voice into text data, which prepares the data in a format that can be analyzed.

[0634] Step 4:

[0635] The device inputs the text data into a generative AI model that analyzes the possibility of sexual harassment remarks. The generative AI model is trained from data on past sexual harassment cases.

[0636] Step 5:

[0637] The device captures the user's facial expressions and tone of voice, and uses an emotion engine to recognize the user's emotions, allowing it to understand the user's mental state and emotions.

[0638] Step 6:

[0639] The device combines the analysis results from the generative AI model with the emotional data from the emotion engine to assess the likelihood of sexual harassment. By combining the two sets of data, a more accurate judgment can be made.

[0640] Step 7:

[0641] If the device determines that a comment is likely to be sexually harassing, it will generate a warning message, dynamically adjusting the content and tone of the message based on the emotional data.

[0642] Step 8:

[0643] The terminal will notify the user of the generated warning message visually and audibly. The user will be alerted by a pop-up display and / or an audio notification.

[0644] Step 9:

[0645] Users will receive a warning message and be made aware that their comments or actions may constitute sexual harassment. Users can correct or retract their comments as necessary.

[0646] Step 10:

[0647] The server collects newly detected sexual harassment cases and emotion data and adds them to the database, which continuously updates the generative AI model and emotion engine, improving the accuracy of the entire system.

[0648] Example 2

[0649] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0650] Sexually harassing remarks have become a problem in modern remote conference and chat environments. However, there are few systems that can detect them in real time and take immediate countermeasures. It has also been pointed out that there is a lack of systems that not only analyze the content of remarks but also recognize the speaker's emotions and provide appropriate feedback. The purpose of this invention is to solve these problems and provide a system that effectively detects and warns about sexually harassing remarks.

[0651] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0652] In this invention, the server includes a data collection means, a voice or text data capture means, a voice recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment expressions, an emotion recognition means for recognizing a user's emotions and integrating them with the analysis results, an alert generation means for generating a warning message when possible sexual harassment expressions are detected, and a notification means for notifying the user of the warning message. This enables effective detection and warning of sexual harassment remarks in real time and appropriate feedback that takes the user's emotions into consideration.

[0653] "Data Collection Measures" refers to the devices and processes used to import and store historical sexual harassment case data in a database.

[0654] "Voice or text data capture means" refers to devices and processes that obtain voice or text data through a microphone or chat app.

[0655] "Speech recognition means" refers to the technology and processes that convert captured voice data into text data.

[0656] A "generative AI model" refers to an artificial intelligence model that is trained based on data from past sexual harassment cases and analyzes text data to detect possible sexual harassment remarks.

[0657] "Analysis means" refers to the devices and processes that use generative AI models to analyze text data and assess potential sexual harassment comments.

[0658] "Emotion recognition means" refers to technology and processes for recognizing emotions from a user's facial expressions and tone of voice and generating emotion data.

[0659] "Alert generation means" refers to a device and process that generates a warning message when it determines that a sexually harassing remark is likely.

[0660] "Notification means" refers to devices and processes for notifying a user of generated warning messages.

[0661] MODE FOR CARRYING OUT THE INVENTION

[0662] The present invention is a system that monitors conversations and chats in real time, detects possible sexual harassment comments, and issues a warning. It also includes a function to recognize the user's emotions and provide appropriate feedback. A specific embodiment of this system will be described below.

[0663] Data collection methods

[0664] The server imports past sexual harassment case data and stores it in a database. This data contains a wide variety of sexual harassment cases and is used to train the generative AI model and emotion recognition engine. The server implements data collection and storage processes using Python and SQL, which significantly improves the accuracy and effectiveness of the system.

[0665] A means of capturing voice or text data

[0666] The device (smartphone or PC) captures voice or text data in real time through a microphone or chat app. Voice data is converted into text using speech recognition technology such as the Google Cloud Speech-to-Text API. This allows the spoken content to be obtained as text and used for subsequent processing.

[0667] Voice recognition means

[0668] The device uses voice recognition technology to convert the captured audio data into text data, using services such as Google Cloud Speech-to-Text, which is then immediately stored internally and passed on to the next analysis step.

[0669] Analytical tools using generative AI models

[0670] The device analyzes text data using a generative AI model to assess the likelihood of sexual harassment. This AI model is trained based on past data on sexual harassment cases and can evaluate the content of statements with high accuracy. For example, it uses advanced generative AI models such as BERT and GPT-3.

[0671] emotion recognition means

[0672] The device uses an emotion recognition engine to recognize emotions from the user's facial expressions and tone of voice, utilizing the Microsoft Azure Emotion API, among other tools. By acquiring emotion data, the device can more accurately understand the intention and nuance of what is being said.

[0673] Alert generation method

[0674] The device evaluates the likelihood of sexual harassment based on the analysis results of the generative AI model and data from the emotion recognition engine. If there is a high probability that the remark is sexual harassment, it automatically generates a warning message. This message is customized to an appropriate tone and content based on the emotion data.

[0675] Notification means

[0676] The terminal generates a warning message and notifies the user. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification. In particular, if the message is urgent, it can immediately draw the user's attention.

[0677] Specific operation example

[0678] Scenario 1: Remote conference operation

[0679] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[0680] 2. The device captures this speech and converts it into text using speech recognition technology (powered by the Google Cloud Speech-to-Text API).

[0681] 3. The device analyzes the comment using a generative AI model (BERT, GPT-3, etc.) and determines that it may be sexual harassment.

[0682] 4. The device uses an emotion engine (such as Azure Emotion API) to recognize User A's emotions.

[0683] 5. The device generates a warning saying, "Beware of potential sexual harassment," and customizes the tone and content accordingly.

[0684] 6. The device pops up a warning message on User A's screen.

[0685] 7. User A corrects himself by saying, "I'm sorry, I made an inappropriate comment."

[0686] Examples of prompt statements

[0687] During a remote meeting, User A says, "You're so cute. How about we go out for dinner sometime?" Capture this remark and analyze it using a sexual harassment remark detection system. Explain the process of generating an appropriate warning message based on the analysis results of the generative AI model and emotion recognition results, and notifying User A.

[0688] This system can detect sexual harassment remarks in real time and provide appropriate warning messages that take emotions into account. This configuration can improve effective communication.

[0689] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0690] Step 1:

[0691] The server imports data on past sexual harassment cases and stores it in a database. The input is data on past sexual harassment cases, and the output is the data stored in the database. A data collection program using Python and SQL reads and saves the initial data.

[0692] What happens: The server periodically checks the data source (e.g., CSV file or external API) and imports any new data it finds. The imported data is then stored in the database.

[0693] Step 2:

[0694] The device captures voice or text data in real time through a microphone or chat app. The input is the user's speech data, and the output is the captured voice file or text data. The data is acquired in real time using JavaScript or Python libraries.

[0695] What it does: When a conversation starts, an app installed on the device activates the microphone to capture voice data in real time, and also captures text messages directly from chat apps.

[0696] Step 3:

[0697] The device converts the voice data into text data using the Google Cloud Speech-to-Text API or similar. The input is the captured voice data, and the output is the converted text data. The device calls a speech recognition API to convert the voice to text.

[0698] Specific operation: The device sends voice data to the API, receives the voice recognition results, stores the results in internal memory, and passes them on to the next process.

[0699] Step 4:

[0700] The device analyzes text data using a generative AI model to assess the likelihood of sexually harassing comments. The input is text data, and the output is a probability score of sexually harassing comments. AI models such as BERT and GPT-3 are used.

[0701] How it works: The device inputs text data into the generative AI model, which then outputs a score indicating the likelihood of the comment being sexually harassing. This score is then used in the next step.

[0702] Step 5:

[0703] The device uses an emotion recognition engine to recognize the user's emotions. The input is the user's facial expression and tone of voice, and the output is the recognized emotion data. It uses the Microsoft Azure Emotion API, etc.

[0704] Specific operation: The device uses the camera and microphone connected to the device to capture the user's facial expressions and tone of voice, and sends them to the emotion recognition API. The obtained emotion data is then stored in the device's internal memory.

[0705] Step 6:

[0706] The device evaluates the probability of sexual harassment remarks based on the analysis results of the generated AI model and emotion recognition data. The input is the probability score of sexual harassment remarks and emotion data, and the output is an alert generation flag. A threshold is set, and if it is exceeded, an alert generation flag is set.

[0707] Specific operation: The analysis results of the text data (probability score) are integrated with the emotion data, and if a certain threshold is exceeded, an alert generation flag is set and the next process is carried out.

[0708] Step 7:

[0709] Generates a warning message when the device has an alert generation flag. The input is the alert generation flag and emotion data, and the output is the warning message. Generates the warning message using a Python script.

[0710] Specific operation: Prepare a warning message template and generate a customized message based on the emotion data. This message is notified to the user in the next step.

[0711] Step 8:

[0712] The terminal notifies the user of the generated warning message. The input is the warning message and the output is the notification to the user. The notification methods are visual (screen popup) and audible (voice notification).

[0713] Specific behavior: The device will display the generated warning message as a pop-up on the screen and, if necessary, provide an audio notification. This is implemented using JavaScript or the Android and iOS notification APIs.

[0714] (Application example 2)

[0715] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0716] In recent years, sexual harassment has become a problem in the workplace and online communities. However, systems that can detect sexual harassment in real time and issue appropriate warnings are not yet widely available. Furthermore, there is a need for technology that provides more accurate and effective feedback by taking user emotions into account. The present invention aims to provide a sexual harassment detection system that solves these problems.

[0717] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0718] In this invention, the server includes a data collection means, a means for capturing voice or text data, a means for converting the captured voice data into text data, a means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, a means for generating a warning message when a sexual harassment remark is detected, a means for notifying the user of the warning message, and an emotion recognition means for recognizing the user's emotion. This makes it possible to detect sexual harassment remarks in real time and issue an appropriate warning. Furthermore, by recognizing the user's emotion, it is possible to customize the warning message according to the situation and provide more effective feedback.

[0719] "Data collection means" is a function that imports data on past sexual harassment cases and stores it in a database.

[0720] A "voice or text data capture means" is a device that acquires voice or text data in real time from a microphone or chat app.

[0721] A "voice recognition means" is a technique or device that converts captured voice data into text data.

[0722] "Analysis tools" are technologies or devices that use generative AI models to analyze text data and assess the likelihood of sexual harassment statements.

[0723] A "generative AI model" is an artificial intelligence model trained based on data on past sexual harassment cases.

[0724] The "alert generation means" is a function that generates a warning message when a sexual harassment remark is detected.

[0725] The "notification means" is a function that notifies the user of the generated warning message visually or audibly.

[0726] The "emotion recognition means" is a function that recognizes the user's emotions from facial expressions, tone of voice, etc.

[0727] A specific embodiment for implementing this invention will be described below. This system monitors comments in real time when users communicate via voice or chat, and issues an immediate warning if possible sexual harassment comments are detected. Furthermore, it recognizes the user's emotions and appropriately customizes the warning message, thereby achieving effective feedback.

[0728] Data collection

[0729] The server imports past sexual harassment case data and stores it in a database. This data is used to train the generative AI model and also helps improve the accuracy of the emotion engine. For example, a dataset could include sexual harassment cases collected from various industries.

[0730] Voice and text data capture

[0731] The device captures voice or text data in real time from a microphone or chat app. In the case of voice data, it needs to be converted into text. For example, a smartphone microphone can be used to capture voice and a chat app can accept input as a text message.

[0732] Voice Recognition

[0733] The device uses voice recognition technology to convert the captured voice data into text data, using Google's voice recognition service as an example. This converted text data is then used for subsequent analysis by generative AI models and sentiment engines.

[0734] Text data analysis

[0735] The device uses a generative AI model to analyze the captured text data and evaluate the possibility of sexual harassment. The generative AI model is trained based on data from past sexual harassment cases and evaluates the content of comments with high accuracy. For example, a comment such as "You're so cute. How about going out for dinner sometime?" is judged to be potentially sexual harassment.

[0736] emotion recognition

[0737] The device uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice. The emotion data is combined with the results of text data analysis to generate the final warning message. For example, it can determine whether the user is speaking in a joking or serious tone.

[0738] Detecting sexual harassment remarks

[0739] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If the remark is judged to be sexually harassing with a high probability, it will activate an alert generation method.

[0740] Alert generation and notification

[0741] If the device determines that a remark is likely to be sexual harassment, it generates a warning message. The generated warning message incorporates emotional data provided by the emotion engine and is customized according to the situation. The device then notifies the user of the warning message visually or audibly using a notification means. For example, a warning such as "Be careful as this may be sexual harassment" may be displayed on the user's screen.

[0742] Specific examples

[0743] 1. Scenario 1: Remote Meeting Operation

[0744] If a user says, "You're so cute. Let's have dinner together sometime," during a remote meeting, the remark is captured and converted into text. A generative AI model then detects potential sexual harassment, and an emotion recognition engine recognizes the user's playful tone. As a result, a warning message appears on the user's screen saying, "Be careful, this may be sexual harassment."

[0745] 2. Scenario 2: Chat operation

[0746] If a user sends a message in a chat room saying, "New employee XX is so beautiful," this text message is captured and the generative AI model detects potential sexual harassment. The emotion recognition engine determines whether the user is speaking in a friendly or teasing manner, and a warning message is displayed on the user's chat screen saying, "Be careful, this may be sexual harassment."

[0747] Prompt Sentence Examples

[0748] Example: "You're so cute, how about we have dinner sometime?"

[0749] Example: "The new employee, Ms. XX, is very beautiful."

[0750] In this way, this system enables the detection of sexually harassing remarks and feedback in real time, taking into account the user's emotions, which is expected to lead to more accurate and effective communication.

[0751] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0752] Step 1: Data collection

[0753] The server imports past sexual harassment case data and stores it in a database. This is primarily used to improve the accuracy of the generative AI model and emotion recognition engine. Specifically, it analyzes past sexual harassment remark data collected from within the company or from public datasets and registers it in the database. The input is past sexual harassment case data, and the output is the data stored in the database.

[0754] Step 2: Capture audio and text data

[0755] The device captures voice and text data in real time from a microphone or chat app. In the case of voice data, it must be converted into text format. Specifically, the device picks up what the user says with a microphone and captures the voice data, or receives text data directly from a chat app. The input is the user's real-time voice or chat message, and the output is the captured voice data or text data.

[0756] Step 3: Voice Recognition

[0757] The device uses speech recognition technology to convert the captured voice data into text data. Here, Google's speech recognition API is used for speech-to-text conversion. The input is the captured voice data, and the output is text data.

[0758] Step 4: Analyze the text data

[0759] The device analyzes text data using a generative AI model to assess the likelihood of sexual harassment. The model is a trained generative AI model that evaluates the content of statements with high accuracy based on past data on sexual harassment cases. The input is text data, and the output is an assessment result indicating the likelihood of sexual harassment.

[0760] Step 5: Emotion Recognition

[0761] The device uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice. For example, it determines whether the user is speaking in a joking or serious tone. The input is the user's facial expression data and tone of voice data, and the output is the recognized emotion data.

[0762] Step 6: Detect sexual harassment comments

[0763] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If it determines with a high probability that the remark was sexually harassing, it activates an alert generation mechanism. The inputs are the analysis results and emotional data, and the output is the evaluation result of the remark being sexually harassing.

[0764] Step 7: Generate an alert

[0765] The device generates a warning message if it determines that a remark is likely to be sexual harassment. The generated warning message is customized by incorporating emotional data provided by an emotion recognition engine. The input is the evaluation result of the sexual harassment remark and the emotional data, and the output is a customized warning message.

[0766] Step 8: Notify users

[0767] The terminal notifies the user of the generated warning message visually or audibly. For example, it notifies the user of the warning by a screen pop-up or a voice notification. The input is the customized warning message, and the output is the warning message notified to the user.

[0768] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0769] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0770] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0771] [Third embodiment]

[0772] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0773] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0774] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0775] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0776] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0777] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0778] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0779] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0780] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0781] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0782] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0783] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0784] Specific Embodiments of the Invention

[0785] The present invention relates to a system that monitors conversations and chats within a company, detects possible sexual harassment comments, and issues a warning in real time. Specific embodiments for implementing this system will be described below.

[0786] Program processing

[0787] Data collection

[0788] The server imports data on past sexual harassment cases and stores it in a centrally managed database, allowing the system to learn in a way that responds to the diversity of situations and word usage.

[0789] Voice and text data capture

[0790] The device (smartphone or PC) captures voice or text data from a microphone or chat app. In the case of voice data, the next step is to convert it into text data.

[0791] Voice Recognition

[0792] The device uses voice recognition technology to convert the captured voice data into text, putting the data in a format that can be analyzed by the generative AI model.

[0793] Text data analysis

[0794] The device analyzes the text data using a generative AI model that has been trained on data from past sexual harassment cases and has the ability to score whether a comment is likely to be sexual harassment.

[0795] Detecting sexual harassment remarks

[0796] The device determines whether the remark is likely to be sexually harassing based on the scores and results obtained from the generative AI model. If the score exceeds a certain threshold during this process, the device proceeds to the next step.

[0797] Generate alerts

[0798] If the device determines that a remark is likely to be sexually harassing, it will generate a warning message, which will be customized depending on the situation and will alert the person making the remark.

[0799] User Notification

[0800] The terminal notifies the user of the generated warning message both visually (e.g., on the screen) and audibly, allowing the user to immediately become aware of their own remarks and make any necessary corrections.

[0801] Specific examples

[0802] Scenario 1: Remote conference operation

[0803] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[0804] 2. The device captures this speech and converts it into text using speech recognition technology.

[0805] 3. The device uses a generative AI model to analyze the comment as potentially sexual harassment.

[0806] 4. The device generates a warning saying, "Caution: potential sexual harassment."

[0807] 5. The device displays a warning message on User A's screen.

[0808] 6. User A realizes what he or she said and corrects himself or herself by saying, "I'm sorry, I made an inappropriate comment."

[0809] Scenario 2: Chat operation

[0810] 1. User B sends a message in a chat room saying, "New employee XX is so beautiful."

[0811] 2. The device captures this text message.

[0812] 3. The device uses a generative AI model to analyze the message and assess its potential for sexual harassment.

[0813] 4. The device generates an alert saying, "Be careful, there is a possibility of sexual harassment."

[0814] 5. The device displays an alert on User B's chat screen.

[0815] 6. User B immediately cancels the message and sends a correction message saying, "I apologize for my inappropriate comment."

[0816] In this way, by introducing this system, it is possible to prevent sexually harassing remarks in everyday communication, thereby maintaining and improving a healthy workplace environment.

[0817] The processing flow will be explained below.

[0818] Step 1:

[0819] The server imports past sexual harassment case data and stores it in a database, which is used to train generative AI models.

[0820] Step 2:

[0821] The device captures voice or text data from a microphone or chat app. In the case of voice data, it needs to be converted into text format for further processing.

[0822] Step 3:

[0823] When the device captures voice data, it uses a speech recognition engine to convert the speech into text data, which can then be analyzed by a generative AI model.

[0824] Step 4:

[0825] The device then inputs the converted text data into a generative AI model that analyzes potential sexual harassment comments. This AI model was trained based on the data collected in step 1.

[0826] Step 5:

[0827] The device evaluates the likelihood of a remark being sexually harassing based on the output from the generated AI model, and if it is determined to be sexually harassing with a high probability, it activates an alert generation method.

[0828] Step 6:

[0829] The device generates a warning message. Specifically, it creates a message indicating the possibility of sexual harassment according to a set format.

[0830] Step 7:

[0831] The device notifies the user of the generated warning message. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification.

[0832] Step 8:

[0833] The user receives a warning message and is made aware that their comments or messages may constitute sexual harassment. The user can correct their comments or take appropriate action.

[0834] Step 9:

[0835] The server continuously updates the generative AI model by collecting newly detected sexual harassment case data and adding it to the database, thereby improving the model's accuracy and effectiveness over time.

[0836] Example 1

[0837] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0838] Preventing sexual harassment in the modern workplace is a difficult task. Traditional methods require time for sexual harassment comments to be reported externally, so a prompt response is required to maintain a healthy workplace. Real-time monitoring and warnings are necessary, especially with the increase in remote work and online chat, but current systems tend to be slow to respond.

[0839] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0840] In this invention, the server includes a means for collecting data, a means for capturing voice or text data, a speech recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, a means for inputting a prompt sentence into the generative AI model, an alert generation means for generating a warning message when sexual harassment remarks are detected, and a notification means for notifying the user of the warning message. This enables real-time collection and analysis of data, and rapid generation and notification of warning messages.

[0841] "Means for collecting data" refers to systems or devices for importing and centrally managing data on past sexual harassment cases and related information.

[0842] "Audio or text data capture means" means a system or device for capturing and storing conversations or text messages in real time.

[0843] "Speech recognition means" refers to the technology or software used to convert captured voice data into text.

[0844] "Analytical means using generative AI models" refers to systems or devices that use AI models trained based on past cases of sexual harassment to analyze text data and evaluate the likelihood of sexual harassment remarks.

[0845] A "means for inputting prompts to a generative AI model" is an input interface or software for providing specific instructions or questions to an AI model.

[0846] An "alert generation means" is a system or device that generates a warning message when a sexually harassing remark is detected.

[0847] The "notification means" is a system or device for visually and audibly notifying the user of the generated warning message.

[0848] MODE FOR CARRYING OUT THE INVENTION

[0849] The present invention relates to a system that monitors conversations and chats within a company, detects possible sexual harassment comments, and issues a warning in real time. Specific embodiments for implementing this system will be described below.

[0850] The server is primarily used to collect data on past sexual harassment cases and store it in a centralized database. This system can use MySQL or PostgreSQL as the database. The collected data is used as training data for the generative AI model, and data preprocessing is performed to improve the model's accuracy. Preprocessing includes filling in missing values ​​and normalizing the data.

[0851] The device (smartphone or personal computer) is used to capture voice or text data in real time from remote meetings or chat apps. Voice can be recorded through a microphone and saved in WAV or MP3 format. Text data can also be obtained using the chat app's API.

[0852] For speech recognition, we use voice recognition technologies such as Google Cloud Speech-to-Text and Amazon Transcribe. This converts the audio data into text, which is then formatted for easy analysis by the generative AI model. The converted text data is then stripped of unnecessary spaces and special characters to prepare it for analysis.

[0853] Text data is analyzed in real time using an analytical method that uses a generative AI model (e.g., OpenAI GPT-4). During this process, a prompt is input into the generative AI model, which evaluates whether the remark constitutes sexual harassment. Including specific examples in the prompt allows for a more accurate judgment.

[0854] Some common prompts are:

[0855] Example prompt 1:

[0856] "Please determine whether the following statements constitute sexual harassment:

[0857] "You're so cute, how about we have dinner sometime?"

[0858] Example prompt 2:

[0859] "Please determine whether the following statements constitute sexual harassment:

[0860] "New employee XX is really beautiful."

[0861] If the analysis results of the generative AI model exceed a threshold, a warning message is generated as an alert generation means. This message includes specific instructions and warns the speaker that the comment is inappropriate.

[0862] The warning message is notified to the user through a notification means, which is visual (for example, a screen display) and audio, so that the user can immediately pay attention to the message.

[0863] As a result, the system can detect sexually harassing remarks in real time and take prompt action to maintain a healthy work environment.

[0864] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0865] Program processing flow

[0866] Step 1:

[0867] The server collects data on past sexual harassment cases and stores it in a centrally managed database.

[0868] Input: Historical sexual harassment case data from a company database.

[0869] Data processing: Preprocessing of collected data (filling in missing values ​​and normalizing data).

[0870] Output: Data stored as a training dataset for a generative AI model.

[0871] How it works: The server uses SQL queries to extract sexual harassment cases from the company's internal database over the past five years, imports the data into a MySQL database, and performs preprocessing such as missing value imputation and normalization to prepare a training dataset.

[0872] Step 2:

[0873] The device captures voice or text data in real time from conversations and chat apps.

[0874] Input: Audio data from a remote meeting or text messages from a chat app.

[0875] Data processing: Audio data is saved in WAV or MP3 format. Text data is saved in a temporary folder.

[0876] Output: The captured audio or text data.

[0877] Specific operation: During remote meetings, the device records audio data in real time in WAV format through the microphone, and uses the chat app's API to capture all messages sent in text format and save them in a temporary folder.

[0878] Step 3:

[0879] The terminal uses voice recognition technology to convert the captured voice data into text.

[0880] Input: Captured audio data (WAV or MP3 format).

[0881] Data processing: Convert speech to text using Google Cloud Speech-to-Text or Amazon Transcribe. Remove unnecessary spaces and special characters.

[0882] Output: The converted text data.

[0883] Specific operation: The device calls the Google Cloud Speech-to-Text API to convert the WAV audio file into a string, then removes unnecessary spaces and special characters from the converted text data in preparation for analysis.

[0884] Step 4:

[0885] The device uses a generative AI model to analyze the text data.

[0886] Input: The converted text data.

[0887] Data processing: A prompt was generated and input into a generative AI model (OpenAI GPT-4). This prompt included instructions to evaluate whether the remark constituted sexual harassment.

[0888] Output: Scoring and analysis results from the generative AI model.

[0889] Specific operation: The device generates the following prompt sentence and inputs it into the AI ​​model:

[0890] "Please determine whether the following statements constitute sexual harassment:

[0891] "You're so cute, how about we have dinner sometime?"

[0892] Obtain the scoring and analysis results returned by the AI ​​model and format the data.

[0893] Step 5:

[0894] The device determines whether a comment is likely to be sexual harassment based on the analysis results obtained from the generative AI model.

[0895] Input: Scoring and analysis results from the AI ​​model.

[0896] Data processing: Determine whether the score exceeds a predefined threshold.

[0897] Output: Triggers an alert generator when a threshold is exceeded.

[0898] Specific operation: The device checks whether the score returned by the AI ​​model is 80 or above (pre-set threshold), and if so, proceeds to the next alert generation step.

[0899] Step 6:

[0900] The terminal generates a warning message if a sexually harassing remark is detected.

[0901] Input: Speech data with scores above a threshold.

[0902] Data processing: Generation of customized warning messages.

[0903] Output: The warning message generated.

[0904] Specific behavior: If the device determines that a comment is sexually harassing, it will generate a warning message stating, "This comment may be inappropriate. Please correct it."

[0905] Step 7:

[0906] The terminal notifies the user of the generated warning message.

[0907] Input: The generated warning message.

[0908] Data processing: Preparation of visual and audio notifications.

[0909] Output: Notification to the user.

[0910] Specific operation: The device will display a warning message on the user's screen and simultaneously issue an audio alert to draw attention to the statement, allowing the user to immediately pay attention to the statement and make corrections.

[0911] (Application example 1)

[0912] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0913] In conventional workplaces, there was no system that could detect sexual harassment remarks in real time and issue a warning. As a result, when a victim received unpleasant remarks, it was difficult to take immediate action, making it difficult to maintain a healthy work environment. There was also a risk that employees who might make sexually harassing remarks would cause problems without realizing that they were making such remarks.

[0914] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0915] In this invention, the server includes a data collection means, a voice or text data capture means, a voice recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, an alert generation means for generating a warning message when sexual harassment remarks are detected, a notification means for notifying a user of the warning message, and a real-time notification means for capturing data in real time using a smart device and immediately notifying a user of the warning message based on the results of the data analysis. This makes it possible to monitor conversations and chats in the office in real time, immediately detect sexual harassment remarks, issue a warning to employees, and encourage them to take prompt action.

[0916] "Data collection means" refers to a means for collecting information such as voice data and text data and saving it in a format required for subsequent processing.

[0917] "Audio or text data capture means" means a device or software for capturing audio or text data in real time.

[0918] "Speech recognition means" is a technology that converts captured voice data into text data.

[0919] The "analysis means" is a means for analyzing collected text data and evaluating and classifying it based on specific conditions.

[0920] A "generative AI model" is an artificial intelligence model that learns from past data and makes predictions and classifications for new data.

[0921] An "alert generator" is a device or software for generating a warning message when a particular condition is met.

[0922] The "notification means" is a means for transmitting the generated warning message to the user.

[0923] A "smart device" is a mobile terminal equipped with advanced computing power and communication functions.

[0924] "Real-time data capture" is the act of collecting data in synchronization with real time.

[0925] "Real-time notification means" refers to a means for transmitting a warning message to the user in a timely manner according to the analysis results.

[0926] This invention relates to a system for monitoring conversations and chats in a workplace and detecting and issuing warnings about sexual harassment comments in real time. Specific embodiments for implementing this system will be described below.

[0927] The system includes a data collection means, a voice or text data capture means, a voice recognition means, an analysis means, an alert generation means, a notification means, real-time data capture using a smart device, and a real-time notification means.

[0928] Server processing

[0929] The server processes data using the following hardware and software:

[0930] Hardware: Server

[0931] Software: Database (MySQL), Generative AI model (OpenAI GPT-3)

[0932] As a means of collecting data, past sexual harassment cases are stored in a database, and the generative AI model learns from this data. New sexual harassment cases are continually added to the database, and it is updated with the latest data.

[0933] What the device is doing

[0934] The terminal uses the following hardware and software:

[0935] Hardware: Smartphone, microphone

[0936] Software: Speech recognition API (Google Speech-to-Text)

[0937] The smartphone's microphone is used to capture voice or text data, and office conversations are captured in real time. The voice data is converted into text data using a speech recognition API, and the converted text data is sent to the server.

[0938] Server text data analysis

[0939] The server uses a generative AI model to analyze text data and score it for potential sexual harassment. This analysis is performed in real time based on a historical database, and a warning message is generated if the score exceeds a certain threshold.

[0940] Alerting and Notifications

[0941] As an alert generation means, the server generates a warning message and immediately notifies the user through a notification means, which is visually and audibly displayed on the smartphone in real time.

[0942] Specific examples

[0943] For example, if User A says, "You're so cute," the smartphone's microphone captures this remark. The speech data is converted into text data using a speech recognition API (Google Speech-to-Text). This text data is sent to a server and analyzed by a generative AI model (OpenAI GPT-3). If the remark is determined to be highly likely to be sexual harassment, a warning message is generated and a message stating, "Please be careful as this may be sexual harassment," is immediately displayed on User A's smartphone.

[0944] Prompt Sentence Examples

[0945] For example, the input prompt for a generative AI model might look like this:

[0946] "New employee XX is so beautiful."

[0947] Based on this prompt, the generative AI model evaluates the statement and issues a warning if necessary.

[0948] As described above, the present invention provides a system for preventing sexual harassment remarks in a workplace environment and maintaining healthy communication.

[0949] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0950] Step 1:

[0951] The device performs audio capture. To obtain audio data, the device (smartphone or PC) constantly monitors the ambient sound using a microphone. The input is the user's voice, and the output is audio data.

[0952] Step 2:

[0953] The device performs speech recognition and converts the speech data into text data. The acquired speech data is converted into text using a speech recognition API (Google Speech-to-Text). The input is speech data, and the output is text data.

[0954] Step 3:

[0955] The terminal sends the converted text data to the server. The text data is sent to the server via the HTTPS protocol. The input is text data, and the output is data sent to the server.

[0956] Step 4:

[0957] The server analyzes the text data using a generative AI model. Based on the received text data, the server uses a generative AI model (OpenAI GPT-3) to score the likelihood of sexual harassment remarks. At this time, a prompt sentence may be used as input. The input is text data, and the output is a score of the likelihood of sexual harassment remarks.

[0958] Step 5:

[0959] The server generates a warning message based on the analysis results. If the scoring result by the generative AI model exceeds a certain threshold, the server generates a customized warning message. The input is the likelihood score of sexual harassment remarks, and the output is a warning message.

[0960] Step 6:

[0961] The server sends a warning message to the terminal. The generated warning message is sent to the terminal via HTTPS protocol and notified in real time. The input is the warning message and the output is the data sent to the terminal.

[0962] Step 7:

[0963] The terminal displays a warning message to the user using a notification means. The terminal notifies the user of the received warning message visually or audibly. Specifically, the warning message is displayed on the screen and an audio notification is played. The input is the sent warning message, and the output is a notification message to the user.

[0964] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0965] Specific Embodiments of the Invention

[0966] This invention relates to a system that monitors conversations and chats, detects possible sexual harassment comments, and issues a warning in real time, and further includes a function to recognize the user's emotions and provide appropriate feedback. Specific embodiments for implementing this system are described below.

[0967] Program processing

[0968] Data collection

[0969] The server imports past sexual harassment case data and stores it in a database, which is used to train the generative AI model and improve the accuracy of the emotion engine.

[0970] Voice and text data capture

[0971] The device (smartphone or PC) captures voice or text data in real time from a microphone or chat app. In the case of voice data, this is then converted into text format.

[0972] Voice Recognition

[0973] The device uses voice recognition technology to convert the captured voice data into text data, which is then used for subsequent analysis by generative AI models and emotion engines.

[0974] Text data analysis

[0975] The device analyzes the text data and assesses the likelihood of sexual harassment using a generative AI model, which is trained on data from past sexual harassment cases.

[0976] emotion recognition

[0977] The device uses an emotion engine to recognize emotions from the user's facial expressions, tone of voice, etc. Emotion data is combined with the results of text data analysis to generate the final warning message.

[0978] Detecting sexual harassment remarks

[0979] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If the remark is judged to be sexually harassing with a high probability, it will activate an alert generation method.

[0980] Generate alerts

[0981] If the device determines that a remark is likely to be sexually harassing, it generates a warning message that is customized to the situation by incorporating emotional data provided by the emotion engine.

[0982] User Notification

[0983] The device notifies the user of the generated warning message. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification. The user is immediately aware of what they said and can make any necessary corrections.

[0984] Specific examples

[0985] Scenario 1: Remote conference operation

[0986] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[0987] 2. The device captures this speech and converts it into text using speech recognition technology.

[0988] 3. The device uses a generative AI model to analyze the comment and determine that it may be sexual harassment.

[0989] 4. The device uses the emotion engine to recognize User A's emotions, for example, determining whether User A is speaking in a joking or serious tone.

[0990] 5. The device generates a warning saying, "Be careful, this may be sexual harassment," and customizes the tone and content based on data from the emotion engine.

[0991] 6. The device displays a warning message on User A's screen.

[0992] 7. User A recognizes what he said and corrects himself by saying, "I'm sorry, I made an inappropriate comment."

[0993] Scenario 2: Chat operation

[0994] 1. User B sends a message in a chat room saying, "New employee XX is so beautiful."

[0995] 2. The device captures this text message.

[0996] 3. The device uses a generative AI model to analyze the message and determine that it may be sexually harassing.

[0997] 4. The device uses an emotion engine to recognize User B's emotions. For example, it determines whether User B is speaking in a friendly manner or teasing the user.

[0998] 5. The device generates a warning saying, "Be careful, this may be sexual harassment," and customizes the tone and content based on data from the emotion engine.

[0999] 6. The device displays an alert on User B's chat screen.

[1000] 7. User B immediately cancels the message and sends a correction message saying, "I apologize for my inappropriate comment."

[1001] This system enables the detection and feedback of sexually harassing remarks taking into account the user's emotions, which is expected to lead to more accurate and effective communication.

[1002] The processing flow will be explained below.

[1003] Step 1:

[1004] The server imports historical sexual harassment case data and stores it in a database, which is used to train generative AI models and emotion engines.

[1005] Step 2:

[1006] The device captures voice or text data from the microphone or chat app, allowing users to collect their conversations and messages in real time.

[1007] Step 3:

[1008] When the device receives voice data, it uses a speech recognition engine to convert the voice into text data, which prepares the data in a format that can be analyzed.

[1009] Step 4:

[1010] The device inputs the text data into a generative AI model that analyzes the possibility of sexual harassment remarks. The generative AI model is trained from data on past sexual harassment cases.

[1011] Step 5:

[1012] The device captures the user's facial expressions and tone of voice, and uses an emotion engine to recognize the user's emotions, allowing it to understand the user's mental state and emotions.

[1013] Step 6:

[1014] The device combines the analysis results from the generative AI model with the emotional data from the emotion engine to assess the likelihood of sexual harassment. By combining the two sets of data, a more accurate judgment can be made.

[1015] Step 7:

[1016] If the device determines that a comment is likely to be sexually harassing, it will generate a warning message, dynamically adjusting the content and tone of the message based on the emotional data.

[1017] Step 8:

[1018] The terminal will notify the user of the generated warning message visually and audibly. The user will be alerted by a pop-up display and / or an audio notification.

[1019] Step 9:

[1020] Users will receive a warning message and be made aware that their comments or actions may constitute sexual harassment. Users can correct or retract their comments as necessary.

[1021] Step 10:

[1022] The server collects newly detected sexual harassment cases and emotion data and adds them to the database, which continuously updates the generative AI model and emotion engine, improving the accuracy of the entire system.

[1023] Example 2

[1024] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1025] Sexually harassing remarks have become a problem in modern remote conference and chat environments. However, there are few systems that can detect them in real time and take immediate countermeasures. It has also been pointed out that there is a lack of systems that not only analyze the content of remarks but also recognize the speaker's emotions and provide appropriate feedback. The purpose of this invention is to solve these problems and provide a system that effectively detects and warns about sexually harassing remarks.

[1026] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1027] In this invention, the server includes a data collection means, a voice or text data capture means, a voice recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment expressions, an emotion recognition means for recognizing a user's emotions and integrating them with the analysis results, an alert generation means for generating a warning message when possible sexual harassment expressions are detected, and a notification means for notifying the user of the warning message. This enables effective detection and warning of sexual harassment remarks in real time and appropriate feedback that takes the user's emotions into consideration.

[1028] "Data Collection Measures" refers to the devices and processes used to import and store historical sexual harassment case data in a database.

[1029] "Voice or text data capture means" refers to devices and processes that obtain voice or text data through a microphone or chat app.

[1030] "Speech recognition means" refers to the technology and processes that convert captured voice data into text data.

[1031] A "generative AI model" refers to an artificial intelligence model that is trained based on data from past sexual harassment cases and analyzes text data to detect possible sexual harassment remarks.

[1032] "Analysis means" refers to the devices and processes that use generative AI models to analyze text data and assess potential sexual harassment comments.

[1033] "Emotion recognition means" refers to technology and processes for recognizing emotions from a user's facial expressions and tone of voice and generating emotion data.

[1034] "Alert generation means" refers to a device and process that generates a warning message when it determines that a sexually harassing remark is likely.

[1035] "Notification means" refers to devices and processes for notifying a user of generated warning messages.

[1036] MODE FOR CARRYING OUT THE INVENTION

[1037] The present invention is a system that monitors conversations and chats in real time, detects possible sexual harassment comments, and issues a warning. It also includes a function to recognize the user's emotions and provide appropriate feedback. A specific embodiment of this system will be described below.

[1038] Data collection methods

[1039] The server imports past sexual harassment case data and stores it in a database. This data contains a wide variety of sexual harassment cases and is used to train the generative AI model and emotion recognition engine. The server implements data collection and storage processes using Python and SQL, which significantly improves the accuracy and effectiveness of the system.

[1040] A means of capturing voice or text data

[1041] The device (smartphone or PC) captures voice or text data in real time through a microphone or chat app. Voice data is converted into text using speech recognition technology such as the Google Cloud Speech-to-Text API. This allows the spoken content to be obtained as text and used for subsequent processing.

[1042] Voice recognition means

[1043] The device uses voice recognition technology to convert the captured audio data into text data, using services such as Google Cloud Speech-to-Text, which is then immediately stored internally and passed on to the next analysis step.

[1044] Analytical tools using generative AI models

[1045] The device analyzes text data using a generative AI model to assess the likelihood of sexual harassment. This AI model is trained based on past data on sexual harassment cases and can evaluate the content of statements with high accuracy. For example, it uses advanced generative AI models such as BERT and GPT-3.

[1046] emotion recognition means

[1047] The device uses an emotion recognition engine to recognize emotions from the user's facial expressions and tone of voice, utilizing the Microsoft Azure Emotion API, among other tools. By acquiring emotion data, the device can more accurately understand the intention and nuance of what is being said.

[1048] Alert generation method

[1049] The device evaluates the likelihood of sexual harassment based on the analysis results of the generative AI model and data from the emotion recognition engine. If there is a high probability that the remark is sexual harassment, it automatically generates a warning message. This message is customized to an appropriate tone and content based on the emotion data.

[1050] Notification means

[1051] The terminal generates a warning message and notifies the user. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification. In particular, if the message is urgent, it can immediately draw the user's attention.

[1052] Specific operation example

[1053] Scenario 1: Remote conference operation

[1054] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[1055] 2. The device captures this speech and converts it into text using speech recognition technology (powered by the Google Cloud Speech-to-Text API).

[1056] 3. The device analyzes the comment using a generative AI model (BERT, GPT-3, etc.) and determines that it may be sexual harassment.

[1057] 4. The device uses an emotion engine (such as Azure Emotion API) to recognize User A's emotions.

[1058] 5. The device generates a warning saying, "Beware of potential sexual harassment," and customizes the tone and content accordingly.

[1059] 6. The device pops up a warning message on User A's screen.

[1060] 7. User A corrects himself by saying, "I'm sorry, I made an inappropriate comment."

[1061] Examples of prompt statements

[1062] During a remote meeting, User A says, "You're so cute. How about we go out for dinner sometime?" Capture this remark and analyze it using a sexual harassment remark detection system. Explain the process of generating an appropriate warning message based on the analysis results of the generative AI model and emotion recognition results, and notifying User A.

[1063] This system can detect sexual harassment remarks in real time and provide appropriate warning messages that take emotions into account. This configuration can improve effective communication.

[1064] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1065] Step 1:

[1066] The server imports data on past sexual harassment cases and stores it in a database. The input is data on past sexual harassment cases, and the output is the data stored in the database. A data collection program using Python and SQL reads and saves the initial data.

[1067] What happens: The server periodically checks the data source (e.g., CSV file or external API) and imports any new data it finds. The imported data is then stored in the database.

[1068] Step 2:

[1069] The device captures voice or text data in real time through a microphone or chat app. The input is the user's speech data, and the output is the captured voice file or text data. The data is acquired in real time using JavaScript or Python libraries.

[1070] What it does: When a conversation starts, an app installed on the device activates the microphone to capture voice data in real time, and also captures text messages directly from chat apps.

[1071] Step 3:

[1072] The device converts the voice data into text data using the Google Cloud Speech-to-Text API or similar. The input is the captured voice data, and the output is the converted text data. The device calls a speech recognition API to convert the voice to text.

[1073] Specific operation: The device sends voice data to the API, receives the voice recognition results, stores the results in internal memory, and passes them on to the next process.

[1074] Step 4:

[1075] The device analyzes text data using a generative AI model to assess the likelihood of sexually harassing comments. The input is text data, and the output is a probability score of sexually harassing comments. AI models such as BERT and GPT-3 are used.

[1076] How it works: The device inputs text data into the generative AI model, which then outputs a score indicating the likelihood of the comment being sexually harassing. This score is then used in the next step.

[1077] Step 5:

[1078] The device uses an emotion recognition engine to recognize the user's emotions. The input is the user's facial expression and tone of voice, and the output is the recognized emotion data. It uses the Microsoft Azure Emotion API, etc.

[1079] Specific operation: The device uses the camera and microphone connected to the device to capture the user's facial expressions and tone of voice, and sends them to the emotion recognition API. The obtained emotion data is then stored in the device's internal memory.

[1080] Step 6:

[1081] The device evaluates the probability of sexual harassment remarks based on the analysis results of the generated AI model and emotion recognition data. The input is the probability score of sexual harassment remarks and emotion data, and the output is an alert generation flag. A threshold is set, and if it is exceeded, an alert generation flag is set.

[1082] Specific operation: The analysis results of the text data (probability score) are integrated with the emotion data, and if a certain threshold is exceeded, an alert generation flag is set and the next process is carried out.

[1083] Step 7:

[1084] Generates a warning message when the device has an alert generation flag. The input is the alert generation flag and emotion data, and the output is the warning message. Generates the warning message using a Python script.

[1085] Specific operation: Prepare a warning message template and generate a customized message based on the emotion data. This message is notified to the user in the next step.

[1086] Step 8:

[1087] The terminal notifies the user of the generated warning message. The input is the warning message and the output is the notification to the user. The notification methods are visual (screen popup) and audible (voice notification).

[1088] Specific behavior: The device will display the generated warning message as a pop-up on the screen and, if necessary, provide an audio notification. This is implemented using JavaScript or the Android and iOS notification APIs.

[1089] (Application example 2)

[1090] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1091] In recent years, sexual harassment has become a problem in the workplace and online communities. However, systems that can detect sexual harassment in real time and issue appropriate warnings are not yet widely available. Furthermore, there is a need for technology that provides more accurate and effective feedback by taking user emotions into account. The present invention aims to provide a sexual harassment detection system that solves these problems.

[1092] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1093] In this invention, the server includes a data collection means, a means for capturing voice or text data, a means for converting the captured voice data into text data, a means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, a means for generating a warning message when a sexual harassment remark is detected, a means for notifying the user of the warning message, and an emotion recognition means for recognizing the user's emotion. This makes it possible to detect sexual harassment remarks in real time and issue an appropriate warning. Furthermore, by recognizing the user's emotion, it is possible to customize the warning message according to the situation and provide more effective feedback.

[1094] "Data collection means" is a function that imports data on past sexual harassment cases and stores it in a database.

[1095] A "voice or text data capture means" is a device that acquires voice or text data in real time from a microphone or chat app.

[1096] A "voice recognition means" is a technique or device that converts captured voice data into text data.

[1097] "Analysis tools" are technologies or devices that use generative AI models to analyze text data and assess the likelihood of sexual harassment statements.

[1098] A "generative AI model" is an artificial intelligence model trained based on data on past sexual harassment cases.

[1099] The "alert generation means" is a function that generates a warning message when a sexual harassment remark is detected.

[1100] The "notification means" is a function that notifies the user of the generated warning message visually or audibly.

[1101] The "emotion recognition means" is a function that recognizes the user's emotions from facial expressions, tone of voice, etc.

[1102] A specific embodiment for implementing this invention will be described below. This system monitors comments in real time when users communicate via voice or chat, and issues an immediate warning if possible sexual harassment comments are detected. Furthermore, it recognizes the user's emotions and appropriately customizes the warning message, thereby achieving effective feedback.

[1103] Data collection

[1104] The server imports past sexual harassment case data and stores it in a database. This data is used to train the generative AI model and also helps improve the accuracy of the emotion engine. For example, a dataset could include sexual harassment cases collected from various industries.

[1105] Voice and text data capture

[1106] The device captures voice or text data in real time from a microphone or chat app. In the case of voice data, it needs to be converted into text. For example, a smartphone microphone can be used to capture voice and a chat app can accept input as a text message.

[1107] Voice Recognition

[1108] The device uses voice recognition technology to convert the captured voice data into text data, using Google's voice recognition service as an example. This converted text data is then used for subsequent analysis by generative AI models and sentiment engines.

[1109] Text data analysis

[1110] The device uses a generative AI model to analyze the captured text data and evaluate the possibility of sexual harassment. The generative AI model is trained based on data from past sexual harassment cases and evaluates the content of comments with high accuracy. For example, a comment such as "You're so cute. How about going out for dinner sometime?" is judged to be potentially sexual harassment.

[1111] emotion recognition

[1112] The device uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice. The emotion data is combined with the results of text data analysis to generate the final warning message. For example, it can determine whether the user is speaking in a joking or serious tone.

[1113] Detecting sexual harassment remarks

[1114] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If the remark is judged to be sexually harassing with a high probability, it will activate an alert generation method.

[1115] Alert generation and notification

[1116] If the device determines that a remark is likely to be sexual harassment, it generates a warning message. The generated warning message incorporates emotional data provided by the emotion engine and is customized according to the situation. The device then notifies the user of the warning message visually or audibly using a notification means. For example, a warning such as "Be careful as this may be sexual harassment" may be displayed on the user's screen.

[1117] Specific examples

[1118] 1. Scenario 1: Remote Meeting Operation

[1119] If a user says, "You're so cute. Let's have dinner together sometime," during a remote meeting, the remark is captured and converted into text. A generative AI model then detects potential sexual harassment, and an emotion recognition engine recognizes the user's playful tone. As a result, a warning message appears on the user's screen saying, "Be careful, this may be sexual harassment."

[1120] 2. Scenario 2: Chat operation

[1121] If a user sends a message in a chat room saying, "New employee XX is so beautiful," this text message is captured and the generative AI model detects potential sexual harassment. The emotion recognition engine determines whether the user is speaking in a friendly or teasing manner, and a warning message is displayed on the user's chat screen saying, "Be careful, this may be sexual harassment."

[1122] Prompt Sentence Examples

[1123] Example: "You're so cute, how about we have dinner sometime?"

[1124] Example: "The new employee, Ms. XX, is very beautiful."

[1125] In this way, this system enables the detection of sexually harassing remarks and feedback in real time, taking into account the user's emotions, which is expected to lead to more accurate and effective communication.

[1126] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1127] Step 1: Data collection

[1128] The server imports past sexual harassment case data and stores it in a database. This is primarily used to improve the accuracy of the generative AI model and emotion recognition engine. Specifically, it analyzes past sexual harassment remark data collected from within the company or from public datasets and registers it in the database. The input is past sexual harassment case data, and the output is the data stored in the database.

[1129] Step 2: Capture audio and text data

[1130] The device captures voice and text data in real time from a microphone or chat app. In the case of voice data, it must be converted into text format. Specifically, the device picks up what the user says with a microphone and captures the voice data, or receives text data directly from a chat app. The input is the user's real-time voice or chat message, and the output is the captured voice data or text data.

[1131] Step 3: Voice Recognition

[1132] The device uses speech recognition technology to convert the captured voice data into text data. Here, Google's speech recognition API is used for speech-to-text conversion. The input is the captured voice data, and the output is text data.

[1133] Step 4: Analyze the text data

[1134] The device analyzes text data using a generative AI model to assess the likelihood of sexual harassment. The model is a trained generative AI model that evaluates the content of statements with high accuracy based on past data on sexual harassment cases. The input is text data, and the output is an assessment result indicating the likelihood of sexual harassment.

[1135] Step 5: Emotion Recognition

[1136] The device uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice. For example, it determines whether the user is speaking in a joking or serious tone. The input is the user's facial expression data and tone of voice data, and the output is the recognized emotion data.

[1137] Step 6: Detect sexual harassment comments

[1138] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If it determines with a high probability that the remark was sexually harassing, it activates an alert generation mechanism. The inputs are the analysis results and emotional data, and the output is the evaluation result of the remark being sexually harassing.

[1139] Step 7: Generate an alert

[1140] The device generates a warning message if it determines that a remark is likely to be sexual harassment. The generated warning message is customized by incorporating emotional data provided by an emotion recognition engine. The input is the evaluation result of the sexual harassment remark and the emotional data, and the output is a customized warning message.

[1141] Step 8: Notify users

[1142] The terminal notifies the user of the generated warning message visually or audibly. For example, it notifies the user of the warning by a screen pop-up or a voice notification. The input is the customized warning message, and the output is the warning message notified to the user.

[1143] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1144] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1145] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1146] [Fourth embodiment]

[1147] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1148] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1149] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1150] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1151] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1152] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1153] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1154] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1155] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1156] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1157] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1158] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1159] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1160] Specific Embodiments of the Invention

[1161] The present invention relates to a system that monitors conversations and chats within a company, detects possible sexual harassment comments, and issues a warning in real time. Specific embodiments for implementing this system will be described below.

[1162] Program processing

[1163] Data collection

[1164] The server imports data on past sexual harassment cases and stores it in a centrally managed database, allowing the system to learn in a way that responds to the diversity of situations and word usage.

[1165] Voice and text data capture

[1166] The device (smartphone or PC) captures voice or text data from a microphone or chat app. In the case of voice data, the next step is to convert it into text data.

[1167] Voice Recognition

[1168] The device uses voice recognition technology to convert the captured voice data into text, putting the data in a format that can be analyzed by the generative AI model.

[1169] Text data analysis

[1170] The device analyzes the text data using a generative AI model that has been trained on data from past sexual harassment cases and has the ability to score whether a comment is likely to be sexual harassment.

[1171] Detecting sexual harassment remarks

[1172] The device determines whether the remark is likely to be sexually harassing based on the scores and results obtained from the generative AI model. If the score exceeds a certain threshold during this process, the device proceeds to the next step.

[1173] Generate alerts

[1174] If the device determines that a remark is likely to be sexually harassing, it will generate a warning message, which will be customized depending on the situation and will alert the person making the remark.

[1175] User Notification

[1176] The terminal notifies the user of the generated warning message both visually (e.g., on the screen) and audibly, allowing the user to immediately become aware of their own remarks and make any necessary corrections.

[1177] Specific examples

[1178] Scenario 1: Remote conference operation

[1179] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[1180] 2. The device captures this speech and converts it into text using speech recognition technology.

[1181] 3. The device uses a generative AI model to analyze the comment as potentially sexual harassment.

[1182] 4. The device generates a warning saying, "Caution: potential sexual harassment."

[1183] 5. The device displays a warning message on User A's screen.

[1184] 6. User A realizes what he or she said and corrects himself or herself by saying, "I'm sorry, I made an inappropriate comment."

[1185] Scenario 2: Chat operation

[1186] 1. User B sends a message in a chat room saying, "New employee XX is so beautiful."

[1187] 2. The device captures this text message.

[1188] 3. The device uses a generative AI model to analyze the message and assess its potential for sexual harassment.

[1189] 4. The device generates an alert saying, "Be careful, there is a possibility of sexual harassment."

[1190] 5. The device displays an alert on User B's chat screen.

[1191] 6. User B immediately cancels the message and sends a correction message saying, "I apologize for my inappropriate comment."

[1192] In this way, by introducing this system, it is possible to prevent sexually harassing remarks in everyday communication, thereby maintaining and improving a healthy workplace environment.

[1193] The processing flow will be explained below.

[1194] Step 1:

[1195] The server imports past sexual harassment case data and stores it in a database, which is used to train generative AI models.

[1196] Step 2:

[1197] The device captures voice or text data from a microphone or chat app. In the case of voice data, it needs to be converted into text format for further processing.

[1198] Step 3:

[1199] When the device captures voice data, it uses a speech recognition engine to convert the speech into text data, which can then be analyzed by a generative AI model.

[1200] Step 4:

[1201] The device then inputs the converted text data into a generative AI model that analyzes potential sexual harassment comments. This AI model was trained based on the data collected in step 1.

[1202] Step 5:

[1203] The device evaluates the likelihood of a remark being sexually harassing based on the output from the generated AI model, and if it is determined to be sexually harassing with a high probability, it activates an alert generation method.

[1204] Step 6:

[1205] The device generates a warning message. Specifically, it creates a message indicating the possibility of sexual harassment according to a set format.

[1206] Step 7:

[1207] The device notifies the user of the generated warning message. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification.

[1208] Step 8:

[1209] The user receives a warning message and is made aware that their comments or messages may constitute sexual harassment. The user can correct their comments or take appropriate action.

[1210] Step 9:

[1211] The server continuously updates the generative AI model by collecting newly detected sexual harassment case data and adding it to the database, thereby improving the model's accuracy and effectiveness over time.

[1212] Example 1

[1213] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1214] Preventing sexual harassment in the modern workplace is a difficult task. Traditional methods require time for sexual harassment comments to be reported externally, so a prompt response is required to maintain a healthy workplace. Real-time monitoring and warnings are necessary, especially with the increase in remote work and online chat, but current systems tend to be slow to respond.

[1215] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1216] In this invention, the server includes a means for collecting data, a means for capturing voice or text data, a speech recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, a means for inputting a prompt sentence into the generative AI model, an alert generation means for generating a warning message when sexual harassment remarks are detected, and a notification means for notifying the user of the warning message. This enables real-time collection and analysis of data, and rapid generation and notification of warning messages.

[1217] "Means for collecting data" refers to systems or devices for importing and centrally managing data on past sexual harassment cases and related information.

[1218] "Audio or text data capture means" means a system or device for capturing and storing conversations or text messages in real time.

[1219] "Speech recognition means" refers to the technology or software used to convert captured voice data into text.

[1220] "Analytical means using generative AI models" refers to systems or devices that use AI models trained based on past cases of sexual harassment to analyze text data and evaluate the likelihood of sexual harassment remarks.

[1221] A "means for inputting prompts to a generative AI model" is an input interface or software for providing specific instructions or questions to an AI model.

[1222] An "alert generation means" is a system or device that generates a warning message when a sexually harassing remark is detected.

[1223] The "notification means" is a system or device for visually and audibly notifying the user of the generated warning message.

[1224] MODE FOR CARRYING OUT THE INVENTION

[1225] The present invention relates to a system that monitors conversations and chats within a company, detects possible sexual harassment comments, and issues a warning in real time. Specific embodiments for implementing this system will be described below.

[1226] The server is primarily used to collect data on past sexual harassment cases and store it in a centralized database. This system can use MySQL or PostgreSQL as the database. The collected data is used as training data for the generative AI model, and data preprocessing is performed to improve the model's accuracy. Preprocessing includes filling in missing values ​​and normalizing the data.

[1227] The device (smartphone or personal computer) is used to capture voice or text data in real time from remote meetings or chat apps. Voice can be recorded through a microphone and saved in WAV or MP3 format. Text data can also be obtained using the chat app's API.

[1228] For speech recognition, we use voice recognition technologies such as Google Cloud Speech-to-Text and Amazon Transcribe. This converts the audio data into text, which is then formatted for easy analysis by the generative AI model. The converted text data is then stripped of unnecessary spaces and special characters to prepare it for analysis.

[1229] Text data is analyzed in real time using an analytical method that uses a generative AI model (e.g., OpenAI GPT-4). During this process, a prompt is input into the generative AI model, which evaluates whether the remark constitutes sexual harassment. Including specific examples in the prompt allows for a more accurate judgment.

[1230] Some common prompts are:

[1231] Example prompt 1:

[1232] "Please determine whether the following statements constitute sexual harassment:

[1233] "You're so cute, how about we have dinner sometime?"

[1234] Example prompt 2:

[1235] "Please determine whether the following statements constitute sexual harassment:

[1236] "New employee XX is really beautiful."

[1237] If the analysis results of the generative AI model exceed a threshold, a warning message is generated as an alert generation means. This message includes specific instructions and warns the speaker that the comment is inappropriate.

[1238] The warning message is notified to the user through a notification means, which is visual (for example, a screen display) and audio, so that the user can immediately pay attention to the message.

[1239] As a result, the system can detect sexually harassing remarks in real time and take prompt action to maintain a healthy work environment.

[1240] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1241] Program processing flow

[1242] Step 1:

[1243] The server collects data on past sexual harassment cases and stores it in a centrally managed database.

[1244] Input: Historical sexual harassment case data from a company database.

[1245] Data processing: Preprocessing of collected data (filling in missing values ​​and normalizing data).

[1246] Output: Data stored as a training dataset for a generative AI model.

[1247] How it works: The server uses SQL queries to extract sexual harassment cases from the company's internal database over the past five years, imports the data into a MySQL database, and performs preprocessing such as missing value imputation and normalization to prepare a training dataset.

[1248] Step 2:

[1249] The device captures voice or text data in real time from conversations and chat apps.

[1250] Input: Audio data from a remote meeting or text messages from a chat app.

[1251] Data processing: Audio data is saved in WAV or MP3 format. Text data is saved in a temporary folder.

[1252] Output: The captured audio or text data.

[1253] Specific operation: During remote meetings, the device records audio data in real time in WAV format through the microphone, and uses the chat app's API to capture all messages sent in text format and save them in a temporary folder.

[1254] Step 3:

[1255] The terminal uses voice recognition technology to convert the captured voice data into text.

[1256] Input: Captured audio data (WAV or MP3 format).

[1257] Data processing: Convert speech to text using Google Cloud Speech-to-Text or Amazon Transcribe. Remove unnecessary spaces and special characters.

[1258] Output: The converted text data.

[1259] Specific operation: The device calls the Google Cloud Speech-to-Text API to convert the WAV audio file into a string, then removes unnecessary spaces and special characters from the converted text data in preparation for analysis.

[1260] Step 4:

[1261] The device uses a generative AI model to analyze the text data.

[1262] Input: The converted text data.

[1263] Data processing: A prompt was generated and input into a generative AI model (OpenAI GPT-4). This prompt included instructions to evaluate whether the remark constituted sexual harassment.

[1264] Output: Scoring and analysis results from the generative AI model.

[1265] Specific operation: The device generates the following prompt sentence and inputs it into the AI ​​model:

[1266] "Please determine whether the following statements constitute sexual harassment:

[1267] "You're so cute, how about we have dinner sometime?"

[1268] Obtain the scoring and analysis results returned by the AI ​​model and format the data.

[1269] Step 5:

[1270] The device determines whether a comment is likely to be sexual harassment based on the analysis results obtained from the generative AI model.

[1271] Input: Scoring and analysis results from the AI ​​model.

[1272] Data processing: Determine whether the score exceeds a predefined threshold.

[1273] Output: Triggers an alert generator when a threshold is exceeded.

[1274] Specific operation: The device checks whether the score returned by the AI ​​model is 80 or above (pre-set threshold), and if so, proceeds to the next alert generation step.

[1275] Step 6:

[1276] The terminal generates a warning message if a sexually harassing remark is detected.

[1277] Input: Speech data with scores above a threshold.

[1278] Data processing: Generation of customized warning messages.

[1279] Output: The warning message generated.

[1280] Specific behavior: If the device determines that a comment is sexually harassing, it will generate a warning message stating, "This comment may be inappropriate. Please correct it."

[1281] Step 7:

[1282] The terminal notifies the user of the generated warning message.

[1283] Input: The generated warning message.

[1284] Data processing: Preparation of visual and audio notifications.

[1285] Output: Notification to the user.

[1286] Specific operation: The device will display a warning message on the user's screen and simultaneously issue an audio alert to draw attention to the statement, allowing the user to immediately pay attention to the statement and make corrections.

[1287] (Application example 1)

[1288] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1289] In conventional workplaces, there was no system that could detect sexual harassment remarks in real time and issue a warning. As a result, when a victim received unpleasant remarks, it was difficult to take immediate action, making it difficult to maintain a healthy work environment. There was also a risk that employees who might make sexually harassing remarks would cause problems without realizing that they were making such remarks.

[1290] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1291] In this invention, the server includes a data collection means, a voice or text data capture means, a voice recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, an alert generation means for generating a warning message when sexual harassment remarks are detected, a notification means for notifying a user of the warning message, and a real-time notification means for capturing data in real time using a smart device and immediately notifying a user of the warning message based on the results of the data analysis. This makes it possible to monitor conversations and chats in the office in real time, immediately detect sexual harassment remarks, issue a warning to employees, and encourage them to take prompt action.

[1292] "Data collection means" refers to a means for collecting information such as voice data and text data and saving it in a format required for subsequent processing.

[1293] "Audio or text data capture means" means a device or software for capturing audio or text data in real time.

[1294] "Speech recognition means" is a technology that converts captured voice data into text data.

[1295] The "analysis means" is a means for analyzing collected text data and evaluating and classifying it based on specific conditions.

[1296] A "generative AI model" is an artificial intelligence model that learns from past data and makes predictions and classifications for new data.

[1297] An "alert generator" is a device or software for generating a warning message when a particular condition is met.

[1298] The "notification means" is a means for transmitting the generated warning message to the user.

[1299] A "smart device" is a mobile terminal equipped with advanced computing power and communication functions.

[1300] "Real-time data capture" is the act of collecting data in synchronization with real time.

[1301] "Real-time notification means" refers to a means for transmitting a warning message to the user in a timely manner according to the analysis results.

[1302] This invention relates to a system for monitoring conversations and chats in a workplace and detecting and issuing warnings about sexual harassment comments in real time. Specific embodiments for implementing this system will be described below.

[1303] The system includes a data collection means, a voice or text data capture means, a voice recognition means, an analysis means, an alert generation means, a notification means, real-time data capture using a smart device, and a real-time notification means.

[1304] Server processing

[1305] The server processes data using the following hardware and software:

[1306] Hardware: Server

[1307] Software: Database (MySQL), Generative AI model (OpenAI GPT-3)

[1308] As a means of collecting data, past sexual harassment cases are stored in a database, and the generative AI model learns from this data. New sexual harassment cases are continually added to the database, and it is updated with the latest data.

[1309] What the device is doing

[1310] The terminal uses the following hardware and software:

[1311] Hardware: Smartphone, microphone

[1312] Software: Speech recognition API (Google Speech-to-Text)

[1313] The smartphone's microphone is used to capture voice or text data, and office conversations are captured in real time. The voice data is converted into text data using a speech recognition API, and the converted text data is sent to the server.

[1314] Server text data analysis

[1315] The server uses a generative AI model to analyze text data and score it for potential sexual harassment. This analysis is performed in real time based on a historical database, and a warning message is generated if the score exceeds a certain threshold.

[1316] Alerting and Notifications

[1317] As an alert generation means, the server generates a warning message and immediately notifies the user through a notification means, which is visually and audibly displayed on the smartphone in real time.

[1318] Specific examples

[1319] For example, if User A says, "You're so cute," the smartphone's microphone captures this remark. The speech data is converted into text data using a speech recognition API (Google Speech-to-Text). This text data is sent to a server and analyzed by a generative AI model (OpenAI GPT-3). If the remark is determined to be highly likely to be sexual harassment, a warning message is generated and a message stating, "Please be careful as this may be sexual harassment," is immediately displayed on User A's smartphone.

[1320] Prompt Sentence Examples

[1321] For example, the input prompt for a generative AI model might look like this:

[1322] "New employee XX is so beautiful."

[1323] Based on this prompt, the generative AI model evaluates the statement and issues a warning if necessary.

[1324] As described above, the present invention provides a system for preventing sexual harassment remarks in a workplace environment and maintaining healthy communication.

[1325] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1326] Step 1:

[1327] The device performs audio capture. To obtain audio data, the device (smartphone or PC) constantly monitors the ambient sound using a microphone. The input is the user's voice, and the output is audio data.

[1328] Step 2:

[1329] The device performs speech recognition and converts the speech data into text data. The acquired speech data is converted into text using a speech recognition API (Google Speech-to-Text). The input is speech data, and the output is text data.

[1330] Step 3:

[1331] The terminal sends the converted text data to the server. The text data is sent to the server via the HTTPS protocol. The input is text data, and the output is data sent to the server.

[1332] Step 4:

[1333] The server analyzes the text data using a generative AI model. Based on the received text data, the server uses a generative AI model (OpenAI GPT-3) to score the likelihood of sexual harassment remarks. At this time, a prompt sentence may be used as input. The input is text data, and the output is a score of the likelihood of sexual harassment remarks.

[1334] Step 5:

[1335] The server generates a warning message based on the analysis results. If the scoring result by the generative AI model exceeds a certain threshold, the server generates a customized warning message. The input is the likelihood score of sexual harassment remarks, and the output is a warning message.

[1336] Step 6:

[1337] The server sends a warning message to the terminal. The generated warning message is sent to the terminal via HTTPS protocol and notified in real time. The input is the warning message and the output is the data sent to the terminal.

[1338] Step 7:

[1339] The terminal displays a warning message to the user using a notification means. The terminal notifies the user of the received warning message visually or audibly. Specifically, the warning message is displayed on the screen and an audio notification is played. The input is the sent warning message, and the output is a notification message to the user.

[1340] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1341] Specific Embodiments of the Invention

[1342] This invention relates to a system that monitors conversations and chats, detects possible sexual harassment comments, and issues a warning in real time, and further includes a function to recognize the user's emotions and provide appropriate feedback. Specific embodiments for implementing this system are described below.

[1343] Program processing

[1344] Data collection

[1345] The server imports past sexual harassment case data and stores it in a database, which is used to train the generative AI model and improve the accuracy of the emotion engine.

[1346] Voice and text data capture

[1347] The device (smartphone or PC) captures voice or text data in real time from a microphone or chat app. In the case of voice data, this is then converted into text format.

[1348] Voice Recognition

[1349] The device uses voice recognition technology to convert the captured voice data into text data, which is then used for subsequent analysis by generative AI models and emotion engines.

[1350] Text data analysis

[1351] The device analyzes the text data and assesses the likelihood of sexual harassment using a generative AI model, which is trained on data from past sexual harassment cases.

[1352] emotion recognition

[1353] The device uses an emotion engine to recognize emotions from the user's facial expressions, tone of voice, etc. Emotion data is combined with the results of text data analysis to generate the final warning message.

[1354] Detecting sexual harassment remarks

[1355] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If the remark is judged to be sexually harassing with a high probability, it will activate an alert generation method.

[1356] Generate alerts

[1357] If the device determines that a remark is likely to be sexually harassing, it generates a warning message that is customized to the situation by incorporating emotional data provided by the emotion engine.

[1358] User Notification

[1359] The device notifies the user of the generated warning message. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification. The user is immediately aware of what they said and can make any necessary corrections.

[1360] Specific examples

[1361] Scenario 1: Remote conference operation

[1362] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[1363] 2. The device captures this speech and converts it into text using speech recognition technology.

[1364] 3. The device uses a generative AI model to analyze the comment and determine that it may be sexual harassment.

[1365] 4. The device uses the emotion engine to recognize User A's emotions, for example, determining whether User A is speaking in a joking or serious tone.

[1366] 5. The device generates a warning saying, "Be careful, this may be sexual harassment," and customizes the tone and content based on data from the emotion engine.

[1367] 6. The device displays a warning message on User A's screen.

[1368] 7. User A recognizes what he said and corrects himself by saying, "I'm sorry, I made an inappropriate comment."

[1369] Scenario 2: Chat operation

[1370] 1. User B sends a message in a chat room saying, "New employee XX is so beautiful."

[1371] 2. The device captures this text message.

[1372] 3. The device uses a generative AI model to analyze the message and determine that it may be sexually harassing.

[1373] 4. The device uses an emotion engine to recognize User B's emotions. For example, it determines whether User B is speaking in a friendly manner or teasing the user.

[1374] 5. The device generates a warning saying, "Be careful, this may be sexual harassment," and customizes the tone and content based on data from the emotion engine.

[1375] 6. The device displays an alert on User B's chat screen.

[1376] 7. User B immediately cancels the message and sends a correction message saying, "I apologize for my inappropriate comment."

[1377] This system enables the detection and feedback of sexually harassing remarks taking into account the user's emotions, which is expected to lead to more accurate and effective communication.

[1378] The processing flow will be explained below.

[1379] Step 1:

[1380] The server imports historical sexual harassment case data and stores it in a database, which is used to train generative AI models and emotion engines.

[1381] Step 2:

[1382] The device captures voice or text data from the microphone or chat app, allowing users to collect their conversations and messages in real time.

[1383] Step 3:

[1384] When the device receives voice data, it uses a speech recognition engine to convert the voice into text data, which prepares the data in a format that can be analyzed.

[1385] Step 4:

[1386] The device inputs the text data into a generative AI model that analyzes the possibility of sexual harassment remarks. The generative AI model is trained from data on past sexual harassment cases.

[1387] Step 5:

[1388] The device captures the user's facial expressions and tone of voice, and uses an emotion engine to recognize the user's emotions, allowing it to understand the user's mental state and emotions.

[1389] Step 6:

[1390] The device combines the analysis results from the generative AI model with the emotional data from the emotion engine to assess the likelihood of sexual harassment. By combining the two sets of data, a more accurate judgment can be made.

[1391] Step 7:

[1392] If the device determines that a comment is likely to be sexually harassing, it will generate a warning message, dynamically adjusting the content and tone of the message based on the emotional data.

[1393] Step 8:

[1394] The terminal will notify the user of the generated warning message visually and audibly. The user will be alerted by a pop-up display and / or an audio notification.

[1395] Step 9:

[1396] Users will receive a warning message and be made aware that their comments or actions may constitute sexual harassment. Users can correct or retract their comments as necessary.

[1397] Step 10:

[1398] The server collects newly detected sexual harassment cases and emotion data and adds them to the database, which continuously updates the generative AI model and emotion engine, improving the accuracy of the entire system.

[1399] Example 2

[1400] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1401] Sexually harassing remarks have become a problem in modern remote conference and chat environments. However, there are few systems that can detect them in real time and take immediate countermeasures. It has also been pointed out that there is a lack of systems that not only analyze the content of remarks but also recognize the speaker's emotions and provide appropriate feedback. The purpose of this invention is to solve these problems and provide a system that effectively detects and warns about sexually harassing remarks.

[1402] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1403] In this invention, the server includes a data collection means, a voice or text data capture means, a voice recognition means for converting the captured voice data into text data, an analysis means for analyzing the text data and using a generative AI model to detect possible sexual harassment expressions, an emotion recognition means for recognizing a user's emotions and integrating them with the analysis results, an alert generation means for generating a warning message when possible sexual harassment expressions are detected, and a notification means for notifying the user of the warning message. This enables effective detection and warning of sexual harassment remarks in real time and appropriate feedback that takes the user's emotions into consideration.

[1404] "Data Collection Measures" refers to the devices and processes used to import and store historical sexual harassment case data in a database.

[1405] "Voice or text data capture means" refers to devices and processes that obtain voice or text data through a microphone or chat app.

[1406] "Speech recognition means" refers to the technology and processes that convert captured voice data into text data.

[1407] A "generative AI model" refers to an artificial intelligence model that is trained based on data from past sexual harassment cases and analyzes text data to detect possible sexual harassment remarks.

[1408] "Analysis means" refers to the devices and processes that use generative AI models to analyze text data and assess potential sexual harassment comments.

[1409] "Emotion recognition means" refers to technology and processes for recognizing emotions from a user's facial expressions and tone of voice and generating emotion data.

[1410] "Alert generation means" refers to a device and process that generates a warning message when it determines that a sexually harassing remark is likely.

[1411] "Notification means" refers to devices and processes for notifying a user of generated warning messages.

[1412] MODE FOR CARRYING OUT THE INVENTION

[1413] The present invention is a system that monitors conversations and chats in real time, detects possible sexual harassment comments, and issues a warning. It also includes a function to recognize the user's emotions and provide appropriate feedback. A specific embodiment of this system will be described below.

[1414] Data collection methods

[1415] The server imports past sexual harassment case data and stores it in a database. This data contains a wide variety of sexual harassment cases and is used to train the generative AI model and emotion recognition engine. The server implements data collection and storage processes using Python and SQL, which significantly improves the accuracy and effectiveness of the system.

[1416] A means of capturing voice or text data

[1417] The device (smartphone or PC) captures voice or text data in real time through a microphone or chat app. Voice data is converted into text using speech recognition technology such as the Google Cloud Speech-to-Text API. This allows the spoken content to be obtained as text and used for subsequent processing.

[1418] Voice recognition means

[1419] The device uses voice recognition technology to convert the captured audio data into text data, using services such as Google Cloud Speech-to-Text, which is then immediately stored internally and passed on to the next analysis step.

[1420] Analytical tools using generative AI models

[1421] The device analyzes text data using a generative AI model to assess the likelihood of sexual harassment. This AI model is trained based on past data on sexual harassment cases and can evaluate the content of statements with high accuracy. For example, it uses advanced generative AI models such as BERT and GPT-3.

[1422] emotion recognition means

[1423] The device uses an emotion recognition engine to recognize emotions from the user's facial expressions and tone of voice, utilizing the Microsoft Azure Emotion API, among other tools. By acquiring emotion data, the device can more accurately understand the intention and nuance of what is being said.

[1424] Alert generation method

[1425] The device evaluates the likelihood of sexual harassment based on the analysis results of the generative AI model and data from the emotion recognition engine. If there is a high probability that the remark is sexual harassment, it automatically generates a warning message. This message is customized to an appropriate tone and content based on the emotion data.

[1426] Notification means

[1427] The terminal generates a warning message and notifies the user. This notification takes the form of a visual display (e.g., a screen pop-up) and an audio notification. In particular, if the message is urgent, it can immediately draw the user's attention.

[1428] Specific operation example

[1429] Scenario 1: Remote conference operation

[1430] 1. User A says during a remote meeting, "You're so cute. How about we have dinner sometime?"

[1431] 2. The device captures this speech and converts it into text using speech recognition technology (powered by the Google Cloud Speech-to-Text API).

[1432] 3. The device analyzes the comment using a generative AI model (BERT, GPT-3, etc.) and determines that it may be sexual harassment.

[1433] 4. The device uses an emotion engine (such as Azure Emotion API) to recognize User A's emotions.

[1434] 5. The device generates a warning saying, "Beware of potential sexual harassment," and customizes the tone and content accordingly.

[1435] 6. The device pops up a warning message on User A's screen.

[1436] 7. User A corrects himself by saying, "I'm sorry, I made an inappropriate comment."

[1437] Examples of prompt statements

[1438] During a remote meeting, User A says, "You're so cute. How about we go out for dinner sometime?" Capture this remark and analyze it using a sexual harassment remark detection system. Explain the process of generating an appropriate warning message based on the analysis results of the generative AI model and emotion recognition results, and notifying User A.

[1439] This system can detect sexual harassment remarks in real time and provide appropriate warning messages that take emotions into account. This configuration can improve effective communication.

[1440] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1441] Step 1:

[1442] The server imports data on past sexual harassment cases and stores it in a database. The input is data on past sexual harassment cases, and the output is the data stored in the database. A data collection program using Python and SQL reads and saves the initial data.

[1443] What happens: The server periodically checks the data source (e.g., CSV file or external API) and imports any new data it finds. The imported data is then stored in the database.

[1444] Step 2:

[1445] The device captures voice or text data in real time through a microphone or chat app. The input is the user's speech data, and the output is the captured voice file or text data. The data is acquired in real time using JavaScript or Python libraries.

[1446] What it does: When a conversation starts, an app installed on the device activates the microphone to capture voice data in real time, and also captures text messages directly from chat apps.

[1447] Step 3:

[1448] The device converts the voice data into text data using the Google Cloud Speech-to-Text API or similar. The input is the captured voice data, and the output is the converted text data. The device calls a speech recognition API to convert the voice to text.

[1449] Specific operation: The device sends voice data to the API, receives the voice recognition results, stores the results in internal memory, and passes them on to the next process.

[1450] Step 4:

[1451] The device analyzes text data using a generative AI model to assess the likelihood of sexually harassing comments. The input is text data, and the output is a probability score of sexually harassing comments. AI models such as BERT and GPT-3 are used.

[1452] How it works: The device inputs text data into the generative AI model, which then outputs a score indicating the likelihood of the comment being sexually harassing. This score is then used in the next step.

[1453] Step 5:

[1454] The device uses an emotion recognition engine to recognize the user's emotions. The input is the user's facial expression and tone of voice, and the output is the recognized emotion data. It uses the Microsoft Azure Emotion API, etc.

[1455] Specific operation: The device uses the camera and microphone connected to the device to capture the user's facial expressions and tone of voice, and sends them to the emotion recognition API. The obtained emotion data is then stored in the device's internal memory.

[1456] Step 6:

[1457] The device evaluates the probability of sexual harassment remarks based on the analysis results of the generated AI model and emotion recognition data. The input is the probability score of sexual harassment remarks and emotion data, and the output is an alert generation flag. A threshold is set, and if it is exceeded, an alert generation flag is set.

[1458] Specific operation: The analysis results of the text data (probability score) are integrated with the emotion data, and if a certain threshold is exceeded, an alert generation flag is set and the next process is carried out.

[1459] Step 7:

[1460] Generates a warning message when the device has an alert generation flag. The input is the alert generation flag and emotion data, and the output is the warning message. Generates the warning message using a Python script.

[1461] Specific operation: Prepare a warning message template and generate a customized message based on the emotion data. This message is notified to the user in the next step.

[1462] Step 8:

[1463] The terminal notifies the user of the generated warning message. The input is the warning message and the output is the notification to the user. The notification methods are visual (screen popup) and audible (voice notification).

[1464] Specific behavior: The device will display the generated warning message as a pop-up on the screen and, if necessary, provide an audio notification. This is implemented using JavaScript or the Android and iOS notification APIs.

[1465] (Application example 2)

[1466] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1467] In recent years, sexual harassment has become a problem in the workplace and online communities. However, systems that can detect sexual harassment in real time and issue appropriate warnings are not yet widely available. Furthermore, there is a need for technology that provides more accurate and effective feedback by taking user emotions into account. The present invention aims to provide a sexual harassment detection system that solves these problems.

[1468] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1469] In this invention, the server includes a data collection means, a means for capturing voice or text data, a means for converting the captured voice data into text data, a means for analyzing the text data and using a generative AI model to detect possible sexual harassment remarks, a means for generating a warning message when a sexual harassment remark is detected, a means for notifying the user of the warning message, and an emotion recognition means for recognizing the user's emotion. This makes it possible to detect sexual harassment remarks in real time and issue an appropriate warning. Furthermore, by recognizing the user's emotion, it is possible to customize the warning message according to the situation and provide more effective feedback.

[1470] "Data collection means" is a function that imports data on past sexual harassment cases and stores it in a database.

[1471] A "voice or text data capture means" is a device that acquires voice or text data in real time from a microphone or chat app.

[1472] A "voice recognition means" is a technique or device that converts captured voice data into text data.

[1473] "Analysis tools" are technologies or devices that use generative AI models to analyze text data and assess the likelihood of sexual harassment statements.

[1474] A "generative AI model" is an artificial intelligence model trained based on data on past sexual harassment cases.

[1475] The "alert generation means" is a function that generates a warning message when a sexual harassment remark is detected.

[1476] The "notification means" is a function that notifies the user of the generated warning message visually or audibly.

[1477] The "emotion recognition means" is a function that recognizes the user's emotions from facial expressions, tone of voice, etc.

[1478] A specific embodiment for implementing this invention will be described below. This system monitors comments in real time when users communicate via voice or chat, and issues an immediate warning if possible sexual harassment comments are detected. Furthermore, it recognizes the user's emotions and appropriately customizes the warning message, thereby achieving effective feedback.

[1479] Data collection

[1480] The server imports past sexual harassment case data and stores it in a database. This data is used to train the generative AI model and also helps improve the accuracy of the emotion engine. For example, a dataset could include sexual harassment cases collected from various industries.

[1481] Voice and text data capture

[1482] The device captures voice or text data in real time from a microphone or chat app. In the case of voice data, it needs to be converted into text. For example, a smartphone microphone can be used to capture voice and a chat app can accept input as a text message.

[1483] Voice Recognition

[1484] The device uses voice recognition technology to convert the captured voice data into text data, using Google's voice recognition service as an example. This converted text data is then used for subsequent analysis by generative AI models and sentiment engines.

[1485] Text data analysis

[1486] The device uses a generative AI model to analyze the captured text data and evaluate the possibility of sexual harassment. The generative AI model is trained based on data from past sexual harassment cases and evaluates the content of comments with high accuracy. For example, a comment such as "You're so cute. How about going out for dinner sometime?" is judged to be potentially sexual harassment.

[1487] emotion recognition

[1488] The device uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice. The emotion data is combined with the results of text data analysis to generate the final warning message. For example, it can determine whether the user is speaking in a joking or serious tone.

[1489] Detecting sexual harassment remarks

[1490] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If the remark is judged to be sexually harassing with a high probability, it will activate an alert generation method.

[1491] Alert generation and notification

[1492] If the device determines that a remark is likely to be sexual harassment, it generates a warning message. The generated warning message incorporates emotional data provided by the emotion engine and is customized according to the situation. The device then notifies the user of the warning message visually or audibly using a notification means. For example, a warning such as "Be careful as this may be sexual harassment" may be displayed on the user's screen.

[1493] Specific examples

[1494] 1. Scenario 1: Remote Meeting Operation

[1495] If a user says, "You're so cute. Let's have dinner together sometime," during a remote meeting, the remark is captured and converted into text. A generative AI model then detects potential sexual harassment, and an emotion recognition engine recognizes the user's playful tone. As a result, a warning message appears on the user's screen saying, "Be careful, this may be sexual harassment."

[1496] 2. Scenario 2: Chat operation

[1497] If a user sends a message in a chat room saying, "New employee XX is so beautiful," this text message is captured and the generative AI model detects potential sexual harassment. The emotion recognition engine determines whether the user is speaking in a friendly or teasing manner, and a warning message is displayed on the user's chat screen saying, "Be careful, this may be sexual harassment."

[1498] Prompt Sentence Examples

[1499] Example: "You're so cute, how about we have dinner sometime?"

[1500] Example: "The new employee, Ms. XX, is very beautiful."

[1501] In this way, this system enables the detection of sexually harassing remarks and feedback in real time, taking into account the user's emotions, which is expected to lead to more accurate and effective communication.

[1502] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1503] Step 1: Data collection

[1504] The server imports past sexual harassment case data and stores it in a database. This is primarily used to improve the accuracy of the generative AI model and emotion recognition engine. Specifically, it analyzes past sexual harassment remark data collected from within the company or from public datasets and registers it in the database. The input is past sexual harassment case data, and the output is the data stored in the database.

[1505] Step 2: Capture audio and text data

[1506] The device captures voice and text data in real time from a microphone or chat app. In the case of voice data, it must be converted into text format. Specifically, the device picks up what the user says with a microphone and captures the voice data, or receives text data directly from a chat app. The input is the user's real-time voice or chat message, and the output is the captured voice data or text data.

[1507] Step 3: Voice Recognition

[1508] The device uses speech recognition technology to convert the captured voice data into text data. Here, Google's speech recognition API is used for speech-to-text conversion. The input is the captured voice data, and the output is text data.

[1509] Step 4: Analyze the text data

[1510] The device analyzes text data using a generative AI model to assess the likelihood of sexual harassment. The model is a trained generative AI model that evaluates the content of statements with high accuracy based on past data on sexual harassment cases. The input is text data, and the output is an assessment result indicating the likelihood of sexual harassment.

[1511] Step 5: Emotion Recognition

[1512] The device uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice. For example, it determines whether the user is speaking in a joking or serious tone. The input is the user's facial expression data and tone of voice data, and the output is the recognized emotion data.

[1513] Step 6: Detect sexual harassment comments

[1514] The device evaluates the likelihood of a remark being sexually harassing based on the analysis results from the generative AI model and the emotional data from the emotion engine. If it determines with a high probability that the remark was sexually harassing, it activates an alert generation mechanism. The inputs are the analysis results and emotional data, and the output is the evaluation result of the remark being sexually harassing.

[1515] Step 7: Generate an alert

[1516] The device generates a warning message if it determines that a remark is likely to be sexual harassment. The generated warning message is customized by incorporating emotional data provided by an emotion recognition engine. The input is the evaluation result of the sexual harassment remark and the emotional data, and the output is a customized warning message.

[1517] Step 8: Notify users

[1518] The terminal notifies the user of the generated warning message visually or audibly. For example, it notifies the user of the warning by a screen pop-up or a voice notification. The input is the customized warning message, and the output is the warning message notified to the user.

[1519] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1520] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1521] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1522] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1523] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1524] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1525] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1526] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1527] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1528] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1529] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1530] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1531] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1532] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1533] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1534] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1535] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1536] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1537] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1538] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1539] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1540] The following is further disclosed regarding the above embodiment.

[1541] (Claim 1)

[1542] data collection means;

[1543] a means for capturing voice or text data;

[1544] speech recognition means for converting the captured voice data into text data;

[1545] an analytical method that uses a generative AI model to analyze text data and detect possible sexual harassment comments;

[1546] an alert generating means for generating a warning message when a sexual harassment remark is detected;

[1547] a notification means for notifying a user of a warning message;

[1548] A system including:

[1549] (Claim 2)

[1550] 10. The system of claim 1, which uses the database to continually learn new cases of sexual harassment.

[1551] (Claim 3)

[1552] 2. The system according to claim 1, wherein the notification means notifies the user of the warning message visually and audibly.

[1553] "Example 1"

[1554] (Claim 1)

[1555] a means of collecting data;

[1556] a means for capturing audio or text data;

[1557] speech recognition means for converting the captured voice data into text data;

[1558] an analytical method that uses a generative AI model to analyze text data and detect possible sexual harassment comments;

[1559] a means for inputting a prompt to the generative AI model;

[1560] an alert generating means for generating a warning message when a sexual harassment remark is detected;

[1561] a notification means for notifying a user of a warning message;

[1562] A system including:

[1563] (Claim 2)

[1564] 10. The system of claim 1, which uses the database to continually learn new cases of sexual harassment.

[1565] (Claim 3)

[1566] 2. The system according to claim 1, wherein the notification means notifies the user of the warning message visually and audibly.

[1567] "Application Example 1"

[1568] (Claim 1)

[1569] data collection means;

[1570] a means for capturing voice or text data;

[1571] speech recognition means for converting the captured voice data into text data;

[1572] an analytical method that uses a generative AI model to analyze text data and detect possible sexual harassment comments;

[1573] an alert generating means for generating a warning message when a sexual harassment remark is detected;

[1574] a notification means for notifying a user of a warning message;

[1575] Real-time data capture using smart devices and

[1576] a real-time notification means for instantly notifying a warning message based on the data analysis results;

[1577] A system including:

[1578] (Claim 2)

[1579] 10. The system of claim 1, which uses the database to continually learn new cases of sexual harassment.

[1580] (Claim 3)

[1581] 2. The system according to claim 1, wherein the notification means notifies the user of the warning message visually and audibly.

[1582] "Example 2: Combining Emotion Engines"

[1583] (Claim 1)

[1584] data collection means;

[1585] a means for capturing voice or text data;

[1586] speech recognition means for converting the captured voice data into text data;

[1587] an analytical means that uses a generative AI model to analyze text data and detect potential sexual harassment expressions;

[1588] emotion recognition means for recognizing a user's emotion and integrating it with the analysis result;

[1589] an alert generating means for generating a warning message when a possible sexual harassment expression is detected;

[1590] a notification means for notifying a user of a warning message;

[1591] A system including:

[1592] (Claim 2)

[1593] 10. The system of claim 1, which uses the database to continually learn new cases of sexual harassment.

[1594] (Claim 3)

[1595] 2. The system according to claim 1, wherein the notification means notifies the user of the warning message visually and audibly.

[1596] "Application example 2 when combining emotion engines"

[1597] (Claim 1)

[1598] data collection means;

[1599] a means for capturing voice or text data;

[1600] speech recognition means for converting the captured voice data into text data;

[1601] an analytical method that uses a generative AI model to analyze text data and detect possible sexual harassment comments;

[1602] an alert generating means for generating a warning message when a sexual harassment remark is detected;

[1603] a notification means for notifying a user of a warning message;

[1604] emotion recognition means for recognizing an emotion of a user;

[1605] A system including:

[1606] (Claim 2)

[1607] 10. The system of claim 1, which uses the database to continually learn new cases of sexual harassment.

[1608] (Claim 3)

[1609] 2. The system according to claim 1, wherein the notification means notifies the user of the warning message visually and audibly. [Explanation of symbols]

[1610] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. data collection means; a means for capturing voice or text data; speech recognition means for converting the captured voice data into text data; an analytical method that uses a generative AI model to analyze text data and detect possible sexual harassment comments; an alert generating means for generating a warning message when a sexual harassment remark is detected; a notification means for notifying a user of a warning message; A system including:

2. 10. The system of claim 1, wherein the system uses a database to continuously learn new cases of sexual harassment.

3. 2. The system according to claim 1, wherein the notification means notifies the user of the warning message visually and audibly.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A