System

A system that collects and analyzes voice data for inappropriate behavior using a generative AI model to detect and prevent maternity harassment, ensuring a safer workplace by issuing real-time alerts.

JP2026030689APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024133673
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Existing systems fail to effectively detect and prevent maternity harassment in the workplace, which can cause significant mental and physical harm to employees, often occurring unconsciously and difficult to address in real time.

Method used

A system that collects voice data, converts it into text, analyzes for inappropriate behavior using a generative AI model, and generates alerts to notify users and administrators of potential maternity harassment.

Benefits of technology

Enables early detection and prevention of inappropriate behavior related to pregnancy, childbirth, and childcare, creating a safer work environment by providing immediate alerts and awareness to perpetrators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026030689000001_ABST
    Figure 2026030689000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting voice data; means for transmitting the collected voice data to a server; means for converting the transmitted voice data into text data; means for analyzing the converted text data and detecting an inappropriate speech or behavior; means for generating an alert message based on a detection result and notifying a terminal of the alert message; and means for displaying the notified alert message on a user interface.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] ---

[0005] The problem that this invention aims to solve is to provide a means for early detection and prevention of inappropriate behavior related to pregnancy, childbirth, and childcare in the workplace, known as maternity harassment. Specifically, because maternity harassment is often committed unconsciously and can cause significant mental and physical damage to victims, the aim is to create a work environment in which maternity harassment can be monitored in real time and an alert can be issued to raise awareness among perpetrators and enable them to take appropriate action. [Means for solving the problem]

[0006] The present invention solves the above-mentioned problems with a system that includes the following means: means for collecting voice data, means for transmitting the collected voice data to a server, means for converting the transmitted voice data into text data, means for analyzing the converted text data and detecting inappropriate speech and behavior, means for generating an alert message based on the detection results and notifying the terminal of the alert message, and means for displaying the notified alert message on a user interface. In particular, by analyzing the text data using a generative AI model and detecting inappropriate speech and behavior, the system can recognize characteristic patterns of maternity harassment with high accuracy. Furthermore, by collecting conversational audio in real time, the system provides a mechanism for immediate response.

[0007] ---

[0008] "Voice data" refers to acoustic information that records a user's speech, and is used for processing such as voice recognition.

[0009] "Collection means" refers to hardware and software components for acquiring audio data, including, specifically, a microphone and an audio collection application.

[0010] A "server" refers to a computer system that receives and processes data from multiple terminals via a network, and is responsible for analyzing and storing voice data.

[0011] "Means of transmission" refers to the communication protocol and equipment used to send the collected audio data to the server, such as using WebSocket or HTTP / 2.

[0012] "Text data" refers to human-readable text information generated by voice recognition technology based on voice data.

[0013] "Means of conversion" refers to speech recognition tools or software for converting voice data into text data, including speech recognition APIs.

[0014] "Means for analysis" refers to the generative AI models and analytical algorithms used to examine text data and identify inappropriate behavior.

[0015] "Means of detection" refers to a processing method for finding inappropriate behavior or statements that may constitute maternity harassment from the analyzed text data.

[0016] An "alert message" refers to a warning message that notifies a user or an administrator when inappropriate behavior is detected.

[0017] "Means of notification" refers to the method or protocol for sending an alert message to the relevant device and notifying the user, such as email or push notification.

[0018] The term "user interface" refers to a display device and operating means for transferring information between a user and a system, and includes screen displays and audio notifications. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] ---

[0041] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, in the workplace at an early stage. This system consists of the following main components for collecting and analyzing voice data and issuing alerts:

[0042] System configuration

[0043] 1. Audio collection method

[0044] The device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[0045] 2. Means of transmitting audio data

[0046] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is transmitted in a secure manner, ensuring privacy.

[0047] 3. Speech-to-text methods

[0048] The server converts the received voice data into text data using a voice recognition API, which uses a cloud service capable of highly accurate recognition (for example, a voice recognition cloud service).

[0049] 4. Methods for analyzing text data

[0050] The server then inputs the converted text data into a generative AI model (such as ChatGPT) and analyzes the text for inappropriate behavior by comparing it with a database of past cases of maternity harassment.

[0051] 5. Measures to detect inappropriate behavior

[0052] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior contains any inappropriate behavior. If the score exceeds a certain threshold, the speech or behavior is deemed to be highly likely to constitute maternity harassment.

[0053] 6. How to Generate an Alert Message

[0054] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The alert message is then sent appropriately to the affected user or administrator.

[0055] 7. Alert notification method

[0056] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[0057] Specific examples

[0058] Conversation scenario

[0059] 1. Collecting comments

[0060] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0061] The device collects this speech through a microphone and transmits the data to a server.

[0062] 2. Speech-to-text conversion and analysis

[0063] The server uses a speech recognition API to convert the collected voice data into text.

[0064] The converted text is analyzed by a generative AI model and scored as a statement that may constitute maternity harassment.

[0065] 3. Alert generation and notification

[0066] Based on the text whose score exceeds the threshold, the server generates an alert message saying, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[0067] The terminal receives this alert message and notifies the user.

[0068] 4. User Behavior

[0069] The user checks the alert message and realizes that his or her remarks were inappropriate.

[0070] As a result, users will reconsider what they said and be more considerate in future conversations.

[0071] ---

[0072] In this way, the system can detect maternity harassment in the workplace at an early stage and prevent inappropriate behavior before it occurs. It is expected that this system will provide a safe and comfortable working environment.

[0073] The processing flow will be explained below.

[0074] ---

[0075] Step 1:

[0076] Audio collection

[0077] The device uses a microphone to collect the user's conversational voice in real time. The voice data is recorded in a specified format (e.g., AAC or Opus). This collection continues from the moment the conversation starts until it stops.

[0078] Step 2:

[0079] Audio data compression and buffering

[0080] The device converts the collected audio data into a compressed format for efficient transmission to the server. The compressed audio data is divided into small data chunks and stored in a buffer for transmission.

[0081] Step 3:

[0082] Sending audio data

[0083] The device transmits the audio data stored in the buffer to the server at regular intervals. This transmission uses a low-latency communication protocol (e.g., WebSocket or HTTP / 2). The transmission is performed in real time and is monitored to ensure there is no interruption of data.

[0084] Step 4:

[0085] Receiving audio data

[0086] The server receives voice data sent from the terminal in real time, stores the received voice data in a fixed buffer, and prepares it for text conversion.

[0087] Step 5:

[0088] Decoding audio data

[0089] The server decodes the received compressed audio data and returns it to its original format. The decoded audio data is stored in temporary storage and passed on to the next processing step.

[0090] Step 6:

[0091] Speech recognition to text

[0092] The server sends the voice data stored in the temporary storage to a voice recognition API, which converts it into text data. The voice recognition API provides highly accurate voice recognition, and the converted text is stored in a database.

[0093] Step 7:

[0094] Preprocessing text data

[0095] The server cleanses the text data obtained from the speech recognition API to remove mistranslations and noise. This preprocessing contributes to improving the accuracy of the text data.

[0096] Step 8:

[0097] Analysis using generative AI models

[0098] The server inputs the preprocessed text data into a generative AI model, which analyzes the text data for inappropriate behavior. The generative AI model then detects inappropriate behavior with high accuracy based on a large amount of training data.

[0099] Step 9:

[0100] Detecting inappropriate behavior

[0101] Based on the analysis results of the generative AI model, the server determines whether the statements contained in the text data constitute inappropriate behavior. The determination is made using a scoring system, and statements that exceed a certain threshold are deemed inappropriate.

[0102] Step 10:

[0103] Generate an alert message

[0104] When inappropriate behavior is detected, the server generates an alert message containing the specific content of the behavior and a warning message. This alert message is sent to the perpetrator and the administrator.

[0105] Step 11:

[0106] Sending an alert message

[0107] The server then sends the generated alert message to the target device. The alert message is distributed in real time and, if necessary, notifies the administrator or harassment prevention department.

[0108] Step 12:

[0109] Receiving and Viewing Alerts

[0110] The terminal receives the alert message sent from the server and displays it on the user interface. The alert message is notified visually or audibly so that the user can immediately recognize it.

[0111] Step 13:

[0112] User Behavior

[0113] The user will check the alert message displayed on their device and realize that their remarks were inappropriate. Based on this realization, the user will be mindful of their behavior and speech in future conversations. It is also expected that the user will receive additional education and training as necessary.

[0114] ---

[0115] Through the above processing steps, the present invention can detect and prevent inappropriate behavior related to pregnancy, childbirth, and childcare in the workplace at an early stage, which is expected to provide a safe and comfortable working environment.

[0116] Example 1

[0117] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0118] In today's workplace, inappropriate verbal and physical behavior related to pregnancy, childbirth, and childcare (known as maternity harassment) is becoming more common, causing increased mental and physical strain on employees. In particular, unintentional comments made by managers and colleagues can cause stress to pregnant women and worsen the workplace environment. There is a need for methods to detect and prevent such inappropriate behavior early on.

[0119] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0120] In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data using a generative AI model to detect inappropriate speech and behavior, means for generating an alert message based on the detection results, means for notifying a terminal of the generated alert message, and means for displaying the notified alert message on a user interface. This enables early detection and prevention of inappropriate speech and behavior related to pregnancy, childbirth, and childcare in the workplace.

[0121] "Audio data" refers to a digital signal of sound captured using an acoustic device such as a microphone.

[0122] A "server" is a computer system for transmitting, receiving, and analyzing voice data.

[0123] "Text data" is character information converted from voice data using voice recognition technology.

[0124] A "generative AI model" is an artificial intelligence model that is trained to perform specific tasks based on input data.

[0125] "Analysis" is the process of determining whether text data contains inappropriate language or behavior.

[0126] "Inappropriate words and actions" are comments or actions that place mental or physical burden on employees in the workplace in relation to pregnancy, childbirth, or childcare.

[0127] An "alert message" is a message that alerts the user when inappropriate speech or behavior is detected.

[0128] A "terminal" is an information processing device such as a smartphone or PC used by a user.

[0129] The "user interface" refers to the screen and audio devices that allow the user to visually and audibly confirm the information displayed on the terminal.

[0130] A "microphone" is a device for converting sound into an electrical signal.

[0131] MODE FOR CARRYING OUT THE INVENTION

[0132] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare (so-called maternity harassment) in the workplace at an early stage. This system is composed of the following main elements for collecting and analyzing voice data and sending alerts.

[0133] System configuration

[0134] 1. Audio collection method

[0135] A device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time. The collected audio data is recorded uncompressed or in an appropriate compressed format (e.g., AAC or MP3) and then transmitted to a server.

[0136] 2. Means of transmitting audio data

[0137] The device sends the collected voice data to the server using a secure communication protocol (e.g., HTTPS or WebSocket). The voice data is encrypted and sent to the server in a privacy-protected manner.

[0138] 3. Speech-to-text methods

[0139] The server converts the transmitted voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The server sends the audio file to the speech recognition API and receives the returned text data.

[0140] 4. Methods for analyzing text data

[0141] The server inputs the converted text data into a generative AI model (e.g., OpenAI's ChatGPT) for analysis. The following prompt is used during analysis:

[0142] Analyze whether the following text is an inappropriate statement about pregnancy or morning sickness, and if so, output a score along with the reason.

[0143] Text: "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0144] The generative AI model analyzes the given text and returns a score and reason for detecting inappropriate behavior.

[0145] 5. Measures to detect inappropriate behavior

[0146] The server performs a scoring process based on the analysis results returned by the generative AI model. If the score exceeds a certain threshold, the remark is deemed to constitute inappropriate behavior (maternity harassment).

[0147] 6. How to Generate an Alert Message

[0148] If inappropriate behavior is detected, the server generates an alert message containing the specific content of the comment and a warning, such as "This comment contains comments about pregnancy and morning sickness. Please be considerate in your comments and behavior."

[0149] 7. Alert notification method

[0150] The terminal receives alert messages sent from the server in real time and displays them on the user interface. Alert messages are notified visually (banners and pop-up notifications) and audibly (voice alerts and chimes) so that users can quickly acknowledge them.

[0151] Specific examples

[0152] Conversation scenario

[0153] 1. Collecting comments

[0154] The user (boss) says, "Mr. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0155] The terminal collects this speech via a microphone and transmits the audio data to a server.

[0156] 2. Speech-to-text conversion and analysis

[0157] The server uses a speech recognition API to convert the collected voice data into text.

[0158] The converted text is parsed using prompts from a generative AI model.

[0159] 3. Detecting inappropriate behavior and generating alerts

[0160] The server generates an alert message containing specific statements and warnings based on text whose score exceeds the threshold.

[0161] The terminal receives this alert message and notifies the user.

[0162] In this way, the system can detect maternity harassment in the workplace early and prevent inappropriate behavior by immediately notifying users, which is expected to provide a safe and comfortable working environment.

[0163] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0164] Program processing steps and detailed explanations

[0165] Step 1: Collecting audio

[0166] The device (user's smartphone or PC) uses a microphone to collect the user's speech in real time. For example, the user might say, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress." This collected speech data (input) is recorded (output) either uncompressed or in an appropriate compressed format (e.g., AAC or MP3).

[0167] Step 2: Sending audio data

[0168] The device sends the collected voice data to the server using a secure communication protocol (e.g., HTTPS or WebSocket). The voice data (input) is encrypted and transmitted to protect the privacy of the data. The server receives this voice data (output).

[0169] Step 3: Speech to Text

[0170] The server sends the received voice data (input) to a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert it into text data. The speech recognition API analyzes the voice data and returns text data (output). The server receives this converted text data.

[0171] Step 4: Analyzing the text data

[0172] The server inputs the converted text data (input) into a generative AI model (e.g., OpenAI's ChatGPT) for analysis. The server sends the following prompt to the generative AI model:

[0173] Analyze whether the following text is an inappropriate statement about pregnancy or morning sickness, and if so, output a score along with the reason.

[0174] Text: "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0175] The generative AI model analyzes this prompt and returns a score indicating inappropriate behavior and the reason (output) to the server.

[0176] Step 5: Detect inappropriate behavior

[0177] The server performs a score based on the analysis results (input) returned from the generative AI model. If this score exceeds a certain threshold, the comment is deemed to be inappropriate behavior (maternity harassment). Specifically, for example, a score of "85 / 100" is output, and a reason such as "Reason: Because mentioning pregnancy or morning sickness involves personal privacy" is displayed.

[0178] Step 6: Generate an alert message

[0179] If the server detects inappropriate behavior, it generates an alert message (input) that includes the specific content of the comment and a warning. For example, it could generate a message (output) stating, "This comment contains comments about pregnancy and morning sickness. Please be considerate in your speech and behavior."

[0180] Step 7: Alert Notification

[0181] The terminal receives the alert message (input) sent from the server in real time and displays it on the user interface. The terminal notifies the user of the alert message visually (banner or pop-up notification) and audibly (voice alert or chime) so that the user can quickly acknowledge it (output). The user acknowledges the alert message and realizes that their comment was inappropriate.

[0182] (Application example 1)

[0183] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0184] In today's workplaces, inappropriate verbal and physical behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, remains a problem. If inappropriate comments and behavior are left unchecked, it can lead to a worsening work environment and have serious consequences for the mental and physical health of victims. However, conventional methods currently make it difficult to quickly detect and prevent such inappropriate behavior. Therefore, a system is needed that can collect and analyze voice data in real time and quickly detect and notify employees of inappropriate behavior in the workplace.

[0185] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0186] In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data using a generative AI model to detect inappropriate speech and behavior, means for generating an alert message based on the detection result and notifying the terminal of the alert message, means for notifying the alert message in real time via WebSocket, and means for displaying the notified alert message on a user interface. This makes it possible to quickly detect inappropriate speech and behavior in the workplace and respond in real time.

[0187] "Voice data" is data that represents voice in digital form, and is information that is collected in real time and is subject to analysis.

[0188] A "collection means" is a system or process for acquiring audio data using a device such as a microphone.

[0189] A "server" is a remote computer system that receives voice data, analyzes it, generates alerts, and performs other processing.

[0190] The "transmission means" is a mechanism for temporarily storing collected voice data and transferring it to a server for subsequent processing.

[0191] "Text data" is character string information converted from voice data using voice recognition technology.

[0192] "Conversion means" refers to technology or systems for converting voice data into text data, including voice recognition APIs.

[0193] "Analysis Methods" refers to the processes and techniques used to identify inappropriate behavior within text data using generative AI models.

[0194] A "generative AI model" is a machine learning model used to analyze text data and is an algorithm for detecting inappropriate behavior.

[0195] "Inappropriate behavior" refers to inappropriate remarks or actions related to pregnancy, childbirth, or childcare in the workplace.

[0196] An "alert message" is a notification message that alerts the user when inappropriate behavior is detected.

[0197] "Notification means" refers to a system or method for transmitting the generated alert message to a user's terminal in real time.

[0198] "User interface" refers to a screen or application that displays the notified alert message so that the user can check it.

[0199] "WebSocket" is a communication protocol for real-time communication between a server and a client.

[0200] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, in the workplace at an early stage. This system consists of the following main components for real-time collection and analysis of voice data and notification of alerts:

[0201] System configuration

[0202] 1. Audio collection method

[0203] This system uses a device (e.g., a smartphone, PC, or dedicated microphone device) to collect the user's speech in real time. The voice data collected through the microphone is recorded in uncompressed or compressed format and then sent to a server.

[0204] 2. Means of transmitting audio data

[0205] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is securely transferred using TLS, ensuring privacy.

[0206] 3. Speech-to-text methods

[0207] The server converts the received voice data into text data in real time using a speech recognition API (for example, Google Cloud Speech-to-Text). Using cloud services enables highly accurate speech recognition.

[0208] 4. Methods for analyzing text data

[0209] The server inputs the converted text data into a generative AI model (e.g., ChatGPT) and analyzes the text for inappropriate behavior. This analysis involves comparing the data with a database of past cases of maternity harassment. The generative AI model's analysis also uses prompt sentences to help detect specific behavior.

[0210] 5. Measures to detect inappropriate behavior

[0211] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior contains any inappropriate behavior. If the score exceeds a certain threshold, the speech or behavior is deemed to be highly likely to constitute maternity harassment.

[0212] 6. How to Generate an Alert Message

[0213] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The alert message is created based on the results of detailed analysis.

[0214] 7. Alert notification method

[0215] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified visually and audibly so that the user can quickly recognize it.

[0216] Specific examples

[0217] Conversation scenario

[0218] Collecting statements:

[0219] The user (boss) says, "Ms. B, I've been taking a lot of time off recently due to morning sickness, but please make sure that it doesn't affect my work progress." Such comments are collected by microphones in the workplace.

[0220] Speech-to-text and analysis:

[0221] The server uses a speech recognition API to convert the collected voice data into text, for example, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0222] The converted text is analyzed by a generative AI model, which scores the statement as inappropriate.

[0223] Alert generation and notification:

[0224] Based on the text whose score exceeds the threshold, the server generates an alert message stating, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[0225] The device receives this alert message via WebSocket and notifies the user in real time.

[0226] Example prompt sentence:

[0227] "The manager said, 'I have kids so it's a problem if you're always late.' Does this statement constitute inappropriate behavior?"

[0228] This system is expected to not only quickly detect inappropriate behavior in the workplace, but also enable real-time responses, thereby providing a safe and comfortable working environment.

[0229] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0230] Step 1:

[0231] The device uses a microphone in the workplace to collect voice data in real time. This allows the user's statements and conversation content to be captured in digital form. The input is voice, and the output is digital voice data. For example, if a user says, "Mr. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure it doesn't affect my work progress," the device will capture this as voice data.

[0232] Step 2:

[0233] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The input is digital audio data, and the output is the audio data sent to the server. Communication is securely encrypted using TLS.

[0234] Step 3:

[0235] The server receives the transmitted voice data and converts it into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is the voice data received by the server, and the output is the converted text data. Specifically, voice data such as "Ms. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure that it doesn't affect my work progress" is converted into text data such as "Ms. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure that it doesn't affect my work progress."

[0236] Step 4:

[0237] The server inputs the converted text data into a generative AI model (e.g., ChatGPT) and analyzes the inappropriate behavior in the text. A prompt is used for the analysis. The input is the converted text data, and the output is the analysis result. Specifically, the prompt, "Does this statement constitute inappropriate behavior?", is input into the generative AI model, and the analysis result is obtained.

[0238] Step 5:

[0239] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior is inappropriate. The input is the analysis result of the generative AI model, and the output is a score. If the score exceeds a certain threshold, the speech or behavior is deemed inappropriate.

[0240] Step 6:

[0241] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The input is the analysis result where the score exceeds the threshold, and the output is the alert message. Specifically, the alert message generated reads, "The remarks include comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[0242] Step 7:

[0243] The terminal receives alert messages sent from the server in real time via WebSocket and displays them on the user interface. The input is the alert message, and the output is the notification displayed on the user interface. Specifically, the user will be notified visually and / or audibly.

[0244] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0245] ---

[0246] This invention is a system that detects inappropriate behavior and emotions related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevents them from occurring. This system consists of the following main components for collecting and analyzing voice data, recognizing emotions, and issuing alerts.

[0247] System configuration

[0248] 1. Audio collection method

[0249] The device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[0250] 2. Means of transmitting audio data

[0251] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is transmitted in a secure manner, ensuring privacy.

[0252] 3. Speech-to-text methods

[0253] The server converts the received voice data into text data using a voice recognition API, which uses a cloud service capable of highly accurate recognition (for example, a voice recognition cloud service).

[0254] 4. Text data and sentiment analysis methods

[0255] The server inputs the converted text and voice data into a generative AI model and emotion engine, analyzes inappropriate behavior in the text, and recognizes the user's emotional state. The analysis is performed by comparing it with a database of past cases of maternity harassment.

[0256] 5. Measures to detect inappropriate behavior and emotions

[0257] The server calculates a score based on the results of analysis by the generative AI model and emotion engine to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed to be highly likely to constitute maternity harassment, and the user's emotional state is also used in the next process.

[0258] 6. How to Generate an Alert Message

[0259] When maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. This alert message will be used to appropriately notify the perpetrator and the administrator. The content of the alert message will also be adjusted based on the user's emotional state.

[0260] 7. Alert notification method

[0261] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[0262] Specific examples

[0263] Conversation scenario

[0264] 1. Collecting comments

[0265] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0266] The device collects this speech through a microphone and transmits the data to a server.

[0267] 2. Speech-to-text conversion and analysis

[0268] The server uses a speech recognition API to convert the collected voice data into text.

[0269] The converted text and speech data are analyzed by a generative AI model and emotion engine to score inappropriate language in the text and the user's emotional state.

[0270] 3. Alert generation and notification

[0271] Based on the text whose score exceeds the threshold, the server generates an alert message saying, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[0272] Depending on the user's emotional state, for example, if the user is emotionally charged, a supplemental message such as "Let's stay calm and continue the conversation" may be added.

[0273] The terminal receives this alert message and notifies the user.

[0274] 4. User Behavior

[0275] The user will receive an alert message, realizing that their remarks were inappropriate, and will also receive feedback about their emotional state.

[0276] As a result, users will reconsider what they have said and be more considerate in future conversations, while also paying attention to controlling their emotions.

[0277] ---

[0278] In this way, the system can detect maternity harassment and related emotional upheavals in the workplace at an early stage and prevent inappropriate behavior before it occurs, which is expected to provide a safe and comfortable working environment.

[0279] The processing flow will be explained below.

[0280] ---

[0281] Step 1:

[0282] Audio collection

[0283] The device uses a microphone to capture the user's conversation in real time, and the audio data is recorded in uncompressed or compressed format and then transmitted to a server.

[0284] Step 2:

[0285] Audio data compression and buffering

[0286] The device converts the collected audio data into a compressed format for efficient transmission to the server. The compressed audio data is divided into small data chunks and stored in a buffer for transmission.

[0287] Step 3:

[0288] Sending audio data

[0289] The device transmits the audio data stored in the buffer to the server at regular intervals. This transmission uses a low-latency communication protocol (e.g., WebSocket or HTTP / 2). The transmission is performed in real time and is monitored to ensure there is no interruption of data.

[0290] Step 4:

[0291] Receiving and decoding audio data

[0292] The server receives the audio data sent from the device in real time, decodes the compressed audio data, and stores it in temporary storage before passing it on to the next processing step.

[0293] Step 5:

[0294] Speech recognition to text

[0295] The server sends the voice data stored in the temporary storage to a voice recognition API, which converts it into text data. The voice recognition API provides highly accurate voice recognition, and the converted text data is stored in a database.

[0296] Step 6:

[0297] Preprocessing text data

[0298] The server cleanses the text data received from the speech recognition API, removing mistranslations and noise, and prepares the cleansed text data for analysis.

[0299] Step 7:

[0300] Analysis using generative AI models and emotion engines

[0301] The server inputs the cleansed text data into a generative AI model and an emotion engine, which analyzes the text for inappropriate behavior and the user's emotional state. The generative AI model detects inappropriate behavior with high accuracy based on a large amount of training data, and the emotion engine estimates the user's emotional state from the voice data.

[0302] Step 8:

[0303] Detecting inappropriate behavior and sentiment

[0304] Based on the analysis results, the server determines whether the statements contained in the text data constitute inappropriate behavior and simultaneously scores the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed inappropriate and the user's emotional state is also recorded.

[0305] Step 9:

[0306] Generate an alert message

[0307] When inappropriate behavior is detected, the server generates an alert message containing specific statements and a warning, and adjusts the content and tone of the alert message based on the user's emotional state and includes additional feedback as needed.

[0308] Step 10:

[0309] Sending an alert message

[0310] The server then sends the generated alert message to the target device. The alert message is distributed in real time and, if necessary, notifies the administrator or harassment prevention department.

[0311] Step 11:

[0312] Receiving and Viewing Alerts

[0313] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[0314] Step 12:

[0315] User Behavior

[0316] The user checks the alert message displayed on their device and realizes that their remarks were inappropriate. They also receive feedback on their emotional state. Based on this recognition, the user can be more considerate in future conversations and pay more attention to controlling their emotions.

[0317] ---

[0318] Through this step, the system of the present invention can detect inappropriate behavior and heightened emotions related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevent inappropriate behavior before it occurs, which is expected to provide a safe and comfortable working environment.

[0319] Example 2

[0320] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0321] In today's work environment, inappropriate behavior and attitudes related to pregnancy, childbirth, and childcare remain a problem. These issues can particularly increase stress and psychological burden on female employees in the workplace. Furthermore, managers often lack the means to identify and address these issues early before they become serious. As a result, maintaining a safe and comfortable work environment is difficult. This invention aims to improve the work environment and reduce the psychological burden on employees by detecting inappropriate behavior and the user's emotional state in the workplace early and responding promptly.

[0322] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0323] In this invention, the server includes a means for converting voice data into text data, a means for analyzing the converted text data and voice data to recognize the user's emotional state, and a means for determining inappropriate behavior and the user's emotional state based on the analysis results, thereby enabling prompt and accurate identification and response to inappropriate behavior related to pregnancy, childbirth, and childcare in the workplace.

[0324] "Audio data" means digital audio signals collected through an audio input device.

[0325] "Text data" refers to text information converted from voice data using a voice recognition API.

[0326] A "generative AI model" is an algorithm or machine learning model that uses artificial intelligence to analyze text data and detect inappropriate behavior.

[0327] An "emotion engine" is software for recognizing a user's emotional state from text data and voice tone.

[0328] An "alert message" is a notification message that is generated to encourage the user to behave considerately when inappropriate behavior or the user's emotional state is determined.

[0329] "Audio input device" refers to a microphone or other audio collection device for capturing speech in real time.

[0330] A "secure transfer protocol" is a communication technology for encrypting and transmitting data securely, and specifically refers to protocols such as HTTPS and WebSocket Secure.

[0331] A "prompt sentence" is the input text presented to a generative AI model, and is the sentence that serves as the starting point for analysis and generation.

[0332] This invention is a system that detects inappropriate behavior and the user's emotional state related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevents such behavior before it occurs. This system consists of the following main elements for collecting and analyzing voice data, recognizing emotions, and issuing alerts.

[0333] System configuration

[0334] Audio collection method

[0335] A device (e.g., a smartphone or a personal computer) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and transmitted to a server.

[0336] Voice data transmission method

[0337] The device sends the collected voice data to the server using an appropriate communication protocol (e.g., WebSocket Secure or HTTPS). The voice data is transmitted securely, ensuring privacy.

[0338] Voice-to-text conversion methods

[0339] The server converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text).

[0340] Text data and sentiment analysis methods

[0341] The server inputs the converted text and voice data into a generative AI model (e.g., GPT-4) and an emotion engine to analyze inappropriate behavior in the text and recognize the user's emotional state. The analysis is performed by comparing the data with a database of past cases of maternity harassment.

[0342] Measures to detect inappropriate behavior and emotions

[0343] The server calculates a score based on the results of analysis by the generative AI model and emotion engine to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed to be highly likely to constitute maternity harassment, and the user's emotional state is also used in the next process.

[0344] How to generate an alert message

[0345] When maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. This alert message will be used to appropriately notify the perpetrator and the administrator. The content of the alert message will also be adjusted based on the user's emotional state.

[0346] Alert notification method

[0347] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[0348] Specific examples

[0349] Conversation scenario

[0350] 1. Collecting comments

[0351] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0352] The device collects this speech through a microphone and transmits the data to a server.

[0353] 2. Speech-to-text conversion and analysis

[0354] The server uses a speech recognition API to convert the collected voice data into text.

[0355] The converted text and voice data are analyzed by a generative AI model (GPT-4) and an emotion engine to score inappropriate language in the text and the user's emotional state.

[0356] 3. Alert generation and notification

[0357] Based on the text whose score exceeds the threshold, the server generates an alert message stating, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[0358] Depending on the user's emotional state, for example, if the user is emotionally charged, a supplemental message such as "Let's stay calm and continue the conversation" is added.

[0359] The terminal receives this alert message and notifies the user.

[0360] 4. User Behavior

[0361] The user will receive an alert message, realizing that their remarks were inappropriate, and will also receive feedback about their emotional state.

[0362] As a result, users will reconsider what they have said and be more considerate in future conversations, while also paying attention to controlling their emotions.

[0363] Prompt Sentence Examples

[0364] "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0365] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0366] Step 1:

[0367] Audio collection

[0368] Input: User's speech

[0369] Specific operation: A user is having a conversation at work. The device (e.g., a smartphone or a personal computer) activates a built-in or external microphone to collect the conversation audio in real time.

[0370] Output: Audio data (digital format)

[0371] Step 2:

[0372] Sending audio data

[0373] Input: Audio data

[0374] Specific operation: The device encrypts the collected voice data using SSL / TLS and sends it to the server's API endpoint using a secure communication protocol (e.g., HTTPS or WebSocket Secure).

[0375] Output: Audio data sent to the server

[0376] Step 3:

[0377] Speech to text

[0378] Input: Audio data sent to the server

[0379] Specific operation: The server passes the received voice data to a speech recognition API (e.g., Google Cloud Speech-to-Text) and converts it into text data. The server receives the text data returned by the API.

[0380] Output: Converted text data

[0381] Step 4:

[0382] Text data and sentiment analysis

[0383] Input: Converted text and audio data

[0384] Specific operation: The server analyzes the text data entered as prompts into a generative AI model (e.g., GPT-4) and identifies inappropriate behavior within the text. In addition, it uses an emotion engine to recognize the user's emotional state from the tone of voice and text data. The analysis results are compared with a database of past maternity harassment cases.

[0385] Output: Analysis results and user emotional state data

[0386] Step 5:

[0387] Detecting inappropriate behavior and sentiment

[0388] Input: Analysis results and user emotional state data

[0389] Specific operation: The server calculates a score for inappropriate behavior and the user's emotional state based on the analysis results from the generative AI model and emotion engine. The score is compared with a pre-set threshold, and behavior that exceeds the threshold is deemed inappropriate.

[0390] Output: Judgment result of inappropriate behavior and emotional state

[0391] Step 6:

[0392] Generate an alert message

[0393] Input: Inappropriate behavior and emotional state assessment results

[0394] Specific actions: Based on the judgment result, the server generates an alert message containing the specific content of the comment and a warning. For example, it creates a message such as, "The comment contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." Based on the user's emotional state, it also adds a supplementary message such as, "Please remain calm and continue the conversation."

[0395] Output: The generated alert message

[0396] Step 7:

[0397] Alert Notification

[0398] Input: The generated alert message

[0399] Specific operation: The terminal receives the alert message sent from the server in real time and displays it on the user interface, specifically notifying the user by means of a pop-up notification, a sound notification, or the like.

[0400] Output: Alert message sent to the user

[0401] Specific examples

[0402] 1. User statement: The user (boss) says, "Ms. B, I've been taking a lot of time off recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0403] 2. What the device does: The built-in microphone collects this speech, encrypts the data, and sends it to a server.

[0404] 3. Server operation: The speech is converted into text using a speech recognition API, and then analyzed using a generative AI model (GPT-4) and an emotion engine.

[0405] 4. Analysis results: Inappropriate behavior is detected and the user is determined to be emotionally charged.

[0406] 5. Alert generation: An alert is generated stating, "This post contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." along with a supplementary message saying, "Please remain calm and continue the discussion."

[0407] 6. Device notification: Alert messages are notified to the user via pop-up and sound.

[0408] 7. User action: The user checks the notification, reviews their comments, and controls their emotions.

[0409] (Application example 2)

[0410] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0411] Inappropriate verbal and physical behavior and the emotional state of employees related to pregnancy, childbirth, and childcare often become problems in the workplace. However, it is difficult to detect these inappropriate verbal and physical behavior and emotional outbursts early and prevent them from occurring. In particular, achieving this in real time has been difficult with conventional technology. The objective of this invention is to provide a system that quickly detects inappropriate verbal and physical behavior and emotional outbursts in the workplace and provides a safe and comfortable working environment.

[0412] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data and detecting inappropriate speech and behavior and the user's emotional state, means for generating an alert message based on the detection result and the emotional state and notifying the terminal of the alert message, and means for displaying the notified alert message on a user interface. This makes it possible to detect inappropriate speech and behavior and heightened emotions in the workplace in real time and prevent them from occurring.

[0413] "Audio data" refers to electronically recorded data of a user's conversation or environmental sounds.

[0414] "Collection means" refers to devices and software for collecting voice data in real time.

[0415] "Transmission means" refers to the communication method or protocol used to send collected voice data to the server.

[0416] "Text data" refers to data obtained by converting voice data into text information using voice recognition technology.

[0417] "Conversion means" refers to the technology or device used to convert voice data into text data.

[0418] "Analysis means" refers to technologies and algorithms used to analyze the converted text data and detect inappropriate behavior or emotional states.

[0419] A "generative AI model" is an AI model trained using machine learning or deep learning, and is used to analyze text data.

[0420] "User's emotional state" refers to the emotional state that the user expresses through their words and actions, such as joy, anger, or anxiety.

[0421] An "alert message" is a warning message generated based on detected inappropriate behavior or emotional state.

[0422] "Notification means" refers to the communication method or protocol used to communicate the generated alert message to the user or administrator.

[0423] "User interface" refers to an interface for visually or audibly displaying an alert message to a user.

[0424] A "prompt sentence" is an initial sentence or instruction input to a generative AI model, and is used to indicate the direction of analysis.

[0425] This invention is a system that detects and prevents inappropriate behavior and emotional upset related to pregnancy, childbirth, and childcare in the workplace at an early stage. This system is mainly composed of components for collecting, transmitting, converting, and analyzing voice data, detecting inappropriate behavior, and generating and notifying alert messages.

[0426] System configuration

[0427] 1. Audio collection method

[0428] The device (e.g., a smartphone or PC) uses a microphone to collect the user's conversational voice in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[0429] 2. Means of transmitting audio data

[0430] The device sends the collected audio data to the server using an appropriate communication protocol (such as HTTP / 2). The audio data is transmitted securely to protect privacy.

[0431] 3. Speech-to-text methods

[0432] The server converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text), a cloud service that provides highly accurate recognition.

[0433] 4. Text data and sentiment analysis methods

[0434] The server then inputs the converted text and audio data into a generative AI model and a sentiment analysis engine (e.g., TextBlob) to analyze the text for inappropriate language and the user's emotional state. The analysis is based on the prompt sentence and identifies inappropriate phrases and heightened emotions.

[0435] As an example, use the following prompt:

[0436] "Turn this audio data into text and look for inappropriate language and heightened emotions. Pay attention to specific phrases and pick out negative comments."

[0437] 5. Measures to detect inappropriate behavior and emotions

[0438] The server uses the generative AI model and sentiment analysis engine to calculate a score from the analyzed text data to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed inappropriate.

[0439] 6. How to Generate an Alert Message

[0440] The server generates an alert message when inappropriate behavior and emotional upset are detected, which includes specific statements and reminders and is tailored to the user's emotional state.

[0441] 7. Alert notification method

[0442] The terminal receives alert messages sent from the server in real time and displays them on the user interface. The alert messages are notified to the user visually and audibly so that they can be quickly recognized.

[0443] Specific examples

[0444] 1. Collecting comments

[0445] The terminal collects the user's (boss') remarks through a microphone.

[0446] For example, a statement such as, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress" is collected.

[0447] 2. Speech-to-text conversion and analysis

[0448] The server uses a speech recognition API to convert the collected voice data into text.

[0449] The converted text is analyzed using a generative AI model and a sentiment analysis engine to score inappropriate behavior and emotional states.

[0450] 3. Alert generation and notification

[0451] The server detects inappropriate comments and generates an alert message saying, "This comment contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." If emotions are running high, a supplemental message is added saying, "Please remain calm and continue the conversation."

[0452] The terminal receives this alert message in real time and notifies the user.

[0453] 4. User Behavior

[0454] The user sees the alert message, realizes that what they said was inappropriate, and also receives feedback about their emotional state.

[0455] This will encourage users to review their future comments and be more considerate in their speech and actions.

[0456] This system will enable early detection of inappropriate behavior and emotional outbursts in the workplace, providing a safe and comfortable working environment.

[0457] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0458] ---

[0459] Step 1:

[0460] The device uses a microphone to collect conversational audio in real time. The input is the user's live voice, and the output is stored on the device as audio data. This audio data can be recorded in uncompressed or compressed format.

[0461] Step 2:

[0462] The device sends the collected audio data to the server using a secure communication protocol such as HTTP / 2. The input is audio data, and the output is audio data sent to the server. The communication is encrypted to protect privacy.

[0463] Step 3:

[0464] The server converts the received voice data into text data using a speech recognition API (Google Cloud Speech-to-Text). The input here is voice data, and the output is text data. This text data is obtained using speech recognition technology.

[0465] Step 4:

[0466] The server inputs the converted text data into a generative AI model and an emotion analysis engine (TextBlob) to analyze the inappropriate behavior and emotional state of the user in the text. The input here is text data, and the output is a score for the inappropriate behavior and emotional state. The generative AI model performs the analysis using prompts. For example, a prompt might be, "Convert this audio data into text and detect inappropriate behavior and heightened emotions. Pay attention to specific phrases and pick out negative comments."

[0467] Step 5:

[0468] The server uses the generative AI model and emotion analysis engine to analyze the results and determine whether the behavior and emotional state are inappropriate. If the score exceeds a certain threshold, the behavior is deemed inappropriate. The input here is the analysis score, and the output is the judgment of inappropriate behavior.

[0469] Step 6:

[0470] The server generates an alert message when inappropriate behavior or heightened emotions are detected. The input is the judgment result of inappropriate behavior, and the output is an alert message. The alert message includes the specific content of the statement and a warning, and if necessary, a supplemental message based on the emotional state is added.

[0471] Step 7:

[0472] The terminal receives alert messages sent from the server in real time and displays them on the user interface. The input here is the alert message, and the output is a notification to the user. The alert message is notified visually and audibly so that the user can quickly recognize it.

[0473] ---

[0474] The above processing steps make it possible to detect inappropriate behavior and emotional outbursts in the workplace in real time and prevent them from occurring.

[0475] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0476] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0477] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0478] [Second embodiment]

[0479] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0480] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0481] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0482] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0483] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0484] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0485] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0486] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0487] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0488] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0489] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0490] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0491] ---

[0492] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, in the workplace at an early stage. This system consists of the following main components for collecting and analyzing voice data and issuing alerts:

[0493] System configuration

[0494] 1. Audio collection method

[0495] The device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[0496] 2. Means of transmitting audio data

[0497] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is transmitted in a secure manner, ensuring privacy.

[0498] 3. Speech-to-text methods

[0499] The server converts the received voice data into text data using a voice recognition API, which uses a cloud service capable of highly accurate recognition (for example, a voice recognition cloud service).

[0500] 4. Methods for analyzing text data

[0501] The server then inputs the converted text data into a generative AI model (such as ChatGPT) and analyzes the text for inappropriate behavior by comparing it with a database of past cases of maternity harassment.

[0502] 5. Measures to detect inappropriate behavior

[0503] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior contains any inappropriate behavior. If the score exceeds a certain threshold, the speech or behavior is deemed to be highly likely to constitute maternity harassment.

[0504] 6. How to Generate an Alert Message

[0505] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The alert message is then sent appropriately to the affected user or administrator.

[0506] 7. Alert notification method

[0507] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[0508] Specific examples

[0509] Conversation scenario

[0510] 1. Collecting comments

[0511] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0512] The device collects this speech through a microphone and transmits the data to a server.

[0513] 2. Speech-to-text conversion and analysis

[0514] The server uses a speech recognition API to convert the collected voice data into text.

[0515] The converted text is analyzed by a generative AI model and scored as a statement that may constitute maternity harassment.

[0516] 3. Alert generation and notification

[0517] Based on the text whose score exceeds the threshold, the server generates an alert message saying, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[0518] The terminal receives this alert message and notifies the user.

[0519] 4. User Behavior

[0520] The user checks the alert message and realizes that his or her remarks were inappropriate.

[0521] As a result, users will reconsider what they said and be more considerate in future conversations.

[0522] ---

[0523] In this way, the system can detect maternity harassment in the workplace at an early stage and prevent inappropriate behavior before it occurs. It is expected that this system will provide a safe and comfortable working environment.

[0524] The processing flow will be explained below.

[0525] ---

[0526] Step 1:

[0527] Audio collection

[0528] The device uses a microphone to collect the user's conversational voice in real time. The voice data is recorded in a specified format (e.g., AAC or Opus). This collection continues from the moment the conversation starts until it stops.

[0529] Step 2:

[0530] Audio data compression and buffering

[0531] The device converts the collected audio data into a compressed format for efficient transmission to the server. The compressed audio data is divided into small data chunks and stored in a buffer for transmission.

[0532] Step 3:

[0533] Sending audio data

[0534] The device transmits the audio data stored in the buffer to the server at regular intervals. This transmission uses a low-latency communication protocol (e.g., WebSocket or HTTP / 2). The transmission is performed in real time and is monitored to ensure there is no interruption of data.

[0535] Step 4:

[0536] Receiving audio data

[0537] The server receives voice data sent from the terminal in real time, stores the received voice data in a fixed buffer, and prepares it for text conversion.

[0538] Step 5:

[0539] Decoding audio data

[0540] The server decodes the received compressed audio data and returns it to its original format. The decoded audio data is stored in temporary storage and passed on to the next processing step.

[0541] Step 6:

[0542] Speech recognition to text

[0543] The server sends the voice data stored in the temporary storage to a voice recognition API, which converts it into text data. The voice recognition API provides highly accurate voice recognition, and the converted text is stored in a database.

[0544] Step 7:

[0545] Preprocessing text data

[0546] The server cleanses the text data obtained from the speech recognition API to remove mistranslations and noise. This preprocessing contributes to improving the accuracy of the text data.

[0547] Step 8:

[0548] Analysis using generative AI models

[0549] The server inputs the preprocessed text data into a generative AI model, which analyzes the text data for inappropriate behavior. The generative AI model then detects inappropriate behavior with high accuracy based on a large amount of training data.

[0550] Step 9:

[0551] Detecting inappropriate behavior

[0552] Based on the analysis results of the generative AI model, the server determines whether the statements contained in the text data constitute inappropriate behavior. The determination is made using a scoring system, and statements that exceed a certain threshold are deemed inappropriate.

[0553] Step 10:

[0554] Generate an alert message

[0555] When inappropriate behavior is detected, the server generates an alert message containing the specific content of the behavior and a warning message. This alert message is sent to the perpetrator and the administrator.

[0556] Step 11:

[0557] Sending an alert message

[0558] The server then sends the generated alert message to the target device. The alert message is distributed in real time and, if necessary, notifies the administrator or harassment prevention department.

[0559] Step 12:

[0560] Receiving and Viewing Alerts

[0561] The terminal receives the alert message sent from the server and displays it on the user interface. The alert message is notified visually or audibly so that the user can immediately recognize it.

[0562] Step 13:

[0563] User Behavior

[0564] The user will check the alert message displayed on their device and realize that their remarks were inappropriate. Based on this realization, the user will be mindful of their behavior and speech in future conversations. It is also expected that the user will receive additional education and training as necessary.

[0565] ---

[0566] Through the above processing steps, the present invention can detect and prevent inappropriate behavior related to pregnancy, childbirth, and childcare in the workplace at an early stage, which is expected to provide a safe and comfortable working environment.

[0567] Example 1

[0568] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0569] In today's workplace, inappropriate verbal and physical behavior related to pregnancy, childbirth, and childcare (known as maternity harassment) is becoming more common, causing increased mental and physical strain on employees. In particular, unintentional comments made by managers and colleagues can cause stress to pregnant women and worsen the workplace environment. There is a need for methods to detect and prevent such inappropriate behavior early on.

[0570] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0571] In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data using a generative AI model to detect inappropriate speech and behavior, means for generating an alert message based on the detection results, means for notifying a terminal of the generated alert message, and means for displaying the notified alert message on a user interface. This enables early detection and prevention of inappropriate speech and behavior related to pregnancy, childbirth, and childcare in the workplace.

[0572] "Audio data" refers to a digital signal of sound captured using an acoustic device such as a microphone.

[0573] A "server" is a computer system for transmitting, receiving, and analyzing voice data.

[0574] "Text data" is character information converted from voice data using voice recognition technology.

[0575] A "generative AI model" is an artificial intelligence model that is trained to perform specific tasks based on input data.

[0576] "Analysis" is the process of determining whether text data contains inappropriate language or behavior.

[0577] "Inappropriate words and actions" are comments or actions that place mental or physical burden on employees in the workplace in relation to pregnancy, childbirth, or childcare.

[0578] An "alert message" is a message that alerts the user when inappropriate speech or behavior is detected.

[0579] A "terminal" is an information processing device such as a smartphone or PC used by a user.

[0580] The "user interface" refers to the screen and audio devices that allow the user to visually and audibly confirm the information displayed on the terminal.

[0581] A "microphone" is a device for converting sound into an electrical signal.

[0582] MODE FOR CARRYING OUT THE INVENTION

[0583] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare (so-called maternity harassment) in the workplace at an early stage. This system is composed of the following main elements for collecting and analyzing voice data and sending alerts.

[0584] System configuration

[0585] 1. Audio collection method

[0586] A device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time. The collected audio data is recorded uncompressed or in an appropriate compressed format (e.g., AAC or MP3) and then transmitted to a server.

[0587] 2. Means of transmitting audio data

[0588] The device sends the collected voice data to the server using a secure communication protocol (e.g., HTTPS or WebSocket). The voice data is encrypted and sent to the server in a privacy-protected manner.

[0589] 3. Speech-to-text methods

[0590] The server converts the transmitted voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The server sends the audio file to the speech recognition API and receives the returned text data.

[0591] 4. Methods for analyzing text data

[0592] The server inputs the converted text data into a generative AI model (e.g., OpenAI's ChatGPT) for analysis. The following prompt is used during analysis:

[0593] Analyze whether the following text is an inappropriate statement about pregnancy or morning sickness, and if so, output a score along with the reason.

[0594] Text: "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0595] The generative AI model analyzes the given text and returns a score and reason for detecting inappropriate behavior.

[0596] 5. Measures to detect inappropriate behavior

[0597] The server performs a scoring process based on the analysis results returned by the generative AI model. If the score exceeds a certain threshold, the remark is deemed to constitute inappropriate behavior (maternity harassment).

[0598] 6. How to Generate an Alert Message

[0599] If inappropriate behavior is detected, the server generates an alert message containing the specific content of the comment and a warning, such as "This comment contains comments about pregnancy and morning sickness. Please be considerate in your comments and behavior."

[0600] 7. Alert notification method

[0601] The terminal receives alert messages sent from the server in real time and displays them on the user interface. Alert messages are notified visually (banners and pop-up notifications) and audibly (voice alerts and chimes) so that users can quickly acknowledge them.

[0602] Specific examples

[0603] Conversation scenario

[0604] 1. Collecting comments

[0605] The user (boss) says, "Mr. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0606] The terminal collects this speech via a microphone and transmits the audio data to a server.

[0607] 2. Speech-to-text conversion and analysis

[0608] The server uses a speech recognition API to convert the collected voice data into text.

[0609] The converted text is parsed using prompts from a generative AI model.

[0610] 3. Detecting inappropriate behavior and generating alerts

[0611] The server generates an alert message containing specific statements and warnings based on text whose score exceeds the threshold.

[0612] The terminal receives this alert message and notifies the user.

[0613] In this way, the system can detect maternity harassment in the workplace early and prevent inappropriate behavior by immediately notifying users, which is expected to provide a safe and comfortable working environment.

[0614] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0615] Program processing steps and detailed explanations

[0616] Step 1: Collecting audio

[0617] The device (user's smartphone or PC) uses a microphone to collect the user's speech in real time. For example, the user might say, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress." This collected speech data (input) is recorded (output) either uncompressed or in an appropriate compressed format (e.g., AAC or MP3).

[0618] Step 2: Sending audio data

[0619] The device sends the collected voice data to the server using a secure communication protocol (e.g., HTTPS or WebSocket). The voice data (input) is encrypted and transmitted to protect the privacy of the data. The server receives this voice data (output).

[0620] Step 3: Speech to Text

[0621] The server sends the received voice data (input) to a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert it into text data. The speech recognition API analyzes the voice data and returns text data (output). The server receives this converted text data.

[0622] Step 4: Analyzing the text data

[0623] The server inputs the converted text data (input) into a generative AI model (e.g., OpenAI's ChatGPT) for analysis. The server sends the following prompt to the generative AI model:

[0624] Analyze whether the following text is an inappropriate statement about pregnancy or morning sickness, and if so, output a score along with the reason.

[0625] Text: "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0626] The generative AI model analyzes this prompt and returns a score indicating inappropriate behavior and the reason (output) to the server.

[0627] Step 5: Detect inappropriate behavior

[0628] The server performs a score based on the analysis results (input) returned from the generative AI model. If this score exceeds a certain threshold, the comment is deemed to be inappropriate behavior (maternity harassment). Specifically, for example, a score of "85 / 100" is output, and a reason such as "Reason: Because mentioning pregnancy or morning sickness involves personal privacy" is displayed.

[0629] Step 6: Generate an alert message

[0630] If the server detects inappropriate behavior, it generates an alert message (input) that includes the specific content of the comment and a warning. For example, it could generate a message (output) stating, "This comment contains comments about pregnancy and morning sickness. Please be considerate in your speech and behavior."

[0631] Step 7: Alert Notification

[0632] The terminal receives the alert message (input) sent from the server in real time and displays it on the user interface. The terminal notifies the user of the alert message visually (banner or pop-up notification) and audibly (voice alert or chime) so that the user can quickly acknowledge it (output). The user acknowledges the alert message and realizes that their comment was inappropriate.

[0633] (Application example 1)

[0634] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0635] In today's workplaces, inappropriate verbal and physical behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, remains a problem. If inappropriate comments and behavior are left unchecked, it can lead to a worsening work environment and have serious consequences for the mental and physical health of victims. However, conventional methods currently make it difficult to quickly detect and prevent such inappropriate behavior. Therefore, a system is needed that can collect and analyze voice data in real time and quickly detect and notify employees of inappropriate behavior in the workplace.

[0636] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0637] In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data using a generative AI model to detect inappropriate speech and behavior, means for generating an alert message based on the detection result and notifying the terminal of the alert message, means for notifying the alert message in real time via WebSocket, and means for displaying the notified alert message on a user interface. This makes it possible to quickly detect inappropriate speech and behavior in the workplace and respond in real time.

[0638] "Voice data" is data that represents voice in digital form, and is information that is collected in real time and is subject to analysis.

[0639] A "collection means" is a system or process for acquiring audio data using a device such as a microphone.

[0640] A "server" is a remote computer system that receives voice data, analyzes it, generates alerts, and performs other processing.

[0641] The "transmission means" is a mechanism for temporarily storing collected voice data and transferring it to a server for subsequent processing.

[0642] "Text data" is character string information converted from voice data using voice recognition technology.

[0643] "Conversion means" refers to technology or systems for converting voice data into text data, including voice recognition APIs.

[0644] "Analysis Methods" refers to the processes and techniques used to identify inappropriate behavior within text data using generative AI models.

[0645] A "generative AI model" is a machine learning model used to analyze text data and is an algorithm for detecting inappropriate behavior.

[0646] "Inappropriate behavior" refers to inappropriate remarks or actions related to pregnancy, childbirth, or childcare in the workplace.

[0647] An "alert message" is a notification message that alerts the user when inappropriate behavior is detected.

[0648] "Notification means" refers to a system or method for transmitting the generated alert message to a user's terminal in real time.

[0649] "User interface" refers to a screen or application that displays the notified alert message so that the user can check it.

[0650] "WebSocket" is a communication protocol for real-time communication between a server and a client.

[0651] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, in the workplace at an early stage. This system consists of the following main components for real-time collection and analysis of voice data and notification of alerts:

[0652] System configuration

[0653] 1. Audio collection method

[0654] This system uses a device (e.g., a smartphone, PC, or dedicated microphone device) to collect the user's speech in real time. The voice data collected through the microphone is recorded in uncompressed or compressed format and then sent to a server.

[0655] 2. Means of transmitting audio data

[0656] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is securely transferred using TLS, ensuring privacy.

[0657] 3. Speech-to-text methods

[0658] The server converts the received voice data into text data in real time using a speech recognition API (for example, Google Cloud Speech-to-Text). Using cloud services enables highly accurate speech recognition.

[0659] 4. Methods for analyzing text data

[0660] The server inputs the converted text data into a generative AI model (e.g., ChatGPT) and analyzes the text for inappropriate behavior. This analysis involves comparing the data with a database of past cases of maternity harassment. The generative AI model's analysis also uses prompt sentences to help detect specific behavior.

[0661] 5. Measures to detect inappropriate behavior

[0662] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior contains any inappropriate behavior. If the score exceeds a certain threshold, the speech or behavior is deemed to be highly likely to constitute maternity harassment.

[0663] 6. How to Generate an Alert Message

[0664] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The alert message is created based on the results of detailed analysis.

[0665] 7. Alert notification method

[0666] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified visually and audibly so that the user can quickly recognize it.

[0667] Specific examples

[0668] Conversation scenario

[0669] Collecting statements:

[0670] The user (boss) says, "Ms. B, I've been taking a lot of time off recently due to morning sickness, but please make sure that it doesn't affect my work progress." Such comments are collected by microphones in the workplace.

[0671] Speech-to-text and analysis:

[0672] The server uses a speech recognition API to convert the collected voice data into text, for example, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0673] The converted text is analyzed by a generative AI model, which scores the statement as inappropriate.

[0674] Alert generation and notification:

[0675] Based on the text whose score exceeds the threshold, the server generates an alert message stating, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[0676] The device receives this alert message via WebSocket and notifies the user in real time.

[0677] Example prompt sentence:

[0678] "The manager said, 'I have kids so it's a problem if you're always late.' Does this statement constitute inappropriate behavior?"

[0679] This system is expected to not only quickly detect inappropriate behavior in the workplace, but also enable real-time responses, thereby providing a safe and comfortable working environment.

[0680] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0681] Step 1:

[0682] The device uses a microphone in the workplace to collect voice data in real time. This allows the user's statements and conversation content to be captured in digital form. The input is voice, and the output is digital voice data. For example, if a user says, "Mr. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure it doesn't affect my work progress," the device will capture this as voice data.

[0683] Step 2:

[0684] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The input is digital audio data, and the output is the audio data sent to the server. Communication is securely encrypted using TLS.

[0685] Step 3:

[0686] The server receives the transmitted voice data and converts it into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is the voice data received by the server, and the output is the converted text data. Specifically, voice data such as "Ms. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure that it doesn't affect my work progress" is converted into text data such as "Ms. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure that it doesn't affect my work progress."

[0687] Step 4:

[0688] The server inputs the converted text data into a generative AI model (e.g., ChatGPT) and analyzes the inappropriate behavior in the text. A prompt is used for the analysis. The input is the converted text data, and the output is the analysis result. Specifically, the prompt, "Does this statement constitute inappropriate behavior?", is input into the generative AI model, and the analysis result is obtained.

[0689] Step 5:

[0690] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior is inappropriate. The input is the analysis result of the generative AI model, and the output is a score. If the score exceeds a certain threshold, the speech or behavior is deemed inappropriate.

[0691] Step 6:

[0692] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The input is the analysis result where the score exceeds the threshold, and the output is the alert message. Specifically, the alert message generated reads, "The remarks include comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[0693] Step 7:

[0694] The terminal receives alert messages sent from the server in real time via WebSocket and displays them on the user interface. The input is the alert message, and the output is the notification displayed on the user interface. Specifically, the user will be notified visually and / or audibly.

[0695] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0696] ---

[0697] This invention is a system that detects inappropriate behavior and emotions related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevents them from occurring. This system consists of the following main components for collecting and analyzing voice data, recognizing emotions, and issuing alerts.

[0698] System configuration

[0699] 1. Audio collection method

[0700] The device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[0701] 2. Means of transmitting audio data

[0702] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is transmitted in a secure manner, ensuring privacy.

[0703] 3. Speech-to-text methods

[0704] The server converts the received voice data into text data using a voice recognition API, which uses a cloud service capable of highly accurate recognition (for example, a voice recognition cloud service).

[0705] 4. Text data and sentiment analysis methods

[0706] The server inputs the converted text and voice data into a generative AI model and emotion engine, analyzes inappropriate behavior in the text, and recognizes the user's emotional state. The analysis is performed by comparing it with a database of past cases of maternity harassment.

[0707] 5. Measures to detect inappropriate behavior and emotions

[0708] The server calculates a score based on the results of analysis by the generative AI model and emotion engine to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed to be highly likely to constitute maternity harassment, and the user's emotional state is also used in the next process.

[0709] 6. How to Generate an Alert Message

[0710] When maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. This alert message will be used to appropriately notify the perpetrator and the administrator. The content of the alert message will also be adjusted based on the user's emotional state.

[0711] 7. Alert notification method

[0712] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[0713] Specific examples

[0714] Conversation scenario

[0715] 1. Collecting comments

[0716] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0717] The device collects this speech through a microphone and transmits the data to a server.

[0718] 2. Speech-to-text conversion and analysis

[0719] The server uses a speech recognition API to convert the collected voice data into text.

[0720] The converted text and speech data are analyzed by a generative AI model and emotion engine to score inappropriate language in the text and the user's emotional state.

[0721] 3. Alert generation and notification

[0722] Based on the text whose score exceeds the threshold, the server generates an alert message saying, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[0723] Depending on the user's emotional state, for example, if the user is emotionally charged, a supplemental message such as "Let's stay calm and continue the conversation" may be added.

[0724] The terminal receives this alert message and notifies the user.

[0725] 4. User Behavior

[0726] The user will receive an alert message, realizing that their remarks were inappropriate, and will also receive feedback about their emotional state.

[0727] As a result, users will reconsider what they have said and be more considerate in future conversations, while also paying attention to controlling their emotions.

[0728] ---

[0729] In this way, the system can detect maternity harassment and related emotional upheavals in the workplace at an early stage and prevent inappropriate behavior before it occurs, which is expected to provide a safe and comfortable working environment.

[0730] The processing flow will be explained below.

[0731] ---

[0732] Step 1:

[0733] Audio collection

[0734] The device uses a microphone to capture the user's conversation in real time, and the audio data is recorded in uncompressed or compressed format and then transmitted to a server.

[0735] Step 2:

[0736] Audio data compression and buffering

[0737] The device converts the collected audio data into a compressed format for efficient transmission to the server. The compressed audio data is divided into small data chunks and stored in a buffer for transmission.

[0738] Step 3:

[0739] Sending audio data

[0740] The device transmits the audio data stored in the buffer to the server at regular intervals. This transmission uses a low-latency communication protocol (e.g., WebSocket or HTTP / 2). The transmission is performed in real time and is monitored to ensure there is no interruption of data.

[0741] Step 4:

[0742] Receiving and decoding audio data

[0743] The server receives the audio data sent from the device in real time, decodes the compressed audio data, and stores it in temporary storage before passing it on to the next processing step.

[0744] Step 5:

[0745] Speech recognition to text

[0746] The server sends the voice data stored in the temporary storage to a voice recognition API, which converts it into text data. The voice recognition API provides highly accurate voice recognition, and the converted text data is stored in a database.

[0747] Step 6:

[0748] Preprocessing text data

[0749] The server cleanses the text data received from the speech recognition API, removing mistranslations and noise, and prepares the cleansed text data for analysis.

[0750] Step 7:

[0751] Analysis using generative AI models and emotion engines

[0752] The server inputs the cleansed text data into a generative AI model and an emotion engine, which analyzes the text for inappropriate behavior and the user's emotional state. The generative AI model detects inappropriate behavior with high accuracy based on a large amount of training data, and the emotion engine estimates the user's emotional state from the voice data.

[0753] Step 8:

[0754] Detecting inappropriate behavior and sentiment

[0755] Based on the analysis results, the server determines whether the statements contained in the text data constitute inappropriate behavior and simultaneously scores the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed inappropriate and the user's emotional state is also recorded.

[0756] Step 9:

[0757] Generate an alert message

[0758] When inappropriate behavior is detected, the server generates an alert message containing specific statements and a warning, and adjusts the content and tone of the alert message based on the user's emotional state and includes additional feedback as needed.

[0759] Step 10:

[0760] Sending an alert message

[0761] The server then sends the generated alert message to the target device. The alert message is distributed in real time and, if necessary, notifies the administrator or harassment prevention department.

[0762] Step 11:

[0763] Receiving and Viewing Alerts

[0764] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[0765] Step 12:

[0766] User Behavior

[0767] The user checks the alert message displayed on their device and realizes that their remarks were inappropriate. They also receive feedback on their emotional state. Based on this recognition, the user can be more considerate in future conversations and pay more attention to controlling their emotions.

[0768] ---

[0769] Through this step, the system of the present invention can detect inappropriate behavior and heightened emotions related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevent inappropriate behavior before it occurs, which is expected to provide a safe and comfortable working environment.

[0770] Example 2

[0771] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0772] In today's work environment, inappropriate behavior and attitudes related to pregnancy, childbirth, and childcare remain a problem. These issues can particularly increase stress and psychological burden on female employees in the workplace. Furthermore, managers often lack the means to identify and address these issues early before they become serious. As a result, maintaining a safe and comfortable work environment is difficult. This invention aims to improve the work environment and reduce the psychological burden on employees by detecting inappropriate behavior and the user's emotional state in the workplace early and responding promptly.

[0773] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0774] In this invention, the server includes a means for converting voice data into text data, a means for analyzing the converted text data and voice data to recognize the user's emotional state, and a means for determining inappropriate behavior and the user's emotional state based on the analysis results, thereby enabling prompt and accurate identification and response to inappropriate behavior related to pregnancy, childbirth, and childcare in the workplace.

[0775] "Audio data" means digital audio signals collected through an audio input device.

[0776] "Text data" refers to text information converted from voice data using a voice recognition API.

[0777] A "generative AI model" is an algorithm or machine learning model that uses artificial intelligence to analyze text data and detect inappropriate behavior.

[0778] An "emotion engine" is software for recognizing a user's emotional state from text data and voice tone.

[0779] An "alert message" is a notification message that is generated to encourage the user to behave considerately when inappropriate behavior or the user's emotional state is determined.

[0780] "Audio input device" refers to a microphone or other audio collection device for capturing speech in real time.

[0781] A "secure transfer protocol" is a communication technology for encrypting and transmitting data securely, and specifically refers to protocols such as HTTPS and WebSocket Secure.

[0782] A "prompt sentence" is the input text presented to a generative AI model, and is the sentence that serves as the starting point for analysis and generation.

[0783] This invention is a system that detects inappropriate behavior and the user's emotional state related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevents such behavior before it occurs. This system consists of the following main elements for collecting and analyzing voice data, recognizing emotions, and issuing alerts.

[0784] System configuration

[0785] Audio collection method

[0786] A device (e.g., a smartphone or a personal computer) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and transmitted to a server.

[0787] Voice data transmission method

[0788] The device sends the collected voice data to the server using an appropriate communication protocol (e.g., WebSocket Secure or HTTPS). The voice data is transmitted securely, ensuring privacy.

[0789] Voice-to-text conversion methods

[0790] The server converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text).

[0791] Text data and sentiment analysis methods

[0792] The server inputs the converted text and voice data into a generative AI model (e.g., GPT-4) and an emotion engine to analyze inappropriate behavior in the text and recognize the user's emotional state. The analysis is performed by comparing the data with a database of past cases of maternity harassment.

[0793] Measures to detect inappropriate behavior and emotions

[0794] The server calculates a score based on the results of analysis by the generative AI model and emotion engine to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed to be highly likely to constitute maternity harassment, and the user's emotional state is also used in the next process.

[0795] How to generate an alert message

[0796] When maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. This alert message will be used to appropriately notify the perpetrator and the administrator. The content of the alert message will also be adjusted based on the user's emotional state.

[0797] Alert notification method

[0798] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[0799] Specific examples

[0800] Conversation scenario

[0801] 1. Collecting comments

[0802] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0803] The device collects this speech through a microphone and transmits the data to a server.

[0804] 2. Speech-to-text conversion and analysis

[0805] The server uses a speech recognition API to convert the collected voice data into text.

[0806] The converted text and voice data are analyzed by a generative AI model (GPT-4) and an emotion engine to score inappropriate language in the text and the user's emotional state.

[0807] 3. Alert generation and notification

[0808] Based on the text whose score exceeds the threshold, the server generates an alert message stating, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[0809] Depending on the user's emotional state, for example, if the user is emotionally charged, a supplemental message such as "Let's stay calm and continue the conversation" is added.

[0810] The terminal receives this alert message and notifies the user.

[0811] 4. User Behavior

[0812] The user will receive an alert message, realizing that their remarks were inappropriate, and will also receive feedback about their emotional state.

[0813] As a result, users will reconsider what they have said and be more considerate in future conversations, while also paying attention to controlling their emotions.

[0814] Prompt Sentence Examples

[0815] "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0816] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0817] Step 1:

[0818] Audio collection

[0819] Input: User's speech

[0820] Specific operation: A user is having a conversation at work. The device (e.g., a smartphone or a personal computer) activates a built-in or external microphone to collect the conversation audio in real time.

[0821] Output: Audio data (digital format)

[0822] Step 2:

[0823] Sending audio data

[0824] Input: Audio data

[0825] Specific operation: The device encrypts the collected voice data using SSL / TLS and sends it to the server's API endpoint using a secure communication protocol (e.g., HTTPS or WebSocket Secure).

[0826] Output: Audio data sent to the server

[0827] Step 3:

[0828] Speech to text

[0829] Input: Audio data sent to the server

[0830] Specific operation: The server passes the received voice data to a speech recognition API (e.g., Google Cloud Speech-to-Text) and converts it into text data. The server receives the text data returned by the API.

[0831] Output: Converted text data

[0832] Step 4:

[0833] Text data and sentiment analysis

[0834] Input: Converted text and audio data

[0835] Specific operation: The server analyzes the text data entered as prompts into a generative AI model (e.g., GPT-4) and identifies inappropriate behavior within the text. In addition, it uses an emotion engine to recognize the user's emotional state from the tone of voice and text data. The analysis results are compared with a database of past maternity harassment cases.

[0836] Output: Analysis results and user emotional state data

[0837] Step 5:

[0838] Detecting inappropriate behavior and sentiment

[0839] Input: Analysis results and user emotional state data

[0840] Specific operation: The server calculates a score for inappropriate behavior and the user's emotional state based on the analysis results from the generative AI model and emotion engine. The score is compared with a pre-set threshold, and behavior that exceeds the threshold is deemed inappropriate.

[0841] Output: Judgment result of inappropriate behavior and emotional state

[0842] Step 6:

[0843] Generate an alert message

[0844] Input: Inappropriate behavior and emotional state assessment results

[0845] Specific actions: Based on the judgment result, the server generates an alert message containing the specific content of the comment and a warning. For example, it creates a message such as, "The comment contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." Based on the user's emotional state, it also adds a supplementary message such as, "Please remain calm and continue the conversation."

[0846] Output: The generated alert message

[0847] Step 7:

[0848] Alert Notification

[0849] Input: The generated alert message

[0850] Specific operation: The terminal receives the alert message sent from the server in real time and displays it on the user interface, specifically notifying the user by means of a pop-up notification, a sound notification, or the like.

[0851] Output: Alert message sent to the user

[0852] Specific examples

[0853] 1. User statement: The user (boss) says, "Ms. B, I've been taking a lot of time off recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0854] 2. What the device does: The built-in microphone collects this speech, encrypts the data, and sends it to a server.

[0855] 3. Server operation: The speech is converted into text using a speech recognition API, and then analyzed using a generative AI model (GPT-4) and an emotion engine.

[0856] 4. Analysis results: Inappropriate behavior is detected and the user is determined to be emotionally charged.

[0857] 5. Alert generation: An alert is generated stating, "This post contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." along with a supplementary message saying, "Please remain calm and continue the discussion."

[0858] 6. Device notification: Alert messages are notified to the user via pop-up and sound.

[0859] 7. User action: The user checks the notification, reviews their comments, and controls their emotions.

[0860] (Application example 2)

[0861] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0862] Inappropriate verbal and physical behavior and the emotional state of employees related to pregnancy, childbirth, and childcare often become problems in the workplace. However, it is difficult to detect these inappropriate verbal and physical behavior and emotional outbursts early and prevent them from occurring. In particular, achieving this in real time has been difficult with conventional technology. The objective of this invention is to provide a system that quickly detects inappropriate verbal and physical behavior and emotional outbursts in the workplace and provides a safe and comfortable working environment.

[0863] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data and detecting inappropriate speech and behavior and the user's emotional state, means for generating an alert message based on the detection result and the emotional state and notifying the terminal of the alert message, and means for displaying the notified alert message on a user interface. This makes it possible to detect inappropriate speech and behavior and heightened emotions in the workplace in real time and prevent them from occurring.

[0864] "Audio data" refers to electronically recorded data of a user's conversation or environmental sounds.

[0865] "Collection means" refers to devices and software for collecting voice data in real time.

[0866] "Transmission means" refers to the communication method or protocol used to send collected voice data to the server.

[0867] "Text data" refers to data obtained by converting voice data into text information using voice recognition technology.

[0868] "Conversion means" refers to the technology or device used to convert voice data into text data.

[0869] "Analysis means" refers to technologies and algorithms used to analyze the converted text data and detect inappropriate behavior or emotional states.

[0870] A "generative AI model" is an AI model trained using machine learning or deep learning, and is used to analyze text data.

[0871] "User's emotional state" refers to the emotional state that the user expresses through their words and actions, such as joy, anger, or anxiety.

[0872] An "alert message" is a warning message generated based on detected inappropriate behavior or emotional state.

[0873] "Notification means" refers to the communication method or protocol used to communicate the generated alert message to the user or administrator.

[0874] "User interface" refers to an interface for visually or audibly displaying an alert message to a user.

[0875] A "prompt sentence" is an initial sentence or instruction input to a generative AI model, and is used to indicate the direction of analysis.

[0876] This invention is a system that detects and prevents inappropriate behavior and emotional upset related to pregnancy, childbirth, and childcare in the workplace at an early stage. This system is mainly composed of components for collecting, transmitting, converting, and analyzing voice data, detecting inappropriate behavior, and generating and notifying alert messages.

[0877] System configuration

[0878] 1. Audio collection method

[0879] The device (e.g., a smartphone or PC) uses a microphone to collect the user's conversational voice in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[0880] 2. Means of transmitting audio data

[0881] The device sends the collected audio data to the server using an appropriate communication protocol (such as HTTP / 2). The audio data is transmitted securely to protect privacy.

[0882] 3. Speech-to-text methods

[0883] The server converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text), a cloud service that provides highly accurate recognition.

[0884] 4. Text data and sentiment analysis methods

[0885] The server then inputs the converted text and audio data into a generative AI model and a sentiment analysis engine (e.g., TextBlob) to analyze the text for inappropriate language and the user's emotional state. The analysis is based on the prompt sentence and identifies inappropriate phrases and heightened emotions.

[0886] As an example, use the following prompt:

[0887] "Turn this audio data into text and look for inappropriate language and heightened emotions. Pay attention to specific phrases and pick out negative comments."

[0888] 5. Measures to detect inappropriate behavior and emotions

[0889] The server uses the generative AI model and sentiment analysis engine to calculate a score from the analyzed text data to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed inappropriate.

[0890] 6. How to Generate an Alert Message

[0891] The server generates an alert message when inappropriate behavior and emotional upset are detected, which includes specific statements and reminders and is tailored to the user's emotional state.

[0892] 7. Alert notification method

[0893] The terminal receives alert messages sent from the server in real time and displays them on the user interface. The alert messages are notified to the user visually and audibly so that they can be quickly recognized.

[0894] Specific examples

[0895] 1. Collecting comments

[0896] The terminal collects the user's (boss') remarks through a microphone.

[0897] For example, a statement such as, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress" is collected.

[0898] 2. Speech-to-text conversion and analysis

[0899] The server uses a speech recognition API to convert the collected voice data into text.

[0900] The converted text is analyzed using a generative AI model and a sentiment analysis engine to score inappropriate behavior and emotional states.

[0901] 3. Alert generation and notification

[0902] The server detects inappropriate comments and generates an alert message saying, "This comment contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." If emotions are running high, a supplemental message is added saying, "Please remain calm and continue the conversation."

[0903] The terminal receives this alert message in real time and notifies the user.

[0904] 4. User Behavior

[0905] The user sees the alert message, realizes that what they said was inappropriate, and also receives feedback about their emotional state.

[0906] This will encourage users to review their future comments and be more considerate in their speech and actions.

[0907] This system will enable early detection of inappropriate behavior and emotional outbursts in the workplace, providing a safe and comfortable working environment.

[0908] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0909] ---

[0910] Step 1:

[0911] The device uses a microphone to collect conversational audio in real time. The input is the user's live voice, and the output is stored on the device as audio data. This audio data can be recorded in uncompressed or compressed format.

[0912] Step 2:

[0913] The device sends the collected audio data to the server using a secure communication protocol such as HTTP / 2. The input is audio data, and the output is audio data sent to the server. The communication is encrypted to protect privacy.

[0914] Step 3:

[0915] The server converts the received voice data into text data using a speech recognition API (Google Cloud Speech-to-Text). The input here is voice data, and the output is text data. This text data is obtained using speech recognition technology.

[0916] Step 4:

[0917] The server inputs the converted text data into a generative AI model and an emotion analysis engine (TextBlob) to analyze the inappropriate behavior and emotional state of the user in the text. The input here is text data, and the output is a score for the inappropriate behavior and emotional state. The generative AI model performs the analysis using prompts. For example, a prompt might be, "Convert this audio data into text and detect inappropriate behavior and heightened emotions. Pay attention to specific phrases and pick out negative comments."

[0918] Step 5:

[0919] The server uses the generative AI model and emotion analysis engine to analyze the results and determine whether the behavior and emotional state are inappropriate. If the score exceeds a certain threshold, the behavior is deemed inappropriate. The input here is the analysis score, and the output is the judgment of inappropriate behavior.

[0920] Step 6:

[0921] The server generates an alert message when inappropriate behavior or heightened emotions are detected. The input is the judgment result of inappropriate behavior, and the output is an alert message. The alert message includes the specific content of the statement and a warning, and if necessary, a supplemental message based on the emotional state is added.

[0922] Step 7:

[0923] The terminal receives alert messages sent from the server in real time and displays them on the user interface. The input here is the alert message, and the output is a notification to the user. The alert message is notified visually and audibly so that the user can quickly recognize it.

[0924] ---

[0925] The above processing steps make it possible to detect inappropriate behavior and emotional outbursts in the workplace in real time and prevent them from occurring.

[0926] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0927] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0928] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0929] [Third embodiment]

[0930] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0931] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0932] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0933] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0934] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0935] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0936] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0937] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0938] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0939] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0940] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0941] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0942] ---

[0943] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, in the workplace at an early stage. This system consists of the following main components for collecting and analyzing voice data and issuing alerts:

[0944] System configuration

[0945] 1. Audio collection method

[0946] The device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[0947] 2. Means of transmitting audio data

[0948] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is transmitted in a secure manner, ensuring privacy.

[0949] 3. Speech-to-text methods

[0950] The server converts the received voice data into text data using a voice recognition API, which uses a cloud service capable of highly accurate recognition (for example, a voice recognition cloud service).

[0951] 4. Methods for analyzing text data

[0952] The server then inputs the converted text data into a generative AI model (such as ChatGPT) and analyzes the text for inappropriate behavior by comparing it with a database of past cases of maternity harassment.

[0953] 5. Measures to detect inappropriate behavior

[0954] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior contains any inappropriate behavior. If the score exceeds a certain threshold, the speech or behavior is deemed to be highly likely to constitute maternity harassment.

[0955] 6. How to Generate an Alert Message

[0956] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The alert message is then sent appropriately to the affected user or administrator.

[0957] 7. Alert notification method

[0958] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[0959] Specific examples

[0960] Conversation scenario

[0961] 1. Collecting comments

[0962] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[0963] The device collects this speech through a microphone and transmits the data to a server.

[0964] 2. Speech-to-text conversion and analysis

[0965] The server uses a speech recognition API to convert the collected voice data into text.

[0966] The converted text is analyzed by a generative AI model and scored as a statement that may constitute maternity harassment.

[0967] 3. Alert generation and notification

[0968] Based on the text whose score exceeds the threshold, the server generates an alert message saying, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[0969] The terminal receives this alert message and notifies the user.

[0970] 4. User Behavior

[0971] The user checks the alert message and realizes that his or her remarks were inappropriate.

[0972] As a result, users will reconsider what they said and be more considerate in future conversations.

[0973] ---

[0974] In this way, the system can detect maternity harassment in the workplace at an early stage and prevent inappropriate behavior before it occurs. It is expected that this system will provide a safe and comfortable working environment.

[0975] The processing flow will be explained below.

[0976] ---

[0977] Step 1:

[0978] Audio collection

[0979] The device uses a microphone to collect the user's conversational voice in real time. The voice data is recorded in a specified format (e.g., AAC or Opus). This collection continues from the moment the conversation starts until it stops.

[0980] Step 2:

[0981] Audio data compression and buffering

[0982] The device converts the collected audio data into a compressed format for efficient transmission to the server. The compressed audio data is divided into small data chunks and stored in a buffer for transmission.

[0983] Step 3:

[0984] Sending audio data

[0985] The device transmits the audio data stored in the buffer to the server at regular intervals. This transmission uses a low-latency communication protocol (e.g., WebSocket or HTTP / 2). The transmission is performed in real time and is monitored to ensure there is no interruption of data.

[0986] Step 4:

[0987] Receiving audio data

[0988] The server receives voice data sent from the terminal in real time, stores the received voice data in a fixed buffer, and prepares it for text conversion.

[0989] Step 5:

[0990] Decoding audio data

[0991] The server decodes the received compressed audio data and returns it to its original format. The decoded audio data is stored in temporary storage and passed on to the next processing step.

[0992] Step 6:

[0993] Speech recognition to text

[0994] The server sends the voice data stored in the temporary storage to a voice recognition API, which converts it into text data. The voice recognition API provides highly accurate voice recognition, and the converted text is stored in a database.

[0995] Step 7:

[0996] Preprocessing text data

[0997] The server cleanses the text data obtained from the speech recognition API to remove mistranslations and noise. This preprocessing contributes to improving the accuracy of the text data.

[0998] Step 8:

[0999] Analysis using generative AI models

[1000] The server inputs the preprocessed text data into a generative AI model, which analyzes the text data for inappropriate behavior. The generative AI model then detects inappropriate behavior with high accuracy based on a large amount of training data.

[1001] Step 9:

[1002] Detecting inappropriate behavior

[1003] Based on the analysis results of the generative AI model, the server determines whether the statements contained in the text data constitute inappropriate behavior. The determination is made using a scoring system, and statements that exceed a certain threshold are deemed inappropriate.

[1004] Step 10:

[1005] Generate an alert message

[1006] When inappropriate behavior is detected, the server generates an alert message containing the specific content of the behavior and a warning message. This alert message is sent to the perpetrator and the administrator.

[1007] Step 11:

[1008] Sending an alert message

[1009] The server then sends the generated alert message to the target device. The alert message is distributed in real time and, if necessary, notifies the administrator or harassment prevention department.

[1010] Step 12:

[1011] Receiving and Viewing Alerts

[1012] The terminal receives the alert message sent from the server and displays it on the user interface. The alert message is notified visually or audibly so that the user can immediately recognize it.

[1013] Step 13:

[1014] User Behavior

[1015] The user will check the alert message displayed on their device and realize that their remarks were inappropriate. Based on this realization, the user will be mindful of their behavior and speech in future conversations. It is also expected that the user will receive additional education and training as necessary.

[1016] ---

[1017] Through the above processing steps, the present invention can detect and prevent inappropriate behavior related to pregnancy, childbirth, and childcare in the workplace at an early stage, which is expected to provide a safe and comfortable working environment.

[1018] Example 1

[1019] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1020] In today's workplace, inappropriate verbal and physical behavior related to pregnancy, childbirth, and childcare (known as maternity harassment) is becoming more common, causing increased mental and physical strain on employees. In particular, unintentional comments made by managers and colleagues can cause stress to pregnant women and worsen the workplace environment. There is a need for methods to detect and prevent such inappropriate behavior early on.

[1021] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1022] In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data using a generative AI model to detect inappropriate speech and behavior, means for generating an alert message based on the detection results, means for notifying a terminal of the generated alert message, and means for displaying the notified alert message on a user interface. This enables early detection and prevention of inappropriate speech and behavior related to pregnancy, childbirth, and childcare in the workplace.

[1023] "Audio data" refers to a digital signal of sound captured using an acoustic device such as a microphone.

[1024] A "server" is a computer system for transmitting, receiving, and analyzing voice data.

[1025] "Text data" is character information converted from voice data using voice recognition technology.

[1026] A "generative AI model" is an artificial intelligence model that is trained to perform specific tasks based on input data.

[1027] "Analysis" is the process of determining whether text data contains inappropriate language or behavior.

[1028] "Inappropriate words and actions" are comments or actions that place mental or physical burden on employees in the workplace in relation to pregnancy, childbirth, or childcare.

[1029] An "alert message" is a message that alerts the user when inappropriate speech or behavior is detected.

[1030] A "terminal" is an information processing device such as a smartphone or PC used by a user.

[1031] The "user interface" refers to the screen and audio devices that allow the user to visually and audibly confirm the information displayed on the terminal.

[1032] A "microphone" is a device for converting sound into an electrical signal.

[1033] MODE FOR CARRYING OUT THE INVENTION

[1034] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare (so-called maternity harassment) in the workplace at an early stage. This system is composed of the following main elements for collecting and analyzing voice data and sending alerts.

[1035] System configuration

[1036] 1. Audio collection method

[1037] A device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time. The collected audio data is recorded uncompressed or in an appropriate compressed format (e.g., AAC or MP3) and then transmitted to a server.

[1038] 2. Means of transmitting audio data

[1039] The device sends the collected voice data to the server using a secure communication protocol (e.g., HTTPS or WebSocket). The voice data is encrypted and sent to the server in a privacy-protected manner.

[1040] 3. Speech-to-text methods

[1041] The server converts the transmitted voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The server sends the audio file to the speech recognition API and receives the returned text data.

[1042] 4. Methods for analyzing text data

[1043] The server inputs the converted text data into a generative AI model (e.g., OpenAI's ChatGPT) for analysis. The following prompt is used during analysis:

[1044] Analyze whether the following text is an inappropriate statement about pregnancy or morning sickness, and if so, output a score along with the reason.

[1045] Text: "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1046] The generative AI model analyzes the given text and returns a score and reason for detecting inappropriate behavior.

[1047] 5. Measures to detect inappropriate behavior

[1048] The server performs a scoring process based on the analysis results returned by the generative AI model. If the score exceeds a certain threshold, the remark is deemed to constitute inappropriate behavior (maternity harassment).

[1049] 6. How to Generate an Alert Message

[1050] If inappropriate behavior is detected, the server generates an alert message containing the specific content of the comment and a warning, such as "This comment contains comments about pregnancy and morning sickness. Please be considerate in your comments and behavior."

[1051] 7. Alert notification method

[1052] The terminal receives alert messages sent from the server in real time and displays them on the user interface. Alert messages are notified visually (banners and pop-up notifications) and audibly (voice alerts and chimes) so that users can quickly acknowledge them.

[1053] Specific examples

[1054] Conversation scenario

[1055] 1. Collecting comments

[1056] The user (boss) says, "Mr. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1057] The terminal collects this speech via a microphone and transmits the audio data to a server.

[1058] 2. Speech-to-text conversion and analysis

[1059] The server uses a speech recognition API to convert the collected voice data into text.

[1060] The converted text is parsed using prompts from a generative AI model.

[1061] 3. Detecting inappropriate behavior and generating alerts

[1062] The server generates an alert message containing specific statements and warnings based on text whose score exceeds the threshold.

[1063] The terminal receives this alert message and notifies the user.

[1064] In this way, the system can detect maternity harassment in the workplace early and prevent inappropriate behavior by immediately notifying users, which is expected to provide a safe and comfortable working environment.

[1065] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1066] Program processing steps and detailed explanations

[1067] Step 1: Collecting audio

[1068] The device (user's smartphone or PC) uses a microphone to collect the user's speech in real time. For example, the user might say, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress." This collected speech data (input) is recorded (output) either uncompressed or in an appropriate compressed format (e.g., AAC or MP3).

[1069] Step 2: Sending audio data

[1070] The device sends the collected voice data to the server using a secure communication protocol (e.g., HTTPS or WebSocket). The voice data (input) is encrypted and transmitted to protect the privacy of the data. The server receives this voice data (output).

[1071] Step 3: Speech to Text

[1072] The server sends the received voice data (input) to a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert it into text data. The speech recognition API analyzes the voice data and returns text data (output). The server receives this converted text data.

[1073] Step 4: Analyzing the text data

[1074] The server inputs the converted text data (input) into a generative AI model (e.g., OpenAI's ChatGPT) for analysis. The server sends the following prompt to the generative AI model:

[1075] Analyze whether the following text is an inappropriate statement about pregnancy or morning sickness, and if so, output a score along with the reason.

[1076] Text: "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1077] The generative AI model analyzes this prompt and returns a score indicating inappropriate behavior and the reason (output) to the server.

[1078] Step 5: Detect inappropriate behavior

[1079] The server performs a score based on the analysis results (input) returned from the generative AI model. If this score exceeds a certain threshold, the comment is deemed to be inappropriate behavior (maternity harassment). Specifically, for example, a score of "85 / 100" is output, and a reason such as "Reason: Because mentioning pregnancy or morning sickness involves personal privacy" is displayed.

[1080] Step 6: Generate an alert message

[1081] If the server detects inappropriate behavior, it generates an alert message (input) that includes the specific content of the comment and a warning. For example, it could generate a message (output) stating, "This comment contains comments about pregnancy and morning sickness. Please be considerate in your speech and behavior."

[1082] Step 7: Alert Notification

[1083] The terminal receives the alert message (input) sent from the server in real time and displays it on the user interface. The terminal notifies the user of the alert message visually (banner or pop-up notification) and audibly (voice alert or chime) so that the user can quickly acknowledge it (output). The user acknowledges the alert message and realizes that their comment was inappropriate.

[1084] (Application example 1)

[1085] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1086] In today's workplaces, inappropriate verbal and physical behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, remains a problem. If inappropriate comments and behavior are left unchecked, it can lead to a worsening work environment and have serious consequences for the mental and physical health of victims. However, conventional methods currently make it difficult to quickly detect and prevent such inappropriate behavior. Therefore, a system is needed that can collect and analyze voice data in real time and quickly detect and notify employees of inappropriate behavior in the workplace.

[1087] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1088] In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data using a generative AI model to detect inappropriate speech and behavior, means for generating an alert message based on the detection result and notifying the terminal of the alert message, means for notifying the alert message in real time via WebSocket, and means for displaying the notified alert message on a user interface. This makes it possible to quickly detect inappropriate speech and behavior in the workplace and respond in real time.

[1089] "Voice data" is data that represents voice in digital form, and is information that is collected in real time and is subject to analysis.

[1090] A "collection means" is a system or process for acquiring audio data using a device such as a microphone.

[1091] A "server" is a remote computer system that receives voice data, analyzes it, generates alerts, and performs other processing.

[1092] The "transmission means" is a mechanism for temporarily storing collected voice data and transferring it to a server for subsequent processing.

[1093] "Text data" is character string information converted from voice data using voice recognition technology.

[1094] "Conversion means" refers to technology or systems for converting voice data into text data, including voice recognition APIs.

[1095] "Analysis Methods" refers to the processes and techniques used to identify inappropriate behavior within text data using generative AI models.

[1096] A "generative AI model" is a machine learning model used to analyze text data and is an algorithm for detecting inappropriate behavior.

[1097] "Inappropriate behavior" refers to inappropriate remarks or actions related to pregnancy, childbirth, or childcare in the workplace.

[1098] An "alert message" is a notification message that alerts the user when inappropriate behavior is detected.

[1099] "Notification means" refers to a system or method for transmitting the generated alert message to a user's terminal in real time.

[1100] "User interface" refers to a screen or application that displays the notified alert message so that the user can check it.

[1101] "WebSocket" is a communication protocol for real-time communication between a server and a client.

[1102] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, in the workplace at an early stage. This system consists of the following main components for real-time collection and analysis of voice data and notification of alerts:

[1103] System configuration

[1104] 1. Audio collection method

[1105] This system uses a device (e.g., a smartphone, PC, or dedicated microphone device) to collect the user's speech in real time. The voice data collected through the microphone is recorded in uncompressed or compressed format and then sent to a server.

[1106] 2. Means of transmitting audio data

[1107] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is securely transferred using TLS, ensuring privacy.

[1108] 3. Speech-to-text methods

[1109] The server converts the received voice data into text data in real time using a speech recognition API (for example, Google Cloud Speech-to-Text). Using cloud services enables highly accurate speech recognition.

[1110] 4. Methods for analyzing text data

[1111] The server inputs the converted text data into a generative AI model (e.g., ChatGPT) and analyzes the text for inappropriate behavior. This analysis involves comparing the data with a database of past cases of maternity harassment. The generative AI model's analysis also uses prompt sentences to help detect specific behavior.

[1112] 5. Measures to detect inappropriate behavior

[1113] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior contains any inappropriate behavior. If the score exceeds a certain threshold, the speech or behavior is deemed to be highly likely to constitute maternity harassment.

[1114] 6. How to Generate an Alert Message

[1115] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The alert message is created based on the results of detailed analysis.

[1116] 7. Alert notification method

[1117] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified visually and audibly so that the user can quickly recognize it.

[1118] Specific examples

[1119] Conversation scenario

[1120] Collecting statements:

[1121] The user (boss) says, "Ms. B, I've been taking a lot of time off recently due to morning sickness, but please make sure that it doesn't affect my work progress." Such comments are collected by microphones in the workplace.

[1122] Speech-to-text and analysis:

[1123] The server uses a speech recognition API to convert the collected voice data into text, for example, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1124] The converted text is analyzed by a generative AI model, which scores the statement as inappropriate.

[1125] Alert generation and notification:

[1126] Based on the text whose score exceeds the threshold, the server generates an alert message stating, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[1127] The device receives this alert message via WebSocket and notifies the user in real time.

[1128] Example prompt sentence:

[1129] "The manager said, 'I have kids so it's a problem if you're always late.' Does this statement constitute inappropriate behavior?"

[1130] This system is expected to not only quickly detect inappropriate behavior in the workplace, but also enable real-time responses, thereby providing a safe and comfortable working environment.

[1131] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1132] Step 1:

[1133] The device uses a microphone in the workplace to collect voice data in real time. This allows the user's statements and conversation content to be captured in digital form. The input is voice, and the output is digital voice data. For example, if a user says, "Mr. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure it doesn't affect my work progress," the device will capture this as voice data.

[1134] Step 2:

[1135] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The input is digital audio data, and the output is the audio data sent to the server. Communication is securely encrypted using TLS.

[1136] Step 3:

[1137] The server receives the transmitted voice data and converts it into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is the voice data received by the server, and the output is the converted text data. Specifically, voice data such as "Ms. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure that it doesn't affect my work progress" is converted into text data such as "Ms. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure that it doesn't affect my work progress."

[1138] Step 4:

[1139] The server inputs the converted text data into a generative AI model (e.g., ChatGPT) and analyzes the inappropriate behavior in the text. A prompt is used for the analysis. The input is the converted text data, and the output is the analysis result. Specifically, the prompt, "Does this statement constitute inappropriate behavior?", is input into the generative AI model, and the analysis result is obtained.

[1140] Step 5:

[1141] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior is inappropriate. The input is the analysis result of the generative AI model, and the output is a score. If the score exceeds a certain threshold, the speech or behavior is deemed inappropriate.

[1142] Step 6:

[1143] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The input is the analysis result where the score exceeds the threshold, and the output is the alert message. Specifically, the alert message generated reads, "The remarks include comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[1144] Step 7:

[1145] The terminal receives alert messages sent from the server in real time via WebSocket and displays them on the user interface. The input is the alert message, and the output is the notification displayed on the user interface. Specifically, the user will be notified visually and / or audibly.

[1146] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1147] ---

[1148] This invention is a system that detects inappropriate behavior and emotions related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevents them from occurring. This system consists of the following main components for collecting and analyzing voice data, recognizing emotions, and issuing alerts.

[1149] System configuration

[1150] 1. Audio collection method

[1151] The device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[1152] 2. Means of transmitting audio data

[1153] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is transmitted in a secure manner, ensuring privacy.

[1154] 3. Speech-to-text methods

[1155] The server converts the received voice data into text data using a voice recognition API, which uses a cloud service capable of highly accurate recognition (for example, a voice recognition cloud service).

[1156] 4. Text data and sentiment analysis methods

[1157] The server inputs the converted text and voice data into a generative AI model and emotion engine, analyzes inappropriate behavior in the text, and recognizes the user's emotional state. The analysis is performed by comparing it with a database of past cases of maternity harassment.

[1158] 5. Measures to detect inappropriate behavior and emotions

[1159] The server calculates a score based on the results of analysis by the generative AI model and emotion engine to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed to be highly likely to constitute maternity harassment, and the user's emotional state is also used in the next process.

[1160] 6. How to Generate an Alert Message

[1161] When maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. This alert message will be used to appropriately notify the perpetrator and the administrator. The content of the alert message will also be adjusted based on the user's emotional state.

[1162] 7. Alert notification method

[1163] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[1164] Specific examples

[1165] Conversation scenario

[1166] 1. Collecting comments

[1167] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1168] The device collects this speech through a microphone and transmits the data to a server.

[1169] 2. Speech-to-text conversion and analysis

[1170] The server uses a speech recognition API to convert the collected voice data into text.

[1171] The converted text and speech data are analyzed by a generative AI model and emotion engine to score inappropriate language in the text and the user's emotional state.

[1172] 3. Alert generation and notification

[1173] Based on the text whose score exceeds the threshold, the server generates an alert message saying, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[1174] Depending on the user's emotional state, for example, if the user is emotionally charged, a supplemental message such as "Let's stay calm and continue the conversation" may be added.

[1175] The terminal receives this alert message and notifies the user.

[1176] 4. User Behavior

[1177] The user will receive an alert message, realizing that their remarks were inappropriate, and will also receive feedback about their emotional state.

[1178] As a result, users will reconsider what they have said and be more considerate in future conversations, while also paying attention to controlling their emotions.

[1179] ---

[1180] In this way, the system can detect maternity harassment and related emotional upheavals in the workplace at an early stage and prevent inappropriate behavior before it occurs, which is expected to provide a safe and comfortable working environment.

[1181] The processing flow will be explained below.

[1182] ---

[1183] Step 1:

[1184] Audio collection

[1185] The device uses a microphone to capture the user's conversation in real time, and the audio data is recorded in uncompressed or compressed format and then transmitted to a server.

[1186] Step 2:

[1187] Audio data compression and buffering

[1188] The device converts the collected audio data into a compressed format for efficient transmission to the server. The compressed audio data is divided into small data chunks and stored in a buffer for transmission.

[1189] Step 3:

[1190] Sending audio data

[1191] The device transmits the audio data stored in the buffer to the server at regular intervals. This transmission uses a low-latency communication protocol (e.g., WebSocket or HTTP / 2). The transmission is performed in real time and is monitored to ensure there is no interruption of data.

[1192] Step 4:

[1193] Receiving and decoding audio data

[1194] The server receives the audio data sent from the device in real time, decodes the compressed audio data, and stores it in temporary storage before passing it on to the next processing step.

[1195] Step 5:

[1196] Speech recognition to text

[1197] The server sends the voice data stored in the temporary storage to a voice recognition API, which converts it into text data. The voice recognition API provides highly accurate voice recognition, and the converted text data is stored in a database.

[1198] Step 6:

[1199] Preprocessing text data

[1200] The server cleanses the text data received from the speech recognition API, removing mistranslations and noise, and prepares the cleansed text data for analysis.

[1201] Step 7:

[1202] Analysis using generative AI models and emotion engines

[1203] The server inputs the cleansed text data into a generative AI model and an emotion engine, which analyzes the text for inappropriate behavior and the user's emotional state. The generative AI model detects inappropriate behavior with high accuracy based on a large amount of training data, and the emotion engine estimates the user's emotional state from the voice data.

[1204] Step 8:

[1205] Detecting inappropriate behavior and sentiment

[1206] Based on the analysis results, the server determines whether the statements contained in the text data constitute inappropriate behavior and simultaneously scores the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed inappropriate and the user's emotional state is also recorded.

[1207] Step 9:

[1208] Generate an alert message

[1209] When inappropriate behavior is detected, the server generates an alert message containing specific statements and a warning, and adjusts the content and tone of the alert message based on the user's emotional state and includes additional feedback as needed.

[1210] Step 10:

[1211] Sending an alert message

[1212] The server then sends the generated alert message to the target device. The alert message is distributed in real time and, if necessary, notifies the administrator or harassment prevention department.

[1213] Step 11:

[1214] Receiving and Viewing Alerts

[1215] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[1216] Step 12:

[1217] User Behavior

[1218] The user checks the alert message displayed on their device and realizes that their remarks were inappropriate. They also receive feedback on their emotional state. Based on this recognition, the user can be more considerate in future conversations and pay more attention to controlling their emotions.

[1219] ---

[1220] Through this step, the system of the present invention can detect inappropriate behavior and heightened emotions related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevent inappropriate behavior before it occurs, which is expected to provide a safe and comfortable working environment.

[1221] Example 2

[1222] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1223] In today's work environment, inappropriate behavior and attitudes related to pregnancy, childbirth, and childcare remain a problem. These issues can particularly increase stress and psychological burden on female employees in the workplace. Furthermore, managers often lack the means to identify and address these issues early before they become serious. As a result, maintaining a safe and comfortable work environment is difficult. This invention aims to improve the work environment and reduce the psychological burden on employees by detecting inappropriate behavior and the user's emotional state in the workplace early and responding promptly.

[1224] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1225] In this invention, the server includes a means for converting voice data into text data, a means for analyzing the converted text data and voice data to recognize the user's emotional state, and a means for determining inappropriate behavior and the user's emotional state based on the analysis results, thereby enabling prompt and accurate identification and response to inappropriate behavior related to pregnancy, childbirth, and childcare in the workplace.

[1226] "Audio data" means digital audio signals collected through an audio input device.

[1227] "Text data" refers to text information converted from voice data using a voice recognition API.

[1228] A "generative AI model" is an algorithm or machine learning model that uses artificial intelligence to analyze text data and detect inappropriate behavior.

[1229] An "emotion engine" is software for recognizing a user's emotional state from text data and voice tone.

[1230] An "alert message" is a notification message that is generated to encourage the user to behave considerately when inappropriate behavior or the user's emotional state is determined.

[1231] "Audio input device" refers to a microphone or other audio collection device for capturing speech in real time.

[1232] A "secure transfer protocol" is a communication technology for encrypting and transmitting data securely, and specifically refers to protocols such as HTTPS and WebSocket Secure.

[1233] A "prompt sentence" is the input text presented to a generative AI model, and is the sentence that serves as the starting point for analysis and generation.

[1234] This invention is a system that detects inappropriate behavior and the user's emotional state related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevents such behavior before it occurs. This system consists of the following main elements for collecting and analyzing voice data, recognizing emotions, and issuing alerts.

[1235] System configuration

[1236] Audio collection method

[1237] A device (e.g., a smartphone or a personal computer) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and transmitted to a server.

[1238] Voice data transmission method

[1239] The device sends the collected voice data to the server using an appropriate communication protocol (e.g., WebSocket Secure or HTTPS). The voice data is transmitted securely, ensuring privacy.

[1240] Voice-to-text conversion methods

[1241] The server converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text).

[1242] Text data and sentiment analysis methods

[1243] The server inputs the converted text and voice data into a generative AI model (e.g., GPT-4) and an emotion engine to analyze inappropriate behavior in the text and recognize the user's emotional state. The analysis is performed by comparing the data with a database of past cases of maternity harassment.

[1244] Measures to detect inappropriate behavior and emotions

[1245] The server calculates a score based on the results of analysis by the generative AI model and emotion engine to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed to be highly likely to constitute maternity harassment, and the user's emotional state is also used in the next process.

[1246] How to generate an alert message

[1247] When maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. This alert message will be used to appropriately notify the perpetrator and the administrator. The content of the alert message will also be adjusted based on the user's emotional state.

[1248] Alert notification method

[1249] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[1250] Specific examples

[1251] Conversation scenario

[1252] 1. Collecting comments

[1253] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1254] The device collects this speech through a microphone and transmits the data to a server.

[1255] 2. Speech-to-text conversion and analysis

[1256] The server uses a speech recognition API to convert the collected voice data into text.

[1257] The converted text and voice data are analyzed by a generative AI model (GPT-4) and an emotion engine to score inappropriate language in the text and the user's emotional state.

[1258] 3. Alert generation and notification

[1259] Based on the text whose score exceeds the threshold, the server generates an alert message stating, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[1260] Depending on the user's emotional state, for example, if the user is emotionally charged, a supplemental message such as "Let's stay calm and continue the conversation" is added.

[1261] The terminal receives this alert message and notifies the user.

[1262] 4. User Behavior

[1263] The user will receive an alert message, realizing that their remarks were inappropriate, and will also receive feedback about their emotional state.

[1264] As a result, users will reconsider what they have said and be more considerate in future conversations, while also paying attention to controlling their emotions.

[1265] Prompt Sentence Examples

[1266] "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1267] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1268] Step 1:

[1269] Audio collection

[1270] Input: User's speech

[1271] Specific operation: A user is having a conversation at work. The device (e.g., a smartphone or a personal computer) activates a built-in or external microphone to collect the conversation audio in real time.

[1272] Output: Audio data (digital format)

[1273] Step 2:

[1274] Sending audio data

[1275] Input: Audio data

[1276] Specific operation: The device encrypts the collected voice data using SSL / TLS and sends it to the server's API endpoint using a secure communication protocol (e.g., HTTPS or WebSocket Secure).

[1277] Output: Audio data sent to the server

[1278] Step 3:

[1279] Speech to text

[1280] Input: Audio data sent to the server

[1281] Specific operation: The server passes the received voice data to a speech recognition API (e.g., Google Cloud Speech-to-Text) and converts it into text data. The server receives the text data returned by the API.

[1282] Output: Converted text data

[1283] Step 4:

[1284] Text data and sentiment analysis

[1285] Input: Converted text and audio data

[1286] Specific operation: The server analyzes the text data entered as prompts into a generative AI model (e.g., GPT-4) and identifies inappropriate behavior within the text. In addition, it uses an emotion engine to recognize the user's emotional state from the tone of voice and text data. The analysis results are compared with a database of past maternity harassment cases.

[1287] Output: Analysis results and user emotional state data

[1288] Step 5:

[1289] Detecting inappropriate behavior and sentiment

[1290] Input: Analysis results and user emotional state data

[1291] Specific operation: The server calculates a score for inappropriate behavior and the user's emotional state based on the analysis results from the generative AI model and emotion engine. The score is compared with a pre-set threshold, and behavior that exceeds the threshold is deemed inappropriate.

[1292] Output: Judgment result of inappropriate behavior and emotional state

[1293] Step 6:

[1294] Generate an alert message

[1295] Input: Inappropriate behavior and emotional state assessment results

[1296] Specific actions: Based on the judgment result, the server generates an alert message containing the specific content of the comment and a warning. For example, it creates a message such as, "The comment contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." Based on the user's emotional state, it also adds a supplementary message such as, "Please remain calm and continue the conversation."

[1297] Output: The generated alert message

[1298] Step 7:

[1299] Alert Notification

[1300] Input: The generated alert message

[1301] Specific operation: The terminal receives the alert message sent from the server in real time and displays it on the user interface, specifically notifying the user by means of a pop-up notification, a sound notification, or the like.

[1302] Output: Alert message sent to the user

[1303] Specific examples

[1304] 1. User statement: The user (boss) says, "Ms. B, I've been taking a lot of time off recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1305] 2. What the device does: The built-in microphone collects this speech, encrypts the data, and sends it to a server.

[1306] 3. Server operation: The speech is converted into text using a speech recognition API, and then analyzed using a generative AI model (GPT-4) and an emotion engine.

[1307] 4. Analysis results: Inappropriate behavior is detected and the user is determined to be emotionally charged.

[1308] 5. Alert generation: An alert is generated stating, "This post contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." along with a supplementary message saying, "Please remain calm and continue the discussion."

[1309] 6. Device notification: Alert messages are notified to the user via pop-up and sound.

[1310] 7. User action: The user checks the notification, reviews their comments, and controls their emotions.

[1311] (Application example 2)

[1312] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1313] Inappropriate verbal and physical behavior and the emotional state of employees related to pregnancy, childbirth, and childcare often become problems in the workplace. However, it is difficult to detect these inappropriate verbal and physical behavior and emotional outbursts early and prevent them from occurring. In particular, achieving this in real time has been difficult with conventional technology. The objective of this invention is to provide a system that quickly detects inappropriate verbal and physical behavior and emotional outbursts in the workplace and provides a safe and comfortable working environment.

[1314] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data and detecting inappropriate speech and behavior and the user's emotional state, means for generating an alert message based on the detection result and the emotional state and notifying the terminal of the alert message, and means for displaying the notified alert message on a user interface. This makes it possible to detect inappropriate speech and behavior and heightened emotions in the workplace in real time and prevent them from occurring.

[1315] "Audio data" refers to electronically recorded data of a user's conversation or environmental sounds.

[1316] "Collection means" refers to devices and software for collecting voice data in real time.

[1317] "Transmission means" refers to the communication method or protocol used to send collected voice data to the server.

[1318] "Text data" refers to data obtained by converting voice data into text information using voice recognition technology.

[1319] "Conversion means" refers to the technology or device used to convert voice data into text data.

[1320] "Analysis means" refers to technologies and algorithms used to analyze the converted text data and detect inappropriate behavior or emotional states.

[1321] A "generative AI model" is an AI model trained using machine learning or deep learning, and is used to analyze text data.

[1322] "User's emotional state" refers to the emotional state that the user expresses through their words and actions, such as joy, anger, or anxiety.

[1323] An "alert message" is a warning message generated based on detected inappropriate behavior or emotional state.

[1324] "Notification means" refers to the communication method or protocol used to communicate the generated alert message to the user or administrator.

[1325] "User interface" refers to an interface for visually or audibly displaying an alert message to a user.

[1326] A "prompt sentence" is an initial sentence or instruction input to a generative AI model, and is used to indicate the direction of analysis.

[1327] This invention is a system that detects and prevents inappropriate behavior and emotional upset related to pregnancy, childbirth, and childcare in the workplace at an early stage. This system is mainly composed of components for collecting, transmitting, converting, and analyzing voice data, detecting inappropriate behavior, and generating and notifying alert messages.

[1328] System configuration

[1329] 1. Audio collection method

[1330] The device (e.g., a smartphone or PC) uses a microphone to collect the user's conversational voice in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[1331] 2. Means of transmitting audio data

[1332] The device sends the collected audio data to the server using an appropriate communication protocol (such as HTTP / 2). The audio data is transmitted securely to protect privacy.

[1333] 3. Speech-to-text methods

[1334] The server converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text), a cloud service that provides highly accurate recognition.

[1335] 4. Text data and sentiment analysis methods

[1336] The server then inputs the converted text and audio data into a generative AI model and a sentiment analysis engine (e.g., TextBlob) to analyze the text for inappropriate language and the user's emotional state. The analysis is based on the prompt sentence and identifies inappropriate phrases and heightened emotions.

[1337] As an example, use the following prompt:

[1338] "Turn this audio data into text and look for inappropriate language and heightened emotions. Pay attention to specific phrases and pick out negative comments."

[1339] 5. Measures to detect inappropriate behavior and emotions

[1340] The server uses the generative AI model and sentiment analysis engine to calculate a score from the analyzed text data to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed inappropriate.

[1341] 6. How to Generate an Alert Message

[1342] The server generates an alert message when inappropriate behavior and emotional upset are detected, which includes specific statements and reminders and is tailored to the user's emotional state.

[1343] 7. Alert notification method

[1344] The terminal receives alert messages sent from the server in real time and displays them on the user interface. The alert messages are notified to the user visually and audibly so that they can be quickly recognized.

[1345] Specific examples

[1346] 1. Collecting comments

[1347] The terminal collects the user's (boss') remarks through a microphone.

[1348] For example, a statement such as, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress" is collected.

[1349] 2. Speech-to-text conversion and analysis

[1350] The server uses a speech recognition API to convert the collected voice data into text.

[1351] The converted text is analyzed using a generative AI model and a sentiment analysis engine to score inappropriate behavior and emotional states.

[1352] 3. Alert generation and notification

[1353] The server detects inappropriate comments and generates an alert message saying, "This comment contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." If emotions are running high, a supplemental message is added saying, "Please remain calm and continue the conversation."

[1354] The terminal receives this alert message in real time and notifies the user.

[1355] 4. User Behavior

[1356] The user sees the alert message, realizes that what they said was inappropriate, and also receives feedback about their emotional state.

[1357] This will encourage users to review their future comments and be more considerate in their speech and actions.

[1358] This system will enable early detection of inappropriate behavior and emotional outbursts in the workplace, providing a safe and comfortable working environment.

[1359] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1360] ---

[1361] Step 1:

[1362] The device uses a microphone to collect conversational audio in real time. The input is the user's live voice, and the output is stored on the device as audio data. This audio data can be recorded in uncompressed or compressed format.

[1363] Step 2:

[1364] The device sends the collected audio data to the server using a secure communication protocol such as HTTP / 2. The input is audio data, and the output is audio data sent to the server. The communication is encrypted to protect privacy.

[1365] Step 3:

[1366] The server converts the received voice data into text data using a speech recognition API (Google Cloud Speech-to-Text). The input here is voice data, and the output is text data. This text data is obtained using speech recognition technology.

[1367] Step 4:

[1368] The server inputs the converted text data into a generative AI model and an emotion analysis engine (TextBlob) to analyze the inappropriate behavior and emotional state of the user in the text. The input here is text data, and the output is a score for the inappropriate behavior and emotional state. The generative AI model performs the analysis using prompts. For example, a prompt might be, "Convert this audio data into text and detect inappropriate behavior and heightened emotions. Pay attention to specific phrases and pick out negative comments."

[1369] Step 5:

[1370] The server uses the generative AI model and emotion analysis engine to analyze the results and determine whether the behavior and emotional state are inappropriate. If the score exceeds a certain threshold, the behavior is deemed inappropriate. The input here is the analysis score, and the output is the judgment of inappropriate behavior.

[1371] Step 6:

[1372] The server generates an alert message when inappropriate behavior or heightened emotions are detected. The input is the judgment result of inappropriate behavior, and the output is an alert message. The alert message includes the specific content of the statement and a warning, and if necessary, a supplemental message based on the emotional state is added.

[1373] Step 7:

[1374] The terminal receives alert messages sent from the server in real time and displays them on the user interface. The input here is the alert message, and the output is a notification to the user. The alert message is notified visually and audibly so that the user can quickly recognize it.

[1375] ---

[1376] The above processing steps make it possible to detect inappropriate behavior and emotional outbursts in the workplace in real time and prevent them from occurring.

[1377] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1378] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1379] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1380] [Fourth embodiment]

[1381] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1382] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1383] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1384] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1385] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1386] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1387] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1388] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1389] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1390] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1391] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1392] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1393] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1394] ---

[1395] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, in the workplace at an early stage. This system consists of the following main components for collecting and analyzing voice data and issuing alerts:

[1396] System configuration

[1397] 1. Audio collection method

[1398] The device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[1399] 2. Means of transmitting audio data

[1400] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is transmitted in a secure manner, ensuring privacy.

[1401] 3. Speech-to-text methods

[1402] The server converts the received voice data into text data using a voice recognition API, which uses a cloud service capable of highly accurate recognition (for example, a voice recognition cloud service).

[1403] 4. Methods for analyzing text data

[1404] The server then inputs the converted text data into a generative AI model (such as ChatGPT) and analyzes the text for inappropriate behavior by comparing it with a database of past cases of maternity harassment.

[1405] 5. Measures to detect inappropriate behavior

[1406] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior contains any inappropriate behavior. If the score exceeds a certain threshold, the speech or behavior is deemed to be highly likely to constitute maternity harassment.

[1407] 6. How to Generate an Alert Message

[1408] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The alert message is then sent appropriately to the affected user or administrator.

[1409] 7. Alert notification method

[1410] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[1411] Specific examples

[1412] Conversation scenario

[1413] 1. Collecting comments

[1414] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1415] The device collects this speech through a microphone and transmits the data to a server.

[1416] 2. Speech-to-text conversion and analysis

[1417] The server uses a speech recognition API to convert the collected voice data into text.

[1418] The converted text is analyzed by a generative AI model and scored as a statement that may constitute maternity harassment.

[1419] 3. Alert generation and notification

[1420] Based on the text whose score exceeds the threshold, the server generates an alert message saying, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[1421] The terminal receives this alert message and notifies the user.

[1422] 4. User Behavior

[1423] The user checks the alert message and realizes that his or her remarks were inappropriate.

[1424] As a result, users will reconsider what they said and be more considerate in future conversations.

[1425] ---

[1426] In this way, the system can detect maternity harassment in the workplace at an early stage and prevent inappropriate behavior before it occurs. It is expected that this system will provide a safe and comfortable working environment.

[1427] The processing flow will be explained below.

[1428] ---

[1429] Step 1:

[1430] Audio collection

[1431] The device uses a microphone to collect the user's conversational voice in real time. The voice data is recorded in a specified format (e.g., AAC or Opus). This collection continues from the moment the conversation starts until it stops.

[1432] Step 2:

[1433] Audio data compression and buffering

[1434] The device converts the collected audio data into a compressed format for efficient transmission to the server. The compressed audio data is divided into small data chunks and stored in a buffer for transmission.

[1435] Step 3:

[1436] Sending audio data

[1437] The device transmits the audio data stored in the buffer to the server at regular intervals. This transmission uses a low-latency communication protocol (e.g., WebSocket or HTTP / 2). The transmission is performed in real time and is monitored to ensure there is no interruption of data.

[1438] Step 4:

[1439] Receiving audio data

[1440] The server receives voice data sent from the terminal in real time, stores the received voice data in a fixed buffer, and prepares it for text conversion.

[1441] Step 5:

[1442] Decoding audio data

[1443] The server decodes the received compressed audio data and returns it to its original format. The decoded audio data is stored in temporary storage and passed on to the next processing step.

[1444] Step 6:

[1445] Speech recognition to text

[1446] The server sends the voice data stored in the temporary storage to a voice recognition API, which converts it into text data. The voice recognition API provides highly accurate voice recognition, and the converted text is stored in a database.

[1447] Step 7:

[1448] Preprocessing text data

[1449] The server cleanses the text data obtained from the speech recognition API to remove mistranslations and noise. This preprocessing contributes to improving the accuracy of the text data.

[1450] Step 8:

[1451] Analysis using generative AI models

[1452] The server inputs the preprocessed text data into a generative AI model, which analyzes the text data for inappropriate behavior. The generative AI model then detects inappropriate behavior with high accuracy based on a large amount of training data.

[1453] Step 9:

[1454] Detecting inappropriate behavior

[1455] Based on the analysis results of the generative AI model, the server determines whether the statements contained in the text data constitute inappropriate behavior. The determination is made using a scoring system, and statements that exceed a certain threshold are deemed inappropriate.

[1456] Step 10:

[1457] Generate an alert message

[1458] When inappropriate behavior is detected, the server generates an alert message containing the specific content of the behavior and a warning message. This alert message is sent to the perpetrator and the administrator.

[1459] Step 11:

[1460] Sending an alert message

[1461] The server then sends the generated alert message to the target device. The alert message is distributed in real time and, if necessary, notifies the administrator or harassment prevention department.

[1462] Step 12:

[1463] Receiving and Viewing Alerts

[1464] The terminal receives the alert message sent from the server and displays it on the user interface. The alert message is notified visually or audibly so that the user can immediately recognize it.

[1465] Step 13:

[1466] User Behavior

[1467] The user will check the alert message displayed on their device and realize that their remarks were inappropriate. Based on this realization, the user will be mindful of their behavior and speech in future conversations. It is also expected that the user will receive additional education and training as necessary.

[1468] ---

[1469] Through the above processing steps, the present invention can detect and prevent inappropriate behavior related to pregnancy, childbirth, and childcare in the workplace at an early stage, which is expected to provide a safe and comfortable working environment.

[1470] Example 1

[1471] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1472] In today's workplace, inappropriate verbal and physical behavior related to pregnancy, childbirth, and childcare (known as maternity harassment) is becoming more common, causing increased mental and physical strain on employees. In particular, unintentional comments made by managers and colleagues can cause stress to pregnant women and worsen the workplace environment. There is a need for methods to detect and prevent such inappropriate behavior early on.

[1473] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1474] In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data using a generative AI model to detect inappropriate speech and behavior, means for generating an alert message based on the detection results, means for notifying a terminal of the generated alert message, and means for displaying the notified alert message on a user interface. This enables early detection and prevention of inappropriate speech and behavior related to pregnancy, childbirth, and childcare in the workplace.

[1475] "Audio data" refers to a digital signal of sound captured using an acoustic device such as a microphone.

[1476] A "server" is a computer system for transmitting, receiving, and analyzing voice data.

[1477] "Text data" is character information converted from voice data using voice recognition technology.

[1478] A "generative AI model" is an artificial intelligence model that is trained to perform specific tasks based on input data.

[1479] "Analysis" is the process of determining whether text data contains inappropriate language or behavior.

[1480] "Inappropriate words and actions" are comments or actions that place mental or physical burden on employees in the workplace in relation to pregnancy, childbirth, or childcare.

[1481] An "alert message" is a message that alerts the user when inappropriate speech or behavior is detected.

[1482] A "terminal" is an information processing device such as a smartphone or PC used by a user.

[1483] The "user interface" refers to the screen and audio devices that allow the user to visually and audibly confirm the information displayed on the terminal.

[1484] A "microphone" is a device for converting sound into an electrical signal.

[1485] MODE FOR CARRYING OUT THE INVENTION

[1486] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare (so-called maternity harassment) in the workplace at an early stage. This system is composed of the following main elements for collecting and analyzing voice data and sending alerts.

[1487] System configuration

[1488] 1. Audio collection method

[1489] A device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time. The collected audio data is recorded uncompressed or in an appropriate compressed format (e.g., AAC or MP3) and then transmitted to a server.

[1490] 2. Means of transmitting audio data

[1491] The device sends the collected voice data to the server using a secure communication protocol (e.g., HTTPS or WebSocket). The voice data is encrypted and sent to the server in a privacy-protected manner.

[1492] 3. Speech-to-text methods

[1493] The server converts the transmitted voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The server sends the audio file to the speech recognition API and receives the returned text data.

[1494] 4. Methods for analyzing text data

[1495] The server inputs the converted text data into a generative AI model (e.g., OpenAI's ChatGPT) for analysis. The following prompt is used during analysis:

[1496] Analyze whether the following text is an inappropriate statement about pregnancy or morning sickness, and if so, output a score along with the reason.

[1497] Text: "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1498] The generative AI model analyzes the given text and returns a score and reason for detecting inappropriate behavior.

[1499] 5. Measures to detect inappropriate behavior

[1500] The server performs a scoring process based on the analysis results returned by the generative AI model. If the score exceeds a certain threshold, the remark is deemed to constitute inappropriate behavior (maternity harassment).

[1501] 6. How to Generate an Alert Message

[1502] If inappropriate behavior is detected, the server generates an alert message containing the specific content of the comment and a warning, such as "This comment contains comments about pregnancy and morning sickness. Please be considerate in your comments and behavior."

[1503] 7. Alert notification method

[1504] The terminal receives alert messages sent from the server in real time and displays them on the user interface. Alert messages are notified visually (banners and pop-up notifications) and audibly (voice alerts and chimes) so that users can quickly acknowledge them.

[1505] Specific examples

[1506] Conversation scenario

[1507] 1. Collecting comments

[1508] The user (boss) says, "Mr. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1509] The terminal collects this speech via a microphone and transmits the audio data to a server.

[1510] 2. Speech-to-text conversion and analysis

[1511] The server uses a speech recognition API to convert the collected voice data into text.

[1512] The converted text is parsed using prompts from a generative AI model.

[1513] 3. Detecting inappropriate behavior and generating alerts

[1514] The server generates an alert message containing specific statements and warnings based on text whose score exceeds the threshold.

[1515] The terminal receives this alert message and notifies the user.

[1516] In this way, the system can detect maternity harassment in the workplace early and prevent inappropriate behavior by immediately notifying users, which is expected to provide a safe and comfortable working environment.

[1517] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1518] Program processing steps and detailed explanations

[1519] Step 1: Collecting audio

[1520] The device (user's smartphone or PC) uses a microphone to collect the user's speech in real time. For example, the user might say, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress." This collected speech data (input) is recorded (output) either uncompressed or in an appropriate compressed format (e.g., AAC or MP3).

[1521] Step 2: Sending audio data

[1522] The device sends the collected voice data to the server using a secure communication protocol (e.g., HTTPS or WebSocket). The voice data (input) is encrypted and transmitted to protect the privacy of the data. The server receives this voice data (output).

[1523] Step 3: Speech to Text

[1524] The server sends the received voice data (input) to a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert it into text data. The speech recognition API analyzes the voice data and returns text data (output). The server receives this converted text data.

[1525] Step 4: Analyzing the text data

[1526] The server inputs the converted text data (input) into a generative AI model (e.g., OpenAI's ChatGPT) for analysis. The server sends the following prompt to the generative AI model:

[1527] Analyze whether the following text is an inappropriate statement about pregnancy or morning sickness, and if so, output a score along with the reason.

[1528] Text: "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1529] The generative AI model analyzes this prompt and returns a score indicating inappropriate behavior and the reason (output) to the server.

[1530] Step 5: Detect inappropriate behavior

[1531] The server performs a score based on the analysis results (input) returned from the generative AI model. If this score exceeds a certain threshold, the comment is deemed to be inappropriate behavior (maternity harassment). Specifically, for example, a score of "85 / 100" is output, and a reason such as "Reason: Because mentioning pregnancy or morning sickness involves personal privacy" is displayed.

[1532] Step 6: Generate an alert message

[1533] If the server detects inappropriate behavior, it generates an alert message (input) that includes the specific content of the comment and a warning. For example, it could generate a message (output) stating, "This comment contains comments about pregnancy and morning sickness. Please be considerate in your speech and behavior."

[1534] Step 7: Alert Notification

[1535] The terminal receives the alert message (input) sent from the server in real time and displays it on the user interface. The terminal notifies the user of the alert message visually (banner or pop-up notification) and audibly (voice alert or chime) so that the user can quickly acknowledge it (output). The user acknowledges the alert message and realizes that their comment was inappropriate.

[1536] (Application example 1)

[1537] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1538] In today's workplaces, inappropriate verbal and physical behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, remains a problem. If inappropriate comments and behavior are left unchecked, it can lead to a worsening work environment and have serious consequences for the mental and physical health of victims. However, conventional methods currently make it difficult to quickly detect and prevent such inappropriate behavior. Therefore, a system is needed that can collect and analyze voice data in real time and quickly detect and notify employees of inappropriate behavior in the workplace.

[1539] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1540] In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data using a generative AI model to detect inappropriate speech and behavior, means for generating an alert message based on the detection result and notifying the terminal of the alert message, means for notifying the alert message in real time via WebSocket, and means for displaying the notified alert message on a user interface. This makes it possible to quickly detect inappropriate speech and behavior in the workplace and respond in real time.

[1541] "Voice data" is data that represents voice in digital form, and is information that is collected in real time and is subject to analysis.

[1542] A "collection means" is a system or process for acquiring audio data using a device such as a microphone.

[1543] A "server" is a remote computer system that receives voice data, analyzes it, generates alerts, and performs other processing.

[1544] The "transmission means" is a mechanism for temporarily storing collected voice data and transferring it to a server for subsequent processing.

[1545] "Text data" is character string information converted from voice data using voice recognition technology.

[1546] "Conversion means" refers to technology or systems for converting voice data into text data, including voice recognition APIs.

[1547] "Analysis Methods" refers to the processes and techniques used to identify inappropriate behavior within text data using generative AI models.

[1548] A "generative AI model" is a machine learning model used to analyze text data and is an algorithm for detecting inappropriate behavior.

[1549] "Inappropriate behavior" refers to inappropriate remarks or actions related to pregnancy, childbirth, or childcare in the workplace.

[1550] An "alert message" is a notification message that alerts the user when inappropriate behavior is detected.

[1551] "Notification means" refers to a system or method for transmitting the generated alert message to a user's terminal in real time.

[1552] "User interface" refers to a screen or application that displays the notified alert message so that the user can check it.

[1553] "WebSocket" is a communication protocol for real-time communication between a server and a client.

[1554] This invention is a system that detects and prevents inappropriate behavior related to pregnancy, childbirth, and childcare, known as maternity harassment, in the workplace at an early stage. This system consists of the following main components for real-time collection and analysis of voice data and notification of alerts:

[1555] System configuration

[1556] 1. Audio collection method

[1557] This system uses a device (e.g., a smartphone, PC, or dedicated microphone device) to collect the user's speech in real time. The voice data collected through the microphone is recorded in uncompressed or compressed format and then sent to a server.

[1558] 2. Means of transmitting audio data

[1559] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is securely transferred using TLS, ensuring privacy.

[1560] 3. Speech-to-text methods

[1561] The server converts the received voice data into text data in real time using a speech recognition API (for example, Google Cloud Speech-to-Text). Using cloud services enables highly accurate speech recognition.

[1562] 4. Methods for analyzing text data

[1563] The server inputs the converted text data into a generative AI model (e.g., ChatGPT) and analyzes the text for inappropriate behavior. This analysis involves comparing the data with a database of past cases of maternity harassment. The generative AI model's analysis also uses prompt sentences to help detect specific behavior.

[1564] 5. Measures to detect inappropriate behavior

[1565] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior contains any inappropriate behavior. If the score exceeds a certain threshold, the speech or behavior is deemed to be highly likely to constitute maternity harassment.

[1566] 6. How to Generate an Alert Message

[1567] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The alert message is created based on the results of detailed analysis.

[1568] 7. Alert notification method

[1569] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified visually and audibly so that the user can quickly recognize it.

[1570] Specific examples

[1571] Conversation scenario

[1572] Collecting statements:

[1573] The user (boss) says, "Ms. B, I've been taking a lot of time off recently due to morning sickness, but please make sure that it doesn't affect my work progress." Such comments are collected by microphones in the workplace.

[1574] Speech-to-text and analysis:

[1575] The server uses a speech recognition API to convert the collected voice data into text, for example, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1576] The converted text is analyzed by a generative AI model, which scores the statement as inappropriate.

[1577] Alert generation and notification:

[1578] Based on the text whose score exceeds the threshold, the server generates an alert message stating, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[1579] The device receives this alert message via WebSocket and notifies the user in real time.

[1580] Example prompt sentence:

[1581] "The manager said, 'I have kids so it's a problem if you're always late.' Does this statement constitute inappropriate behavior?"

[1582] This system is expected to not only quickly detect inappropriate behavior in the workplace, but also enable real-time responses, thereby providing a safe and comfortable working environment.

[1583] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1584] Step 1:

[1585] The device uses a microphone in the workplace to collect voice data in real time. This allows the user's statements and conversation content to be captured in digital form. The input is voice, and the output is digital voice data. For example, if a user says, "Mr. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure it doesn't affect my work progress," the device will capture this as voice data.

[1586] Step 2:

[1587] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The input is digital audio data, and the output is the audio data sent to the server. Communication is securely encrypted using TLS.

[1588] Step 3:

[1589] The server receives the transmitted voice data and converts it into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is the voice data received by the server, and the output is the converted text data. Specifically, voice data such as "Ms. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure that it doesn't affect my work progress" is converted into text data such as "Ms. B, I've been taking a lot of time off work recently because of morning sickness, but please make sure that it doesn't affect my work progress."

[1590] Step 4:

[1591] The server inputs the converted text data into a generative AI model (e.g., ChatGPT) and analyzes the inappropriate behavior in the text. A prompt is used for the analysis. The input is the converted text data, and the output is the analysis result. Specifically, the prompt, "Does this statement constitute inappropriate behavior?", is input into the generative AI model, and the analysis result is obtained.

[1592] Step 5:

[1593] The server calculates a score based on the results of the analysis by the generative AI model to determine whether the speech or behavior is inappropriate. The input is the analysis result of the generative AI model, and the output is a score. If the score exceeds a certain threshold, the speech or behavior is deemed inappropriate.

[1594] Step 6:

[1595] If maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. The input is the analysis result where the score exceeds the threshold, and the output is the alert message. Specifically, the alert message generated reads, "The remarks include comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[1596] Step 7:

[1597] The terminal receives alert messages sent from the server in real time via WebSocket and displays them on the user interface. The input is the alert message, and the output is the notification displayed on the user interface. Specifically, the user will be notified visually and / or audibly.

[1598] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1599] ---

[1600] This invention is a system that detects inappropriate behavior and emotions related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevents them from occurring. This system consists of the following main components for collecting and analyzing voice data, recognizing emotions, and issuing alerts.

[1601] System configuration

[1602] 1. Audio collection method

[1603] The device (e.g., a smartphone or PC) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[1604] 2. Means of transmitting audio data

[1605] The device sends the collected audio data to the server using an appropriate communication protocol (e.g., WebSocket or HTTP / 2). The audio data is transmitted in a secure manner, ensuring privacy.

[1606] 3. Speech-to-text methods

[1607] The server converts the received voice data into text data using a voice recognition API, which uses a cloud service capable of highly accurate recognition (for example, a voice recognition cloud service).

[1608] 4. Text data and sentiment analysis methods

[1609] The server inputs the converted text and voice data into a generative AI model and emotion engine, analyzes inappropriate behavior in the text, and recognizes the user's emotional state. The analysis is performed by comparing it with a database of past cases of maternity harassment.

[1610] 5. Measures to detect inappropriate behavior and emotions

[1611] The server calculates a score based on the results of analysis by the generative AI model and emotion engine to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed to be highly likely to constitute maternity harassment, and the user's emotional state is also used in the next process.

[1612] 6. How to Generate an Alert Message

[1613] When maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. This alert message will be used to appropriately notify the perpetrator and the administrator. The content of the alert message will also be adjusted based on the user's emotional state.

[1614] 7. Alert notification method

[1615] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[1616] Specific examples

[1617] Conversation scenario

[1618] 1. Collecting comments

[1619] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1620] The device collects this speech through a microphone and transmits the data to a server.

[1621] 2. Speech-to-text conversion and analysis

[1622] The server uses a speech recognition API to convert the collected voice data into text.

[1623] The converted text and speech data are analyzed by a generative AI model and emotion engine to score inappropriate language in the text and the user's emotional state.

[1624] 3. Alert generation and notification

[1625] Based on the text whose score exceeds the threshold, the server generates an alert message saying, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[1626] Depending on the user's emotional state, for example, if the user is emotionally charged, a supplemental message such as "Let's stay calm and continue the conversation" may be added.

[1627] The terminal receives this alert message and notifies the user.

[1628] 4. User Behavior

[1629] The user will receive an alert message, realizing that their remarks were inappropriate, and will also receive feedback about their emotional state.

[1630] As a result, users will reconsider what they have said and be more considerate in future conversations, while also paying attention to controlling their emotions.

[1631] ---

[1632] In this way, the system can detect maternity harassment and related emotional upheavals in the workplace at an early stage and prevent inappropriate behavior before it occurs, which is expected to provide a safe and comfortable working environment.

[1633] The processing flow will be explained below.

[1634] ---

[1635] Step 1:

[1636] Audio collection

[1637] The device uses a microphone to capture the user's conversation in real time, and the audio data is recorded in uncompressed or compressed format and then transmitted to a server.

[1638] Step 2:

[1639] Audio data compression and buffering

[1640] The device converts the collected audio data into a compressed format for efficient transmission to the server. The compressed audio data is divided into small data chunks and stored in a buffer for transmission.

[1641] Step 3:

[1642] Sending audio data

[1643] The device transmits the audio data stored in the buffer to the server at regular intervals. This transmission uses a low-latency communication protocol (e.g., WebSocket or HTTP / 2). The transmission is performed in real time and is monitored to ensure there is no interruption of data.

[1644] Step 4:

[1645] Receiving and decoding audio data

[1646] The server receives the audio data sent from the device in real time, decodes the compressed audio data, and stores it in temporary storage before passing it on to the next processing step.

[1647] Step 5:

[1648] Speech recognition to text

[1649] The server sends the voice data stored in the temporary storage to a voice recognition API, which converts it into text data. The voice recognition API provides highly accurate voice recognition, and the converted text data is stored in a database.

[1650] Step 6:

[1651] Preprocessing text data

[1652] The server cleanses the text data received from the speech recognition API, removing mistranslations and noise, and prepares the cleansed text data for analysis.

[1653] Step 7:

[1654] Analysis using generative AI models and emotion engines

[1655] The server inputs the cleansed text data into a generative AI model and an emotion engine, which analyzes the text for inappropriate behavior and the user's emotional state. The generative AI model detects inappropriate behavior with high accuracy based on a large amount of training data, and the emotion engine estimates the user's emotional state from the voice data.

[1656] Step 8:

[1657] Detecting inappropriate behavior and sentiment

[1658] Based on the analysis results, the server determines whether the statements contained in the text data constitute inappropriate behavior and simultaneously scores the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed inappropriate and the user's emotional state is also recorded.

[1659] Step 9:

[1660] Generate an alert message

[1661] When inappropriate behavior is detected, the server generates an alert message containing specific statements and a warning, and adjusts the content and tone of the alert message based on the user's emotional state and includes additional feedback as needed.

[1662] Step 10:

[1663] Sending an alert message

[1664] The server then sends the generated alert message to the target device. The alert message is distributed in real time and, if necessary, notifies the administrator or harassment prevention department.

[1665] Step 11:

[1666] Receiving and Viewing Alerts

[1667] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[1668] Step 12:

[1669] User Behavior

[1670] The user checks the alert message displayed on their device and realizes that their remarks were inappropriate. They also receive feedback on their emotional state. Based on this recognition, the user can be more considerate in future conversations and pay more attention to controlling their emotions.

[1671] ---

[1672] Through this step, the system of the present invention can detect inappropriate behavior and heightened emotions related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevent inappropriate behavior before it occurs, which is expected to provide a safe and comfortable working environment.

[1673] Example 2

[1674] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1675] In today's work environment, inappropriate behavior and attitudes related to pregnancy, childbirth, and childcare remain a problem. These issues can particularly increase stress and psychological burden on female employees in the workplace. Furthermore, managers often lack the means to identify and address these issues early before they become serious. As a result, maintaining a safe and comfortable work environment is difficult. This invention aims to improve the work environment and reduce the psychological burden on employees by detecting inappropriate behavior and the user's emotional state in the workplace early and responding promptly.

[1676] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1677] In this invention, the server includes a means for converting voice data into text data, a means for analyzing the converted text data and voice data to recognize the user's emotional state, and a means for determining inappropriate behavior and the user's emotional state based on the analysis results, thereby enabling prompt and accurate identification and response to inappropriate behavior related to pregnancy, childbirth, and childcare in the workplace.

[1678] "Audio data" means digital audio signals collected through an audio input device.

[1679] "Text data" refers to text information converted from voice data using a voice recognition API.

[1680] A "generative AI model" is an algorithm or machine learning model that uses artificial intelligence to analyze text data and detect inappropriate behavior.

[1681] An "emotion engine" is software for recognizing a user's emotional state from text data and voice tone.

[1682] An "alert message" is a notification message that is generated to encourage the user to behave considerately when inappropriate behavior or the user's emotional state is determined.

[1683] "Audio input device" refers to a microphone or other audio collection device for capturing speech in real time.

[1684] A "secure transfer protocol" is a communication technology for encrypting and transmitting data securely, and specifically refers to protocols such as HTTPS and WebSocket Secure.

[1685] A "prompt sentence" is the input text presented to a generative AI model, and is the sentence that serves as the starting point for analysis and generation.

[1686] This invention is a system that detects inappropriate behavior and the user's emotional state related to pregnancy, childbirth, and childcare in the workplace at an early stage and prevents such behavior before it occurs. This system consists of the following main elements for collecting and analyzing voice data, recognizing emotions, and issuing alerts.

[1687] System configuration

[1688] Audio collection method

[1689] A device (e.g., a smartphone or a personal computer) uses a microphone to collect the user's speech in real time, which is then recorded in uncompressed or compressed format and transmitted to a server.

[1690] Voice data transmission method

[1691] The device sends the collected voice data to the server using an appropriate communication protocol (e.g., WebSocket Secure or HTTPS). The voice data is transmitted securely, ensuring privacy.

[1692] Voice-to-text conversion methods

[1693] The server converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text).

[1694] Text data and sentiment analysis methods

[1695] The server inputs the converted text and voice data into a generative AI model (e.g., GPT-4) and an emotion engine to analyze inappropriate behavior in the text and recognize the user's emotional state. The analysis is performed by comparing the data with a database of past cases of maternity harassment.

[1696] Measures to detect inappropriate behavior and emotions

[1697] The server calculates a score based on the results of analysis by the generative AI model and emotion engine to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed to be highly likely to constitute maternity harassment, and the user's emotional state is also used in the next process.

[1698] How to generate an alert message

[1699] When maternity harassment is detected, the server generates an alert message containing the specific content of the remarks and a warning. This alert message will be used to appropriately notify the perpetrator and the administrator. The content of the alert message will also be adjusted based on the user's emotional state.

[1700] Alert notification method

[1701] The terminal receives the alert message sent from the server in real time and displays it on the user interface. The alert message is notified by visual and audible means so that the user can quickly recognize it.

[1702] Specific examples

[1703] Conversation scenario

[1704] 1. Collecting comments

[1705] The user (boss) says, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1706] The device collects this speech through a microphone and transmits the data to a server.

[1707] 2. Speech-to-text conversion and analysis

[1708] The server uses a speech recognition API to convert the collected voice data into text.

[1709] The converted text and voice data are analyzed by a generative AI model (GPT-4) and an emotion engine to score inappropriate language in the text and the user's emotional state.

[1710] 3. Alert generation and notification

[1711] Based on the text whose score exceeds the threshold, the server generates an alert message stating, "This text contains comments about pregnancy and morning sickness. Please be considerate in your words and actions."

[1712] Depending on the user's emotional state, for example, if the user is emotionally charged, a supplemental message such as "Let's stay calm and continue the conversation" is added.

[1713] The terminal receives this alert message and notifies the user.

[1714] 4. User Behavior

[1715] The user will receive an alert message, realizing that their remarks were inappropriate, and will also receive feedback about their emotional state.

[1716] As a result, users will reconsider what they have said and be more considerate in future conversations, while also paying attention to controlling their emotions.

[1717] Prompt Sentence Examples

[1718] "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1719] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1720] Step 1:

[1721] Audio collection

[1722] Input: User's speech

[1723] Specific operation: A user is having a conversation at work. The device (e.g., a smartphone or a personal computer) activates a built-in or external microphone to collect the conversation audio in real time.

[1724] Output: Audio data (digital format)

[1725] Step 2:

[1726] Sending audio data

[1727] Input: Audio data

[1728] Specific operation: The device encrypts the collected voice data using SSL / TLS and sends it to the server's API endpoint using a secure communication protocol (e.g., HTTPS or WebSocket Secure).

[1729] Output: Audio data sent to the server

[1730] Step 3:

[1731] Speech to text

[1732] Input: Audio data sent to the server

[1733] Specific operation: The server passes the received voice data to a speech recognition API (e.g., Google Cloud Speech-to-Text) and converts it into text data. The server receives the text data returned by the API.

[1734] Output: Converted text data

[1735] Step 4:

[1736] Text data and sentiment analysis

[1737] Input: Converted text and audio data

[1738] Specific operation: The server analyzes the text data entered as prompts into a generative AI model (e.g., GPT-4) and identifies inappropriate behavior within the text. In addition, it uses an emotion engine to recognize the user's emotional state from the tone of voice and text data. The analysis results are compared with a database of past maternity harassment cases.

[1739] Output: Analysis results and user emotional state data

[1740] Step 5:

[1741] Detecting inappropriate behavior and sentiment

[1742] Input: Analysis results and user emotional state data

[1743] Specific operation: The server calculates a score for inappropriate behavior and the user's emotional state based on the analysis results from the generative AI model and emotion engine. The score is compared with a pre-set threshold, and behavior that exceeds the threshold is deemed inappropriate.

[1744] Output: Judgment result of inappropriate behavior and emotional state

[1745] Step 6:

[1746] Generate an alert message

[1747] Input: Inappropriate behavior and emotional state assessment results

[1748] Specific actions: Based on the judgment result, the server generates an alert message containing the specific content of the comment and a warning. For example, it creates a message such as, "The comment contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." Based on the user's emotional state, it also adds a supplementary message such as, "Please remain calm and continue the conversation."

[1749] Output: The generated alert message

[1750] Step 7:

[1751] Alert Notification

[1752] Input: The generated alert message

[1753] Specific operation: The terminal receives the alert message sent from the server in real time and displays it on the user interface, specifically notifying the user by means of a pop-up notification, a sound notification, or the like.

[1754] Output: Alert message sent to the user

[1755] Specific examples

[1756] 1. User statement: The user (boss) says, "Ms. B, I've been taking a lot of time off recently due to morning sickness, but please make sure it doesn't affect my work progress."

[1757] 2. What the device does: The built-in microphone collects this speech, encrypts the data, and sends it to a server.

[1758] 3. Server operation: The speech is converted into text using a speech recognition API, and then analyzed using a generative AI model (GPT-4) and an emotion engine.

[1759] 4. Analysis results: Inappropriate behavior is detected and the user is determined to be emotionally charged.

[1760] 5. Alert generation: An alert is generated stating, "This post contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." along with a supplementary message saying, "Please remain calm and continue the discussion."

[1761] 6. Device notification: Alert messages are notified to the user via pop-up and sound.

[1762] 7. User action: The user checks the notification, reviews their comments, and controls their emotions.

[1763] (Application example 2)

[1764] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1765] Inappropriate verbal and physical behavior and the emotional state of employees related to pregnancy, childbirth, and childcare often become problems in the workplace. However, it is difficult to detect these inappropriate verbal and physical behavior and emotional outbursts early and prevent them from occurring. In particular, achieving this in real time has been difficult with conventional technology. The objective of this invention is to provide a system that quickly detects inappropriate verbal and physical behavior and emotional outbursts in the workplace and provides a safe and comfortable working environment.

[1766] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data to the server, means for converting the transmitted voice data into text data, means for analyzing the converted text data and detecting inappropriate speech and behavior and the user's emotional state, means for generating an alert message based on the detection result and the emotional state and notifying the terminal of the alert message, and means for displaying the notified alert message on a user interface. This makes it possible to detect inappropriate speech and behavior and heightened emotions in the workplace in real time and prevent them from occurring.

[1767] "Audio data" refers to electronically recorded data of a user's conversation or environmental sounds.

[1768] "Collection means" refers to devices and software for collecting voice data in real time.

[1769] "Transmission means" refers to the communication method or protocol used to send collected voice data to the server.

[1770] "Text data" refers to data obtained by converting voice data into text information using voice recognition technology.

[1771] "Conversion means" refers to the technology or device used to convert voice data into text data.

[1772] "Analysis means" refers to technologies and algorithms used to analyze the converted text data and detect inappropriate behavior or emotional states.

[1773] A "generative AI model" is an AI model trained using machine learning or deep learning, and is used to analyze text data.

[1774] "User's emotional state" refers to the emotional state that the user expresses through their words and actions, such as joy, anger, or anxiety.

[1775] An "alert message" is a warning message generated based on detected inappropriate behavior or emotional state.

[1776] "Notification means" refers to the communication method or protocol used to communicate the generated alert message to the user or administrator.

[1777] "User interface" refers to an interface for visually or audibly displaying an alert message to a user.

[1778] A "prompt sentence" is an initial sentence or instruction input to a generative AI model, and is used to indicate the direction of analysis.

[1779] This invention is a system that detects and prevents inappropriate behavior and emotional upset related to pregnancy, childbirth, and childcare in the workplace at an early stage. This system is mainly composed of components for collecting, transmitting, converting, and analyzing voice data, detecting inappropriate behavior, and generating and notifying alert messages.

[1780] System configuration

[1781] 1. Audio collection method

[1782] The device (e.g., a smartphone or PC) uses a microphone to collect the user's conversational voice in real time, which is then recorded in uncompressed or compressed format and sent to a server.

[1783] 2. Means of transmitting audio data

[1784] The device sends the collected audio data to the server using an appropriate communication protocol (such as HTTP / 2). The audio data is transmitted securely to protect privacy.

[1785] 3. Speech-to-text methods

[1786] The server converts the received voice data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text), a cloud service that provides highly accurate recognition.

[1787] 4. Text data and sentiment analysis methods

[1788] The server then inputs the converted text and audio data into a generative AI model and a sentiment analysis engine (e.g., TextBlob) to analyze the text for inappropriate language and the user's emotional state. The analysis is based on the prompt sentence and identifies inappropriate phrases and heightened emotions.

[1789] As an example, use the following prompt:

[1790] "Turn this audio data into text and look for inappropriate language and heightened emotions. Pay attention to specific phrases and pick out negative comments."

[1791] 5. Measures to detect inappropriate behavior and emotions

[1792] The server uses the generative AI model and sentiment analysis engine to calculate a score from the analyzed text data to determine inappropriate behavior and the user's emotional state. If the score exceeds a certain threshold, the behavior is deemed inappropriate.

[1793] 6. How to Generate an Alert Message

[1794] The server generates an alert message when inappropriate behavior and emotional upset are detected, which includes specific statements and reminders and is tailored to the user's emotional state.

[1795] 7. Alert notification method

[1796] The terminal receives alert messages sent from the server in real time and displays them on the user interface. The alert messages are notified to the user visually and audibly so that they can be quickly recognized.

[1797] Specific examples

[1798] 1. Collecting comments

[1799] The terminal collects the user's (boss') remarks through a microphone.

[1800] For example, a statement such as, "Ms. B, I've been taking a lot of time off work recently due to morning sickness, but please make sure it doesn't affect my work progress" is collected.

[1801] 2. Speech-to-text conversion and analysis

[1802] The server uses a speech recognition API to convert the collected voice data into text.

[1803] The converted text is analyzed using a generative AI model and a sentiment analysis engine to score inappropriate behavior and emotional states.

[1804] 3. Alert generation and notification

[1805] The server detects inappropriate comments and generates an alert message saying, "This comment contains comments about pregnancy and morning sickness. Please be considerate in your words and actions." If emotions are running high, a supplemental message is added saying, "Please remain calm and continue the conversation."

[1806] The terminal receives this alert message in real time and notifies the user.

[1807] 4. User Behavior

[1808] The user sees the alert message, realizes that what they said was inappropriate, and also receives feedback about their emotional state.

[1809] This will encourage users to review their future comments and be more considerate in their speech and actions.

[1810] This system will enable early detection of inappropriate behavior and emotional outbursts in the workplace, providing a safe and comfortable working environment.

[1811] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1812] ---

[1813] Step 1:

[1814] The device uses a microphone to collect conversational audio in real time. The input is the user's live voice, and the output is stored on the device as audio data. This audio data can be recorded in uncompressed or compressed format.

[1815] Step 2:

[1816] The device sends the collected audio data to the server using a secure communication protocol such as HTTP / 2. The input is audio data, and the output is audio data sent to the server. The communication is encrypted to protect privacy.

[1817] Step 3:

[1818] The server converts the received voice data into text data using a speech recognition API (Google Cloud Speech-to-Text). The input here is voice data, and the output is text data. This text data is obtained using speech recognition technology.

[1819] Step 4:

[1820] The server inputs the converted text data into a generative AI model and an emotion analysis engine (TextBlob) to analyze the inappropriate behavior and emotional state of the user in the text. The input here is text data, and the output is a score for the inappropriate behavior and emotional state. The generative AI model performs the analysis using prompts. For example, a prompt might be, "Convert this audio data into text and detect inappropriate behavior and heightened emotions. Pay attention to specific phrases and pick out negative comments."

[1821] Step 5:

[1822] The server uses the generative AI model and emotion analysis engine to analyze the results and determine whether the behavior and emotional state are inappropriate. If the score exceeds a certain threshold, the behavior is deemed inappropriate. The input here is the analysis score, and the output is the judgment of inappropriate behavior.

[1823] Step 6:

[1824] The server generates an alert message when inappropriate behavior or heightened emotions are detected. The input is the judgment result of inappropriate behavior, and the output is an alert message. The alert message includes the specific content of the statement and a warning, and if necessary, a supplemental message based on the emotional state is added.

[1825] Step 7:

[1826] The terminal receives alert messages sent from the server in real time and displays them on the user interface. The input here is the alert message, and the output is a notification to the user. The alert message is notified visually and audibly so that the user can quickly recognize it.

[1827] ---

[1828] The above processing steps make it possible to detect inappropriate behavior and emotional outbursts in the workplace in real time and prevent them from occurring.

[1829] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1830] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1831] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1832] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1833] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1834] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1835] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1836] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1837] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1838] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1839] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1840] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1841] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1842] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1843] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1844] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1845] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1846] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1847] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1848] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1849] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1850] The following is further disclosed regarding the above embodiment.

[1851] ---

[1852] (Claim 1)

[1853] means for collecting audio data;

[1854] means for transmitting the collected voice data to a server;

[1855] means for converting the transmitted voice data into text data;

[1856] A means for analyzing the converted text data and detecting inappropriate speech and behavior;

[1857] means for generating an alert message based on the detection result and notifying the terminal of the alert message;

[1858] The system includes means for displaying the notified alert message on a user interface.

[1859] (Claim 2)

[1860] The system according to claim 1, characterized in that the analysis means analyzes text data using a generative AI model to detect inappropriate speech and behavior.

[1861] (Claim 3)

[1862] 2. The system of claim 1, wherein the collecting means includes a microphone for collecting speech in real time.

[1863] "Example 1"

[1864] (Claim 1)

[1865] means for collecting audio data;

[1866] means for transmitting the collected voice data to a server;

[1867] means for converting the transmitted voice data into text data;

[1868] A means for analyzing the converted text data using a generative AI model to detect inappropriate speech and behavior;

[1869] means for generating an alert message based on the detection results;

[1870] means for notifying a terminal of the generated alert message;

[1871] The system includes means for displaying the notified alert message on a user interface.

[1872] (Claim 2)

[1873] The system described in claim 1, characterized in that it analyzes text data using a generative AI model and detects inappropriate behavior.

[1874] (Claim 3)

[1875] 10. The system of claim 1, further comprising a microphone for collecting speech in real time.

[1876] "Application Example 1"

[1877] (Claim 1)

[1878] means for collecting audio data;

[1879] means for transmitting the collected voice data to a server;

[1880] means for converting the transmitted voice data into text data;

[1881] A means for analyzing the converted text data and detecting inappropriate speech and behavior;

[1882] means for generating an alert message based on the detection result and notifying the terminal of the alert message;

[1883] means for displaying the notified alert message on a user interface;

[1884] A system that includes a means to notify alert messages in real time via WebSocket.

[1885] (Claim 2)

[1886] The system according to claim 1, characterized in that the analysis means analyzes text data using a generative AI model to detect inappropriate speech and behavior.

[1887] (Claim 3)

[1888] 2. The system of claim 1, wherein the collecting means includes a microphone for collecting speech in real time.

[1889] "Example 2: Combining Emotion Engines"

[1890] (Claim 1)

[1891] means for collecting audio data;

[1892] means for transmitting the collected voice data to a server;

[1893] means for converting the transmitted voice data into text data;

[1894] means for analyzing the converted text data and voice data to recognize the emotional state of the user;

[1895] means for determining inappropriate behavior and the emotional state of the user based on the analysis results;

[1896] means for generating an alert message based on the determination result and notifying the terminal of the alert message;

[1897] means for displaying the notified alert message on a user interface;

[1898] A system including:

[1899] (Claim 2)

[1900] The system according to claim 1, characterized in that the analysis means analyzes text data and voice data using a generative AI model to determine inappropriate behavior and the user's emotional state.

[1901] (Claim 3)

[1902] 2. The system according to claim 1, wherein the collecting means includes a voice input device for collecting conversational voice in real time.

[1903] "Application example 2 when combining emotion engines"

[1904] ---

[1905] (Claim 1)

[1906] means for collecting audio data;

[1907] means for transmitting the collected voice data to a server;

[1908] means for converting the transmitted voice data into text data;

[1909] A means for analyzing the converted text data and detecting inappropriate speech and behavior;

[1910] means for recognizing and analyzing the emotional state of a user;

[1911] means for generating an alert message based on the detection result and the emotional state and notifying the terminal of the alert message;

[1912] The system includes means for displaying the notified alert message on a user interface.

[1913] (Claim 2)

[1914] The system according to claim 1, characterized in that the analysis means analyzes text data using a generative AI model to detect inappropriate behavior and the user's emotional state.

[1915] (Claim 3)

[1916] 2. The system according to claim 1, wherein the collecting means includes a microphone for collecting conversational voice in real time and performs analysis based on a prompt sentence. [Explanation of symbols]

[1917] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for collecting audio data; means for transmitting the collected voice data to a server; means for converting the transmitted voice data into text data; A means for analyzing the converted text data and detecting inappropriate speech and behavior; means for generating an alert message based on the detection result and notifying the terminal of the alert message; The system includes means for displaying the notified alert message on a user interface.

2. The system according to claim 1, wherein the analysis means analyzes text data using a generative AI model to detect inappropriate speech and behavior.

3. 2. The system of claim 1, wherein said collecting means includes a microphone for collecting speech in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A