system

The system addresses the challenge of real-time harassment detection by using AI to analyze conversation content, send alerts, and learn from user registrations, enhancing the effectiveness of harassment prevention.

JP2026073613APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing systems struggle to detect harassing utterances in real-time and prompt the speaker effectively.

Method used

A system comprising an analysis unit, detection unit, notification unit, and learning unit that uses AI to analyze conversation content, detect potentially harassing remarks, send alerts, and improve detection accuracy through user registration of offensive remarks.

Benefits of technology

Enables real-time detection and alerting of harassing remarks, encouraging perpetrators to recognize their behavior and improving detection accuracy over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073613000001_ABST
    Figure 2026073613000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to detect harassing remarks in real time and alert the person making the remark. [Solution] The system according to the embodiment comprises an analysis unit, a detection unit, a notification unit, a registration unit, and a learning unit. The analysis unit analyzes conversation content in real time. The detection unit detects potentially harassing remarks from the conversation content analyzed by the analysis unit. The notification unit sends a message to the speaker to warn them about the remarks detected by the detection unit. The registration unit registers remarks that are found offensive in everyday life. The learning unit learns the remarks registered by the registration unit and improves detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, it is difficult to detect harassing utterances in real time and prompt the speaker, and there is room for improvement.

[0005] The system according to the embodiment aims to detect harassing utterances in real time and prompt the speaker.

Means for Solving the Problems

[0006] The system according to this embodiment comprises an analysis unit, a detection unit, a notification unit, a registration unit, and a learning unit. The analysis unit analyzes conversation content in real time. The detection unit detects potentially harassing remarks from the conversation content analyzed by the analysis unit. The notification unit sends a message to the speaker to warn them about the remarks detected by the detection unit. The registration unit registers remarks that are found offensive in everyday life. The learning unit learns the remarks registered by the registration unit and improves detection accuracy. [Effects of the Invention]

[0007] The system according to this embodiment can detect harassing remarks in real time and alert the person making the remark. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10]This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The harassment prevention system according to an embodiment of the present invention is a system that encourages perpetrators of harassment to recognize their own behavior and review their words and actions. The harassment prevention system detects potentially harassing remarks and alerts the speaker, and when an employee receives a remark they find offensive in their daily life, the content of that remark is registered, allowing the AI ​​to learn and improve detection accuracy. This contributes to creating a workplace environment where all employees can work with peace of mind. For example, the harassment prevention system uses AI to analyze conversation content in real time and detect potentially harassing remarks. For example, during a meeting or chat exchange, the AI ​​determines whether specific words or phrases constitute harassment. Next, the harassment prevention system automatically sends an alert message to the speaker regarding the detected remark. This gives the speaker an opportunity to review their words and actions. Furthermore, when an employee receives a remark they find offensive in their daily life, they register the content of that remark in the system. The registered remarks are learned by the AI ​​and used to improve future detection accuracy. For example, by repeatedly registering specific words or phrases, the AI ​​can determine that these words or phrases are highly likely to constitute harassment, enabling faster and more accurate detection. This mechanism encourages perpetrators of harassment to recognize their own behavior and reconsider their actions. Furthermore, by registering remarks that victims find offensive, the AI ​​learns and improves detection accuracy. This contributes to creating a workplace environment where all employees can work with peace of mind. For instance, if someone says "You're always late" during a meeting, the AI ​​will detect the remark and send a message to the speaker warning them that "That remark may be harassment." Also, if a victim registers with the system that "I found my boss's remark 'You're useless' offensive," the AI ​​learns from that remark and can quickly detect similar remarks in the future. In this way, the harassment prevention system encourages perpetrators of harassment to recognize their own behavior and reconsider their actions, while also providing a mechanism for the AI ​​to learn from and improve detection accuracy through the registration of remarks that victims find offensive. This contributes to creating a workplace environment where all employees can work with peace of mind.This allows harassment prevention systems to encourage perpetrators of harassment to recognize their actions and reconsider their behavior, thereby contributing to the improvement of the workplace environment.

[0029] The harassment prevention system according to this embodiment comprises an analysis unit, a detection unit, a notification unit, a registration unit, and a learning unit. The analysis unit analyzes conversation content in real time. The analysis unit uses AI to determine, for example, whether specific words or phrases constitute harassment in a meeting or chat exchange. The analysis unit can analyze conversation content using, for example, natural language processing technology to determine whether specific words or phrases constitute harassment. The analysis unit can also use a machine learning model to analyze conversation content and detect potentially harassing remarks. Furthermore, the analysis unit can use a rule-based system to determine whether specific words or phrases constitute harassment. The detection unit detects potentially harassing remarks from the conversation content analyzed by the analysis unit. The detection unit can, for example, use natural language processing technology to detect whether specific words or phrases constitute harassment from the conversation content analyzed by the analysis unit. The detection unit can also use a machine learning model to detect potentially harassing remarks from the conversation content analyzed by the analysis unit. Furthermore, the detection unit can use a rule-based system to detect whether specific words or phrases constitute harassment from the conversation content analyzed by the analysis unit. The notification unit sends a message to the speaker to alert them to the statement detected by the detection unit. For example, the notification unit can automatically send a message to the speaker informing them that "that statement may be harassment" in response to a statement detected by the detection unit. The notification unit can also send a voice message to the speaker in response to a statement detected by the detection unit. Furthermore, the notification unit can send a text message to the speaker in response to a statement detected by the detection unit. The registration unit registers statements that are found offensive in everyday life. For example, the registration unit can register statements that victims find offensive in the system. The registration unit can also register statements that victims find offensive in voice. Furthermore, the registration unit can register statements that victims find offensive in text messages.The learning unit learns the content of statements registered by the registration unit to improve detection accuracy. For example, the learning unit can improve detection accuracy by inputting the content of statements registered by the registration unit into a machine learning model and learning from it. The learning unit can also improve detection accuracy by learning the content of statements registered by the registration unit using natural language processing technology. Furthermore, the learning unit can improve detection accuracy by inputting the content of statements registered by the registration unit into a rule-based system and learning from it. As a result, the harassment prevention system according to this embodiment can encourage perpetrators of harassment to recognize their own behavior and review their words and actions, thereby contributing to the improvement of the workplace environment.

[0030] The analysis unit analyzes conversation content in real time. For example, in meetings or chat exchanges, the analysis unit uses AI to determine whether specific words or phrases constitute harassment. Specifically, the analysis unit uses natural language processing technology to analyze the context and emotions of the conversation and determine whether specific words or phrases constitute harassment. Natural language processing technology includes morphological analysis, grammatical analysis, and sentiment analysis, which are combined to analyze the conversation content in detail. For example, morphological analysis is used to break down words in a sentence, grammatical analysis is used to understand the structure of the sentence, and sentiment analysis is used to infer the speaker's emotions. Furthermore, the analysis unit uses machine learning models to detect potentially harassing remarks based on past data. Machine learning models have the ability to learn from large amounts of conversation data and identify patterns that constitute harassment. This allows the analysis unit to detect new signs of harassment. It can also use a rule-based system to determine whether specific words or phrases constitute harassment. The rule-based system evaluates the conversation content based on predefined rules and identifies potentially harassing remarks. This allows the analysis unit to analyze conversation content from multiple perspectives, enabling early detection of harassment.

[0031] The detection unit detects potentially harassing remarks from the conversation content analyzed by the analysis unit. For example, the detection unit can use natural language processing techniques to detect whether specific words or phrases constitute harassment from the conversation content analyzed by the analysis unit. Specifically, the detection unit identifies potentially harassing remarks based on the analysis results provided by the analysis unit. Natural language processing techniques include text classification, keyword extraction, and sentiment analysis, which are combined to detect the content of remarks in detail. For example, text classification is used to categorize remarks, keyword extraction is used to extract specific words or phrases, and sentiment analysis is used to evaluate the speaker's emotions. Furthermore, the detection unit can also use machine learning models to detect potentially harassing remarks from the conversation content analyzed by the analysis unit. Machine learning models have the ability to learn from large amounts of conversation data and identify patterns that constitute harassment. This allows the detection unit to detect new signs of harassment. Additionally, rule-based systems can be used to detect whether specific words or phrases constitute harassment from the conversation content analyzed by the analysis unit. The rule-based system evaluates conversation content based on predefined rules and identifies potentially harassing remarks. This allows the detection unit to detect conversation content from multiple angles, enabling early detection of harassment.

[0032] The notification unit sends a message to the speaker to alert them to the statement detected by the detection unit. Specifically, the notification unit can automatically send a message to the speaker informing them that "that statement may be harassment" in response to a statement detected by the detection unit. The notification unit can, for example, send a message to the speaker's device in real time to alert them. Message delivery methods include text messages, voice messages, and pop-up notifications. Text messages are sent directly to the speaker's device, allowing the speaker to see the message. Voice messages provide an audible alert from the speaker's device, allowing the speaker to respond immediately. Pop-up notifications appear in a pop-up format on the speaker's device, allowing the speaker to see the alert message. Furthermore, the notification unit can send alert messages not only to the speaker but also to managers and supervisors. This allows managers and supervisors to understand the situation and take appropriate action. The notification unit can also record the speaker's response and provide data for later analysis. This allows the notification unit to alert the speaker quickly and appropriately, preventing the recurrence of harassment.

[0033] The registration unit registers statements that users find offensive in their daily lives. Specifically, the registration unit allows victims to register statements they find offensive in the system. Registration methods include text input, voice input, and sending text messages. Victims can use text input to describe offensive statements in detail and register them in the system. When using voice input, victims can record offensive statements aloud and register them in the system. When using text messages, victims can send offensive statements to the system as text messages and register them. Furthermore, the registration unit classifies the statements registered by victims and stores them in a database for use by the analysis and detection units. The registered statements are used as training data for the analysis and detection units to improve the accuracy of harassment detection. This allows the registration unit to efficiently register statements that victims find offensive and improve the overall harassment detection capability of the system. In addition, the registration unit can anonymize statements registered by victims to protect their privacy. This allows victims to register offensive statements with peace of mind.

[0034] The learning unit improves detection accuracy by learning from the utterances registered by the registration unit. Specifically, the learning unit can improve detection accuracy by inputting the utterances registered by the registration unit into a machine learning model and learning from it. The machine learning model has the ability to learn from large amounts of data and identify patterns that constitute harassment. This allows the learning unit to detect new signs of harassment. The learning unit can also improve detection accuracy by learning from the utterances registered by the registration unit using natural language processing techniques. Natural language processing techniques include morphological analysis, grammatical analysis, and sentiment analysis, and these are combined to analyze and learn from the utterances in detail. Furthermore, the learning unit can improve detection accuracy by inputting the utterances registered by the registration unit into a rule-based system and learning from it. The rule-based system evaluates the utterances based on predefined rules and identifies utterances that may constitute harassment. This allows the learning unit to learn from utterances using a multifaceted approach and improve the accuracy of harassment detection. In addition, the learning unit can maintain detection accuracy by periodically retraining the model based on the latest data. This allows the learning unit to always perform highly accurate harassment detection based on the latest information, improving the reliability and security of the entire system.

[0035] The analysis unit can use AI to determine whether specific words or phrases in meetings or chat interactions constitute harassment. For example, the analysis unit can analyze conversation content using natural language processing technology to determine whether specific words or phrases constitute harassment. Furthermore, the analysis unit can use machine learning models to determine whether specific words or phrases in meetings or chat interactions constitute harassment. In addition, the analysis unit can use rule-based systems to determine whether specific words or phrases constitute harassment. This enables real-time harassment detection by using AI to determine whether specific words or phrases in meetings or chat interactions constitute harassment. Specific words and phrases include, but are not limited to, discriminatory words and insulting phrases. AI-based determinations include, but are not limited to, machine learning models and rule-based systems. Some or all of the above-described processes in the analysis unit may be performed using, for example, generative AI, or without generative AI. For example, the analysis unit can input the conversation content into the generating AI and have the AI ​​determine whether specific words or phrases constitute harassment.

[0036] The notification unit can automatically send a message to the speaker to alert them to detected remarks. For example, the notification unit can automatically send a message to the speaker stating, "That remark may be harassment." The notification unit can also send a voice message to the speaker to alert them to detected remarks. Furthermore, the notification unit can send a text message to the speaker to alert them to detected remarks. By automatically sending a message to the speaker to alert them to detected remarks, the notification unit can provide the speaker with an opportunity to review their words and actions. Automatic sending includes, but is not limited to, immediate sending or sending when certain conditions are met. Some or all of the above processing in the notification unit may be performed using, for example, a generation AI, or without a generation AI. For example, the notification unit can input the detected remarks into a generation AI and have the generation AI generate a warning message.

[0037] The registration unit can register statements that victims find offensive to the system. For example, the registration unit can register statements that victims find offensive to the system. The registration unit can also register statements that victims find offensive to the system via voice. Furthermore, the registration unit can register statements that victims find offensive to the system via text message. By registering statements that victims find offensive to the system, the AI ​​can learn and improve its detection accuracy. Offensive statements include, but are not limited to, specific words, phrases, or contexts. Some or all of the processing described above in the registration unit may be performed, for example, using a generative AI, or without using a generative AI. For example, the registration unit can input statements that victims find offensive to a generative AI and have the generative AI register the statements.

[0038] The learning unit can learn from registered statements and use this information to improve future detection accuracy. For example, the learning unit can improve detection accuracy by inputting registered statements into a machine learning model and having it learn. Alternatively, the learning unit can improve detection accuracy by learning from registered statements using natural language processing techniques. Furthermore, the learning unit can improve detection accuracy by inputting registered statements into a rule-based system and having it learn. This allows for more accurate harassment detection by learning from registered statements and improving future detection accuracy. Detection accuracy includes, but is not limited to, accuracy, recall, and F-score. Some or all of the above-described processes in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input registered statements into a generative AI and have the generative AI perform the learning.

[0039] The analysis unit can improve the accuracy of its analysis by considering the context of the statements when analyzing conversation content. For example, the analysis unit can analyze the context before and after a statement to determine whether a particular word constitutes harassment. It can also analyze the flow of the conversation to determine whether a statement is taken as a joke. Furthermore, the analysis unit can consider the relationship between the speaker and the listener to analyze the intent of the statement. By considering the context of the statements, the analysis accuracy is improved, enabling more accurate harassment detection. Specific methods for considering context include, but are not limited to, the content of the preceding and following statements and related topics. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input the conversation content into a generative AI and have the generative AI perform an analysis that considers the context.

[0040] The analysis unit can perform analysis by referring to the speaker's past statement history when analyzing conversation content. For example, the analysis unit can check whether the speaker has made similar statements in the past to determine the possibility of harassment. The analysis unit can also analyze whether specific words or phrases are frequently used from the speaker's past statement history. Furthermore, the analysis unit can analyze the intent and background of a statement based on the speaker's past statement history. This allows for a more accurate analysis of the intent and background of a statement by referring to the speaker's past statement history. Specific methods for referring to the statement history include, but are not limited to, past statement content and frequency. Some or all of the above processing in the analysis unit may be performed using, for example, a generating AI, or without a generating AI. For example, the analysis unit can input the speaker's past statement history into a generating AI and have the generating AI perform the analysis.

[0041] The analysis unit can perform analysis of conversation content while considering the speaker's geographical location information. For example, if the speaker is in a specific region, the analysis unit will consider the culture and customs of that region during the analysis. The analysis unit can also determine whether specific words or phrases are region-specific based on the speaker's geographical location information. Furthermore, the analysis unit can analyze the background and intent of the speech based on the speaker's geographical location information. This makes it possible to perform analysis that reflects region-specific culture and customs by considering the speaker's geographical location information. Specific methods for obtaining geographical location information include, but are not limited to, GPS data and IP addresses. Some or all of the above-described processes in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input the speaker's geographical location information into a generative AI and have the generative AI perform the analysis.

[0042] The analysis unit can analyze the speaker's social media activity and analyze relevant statements when analyzing conversation content. For example, the analysis unit can analyze the speaker's social media posting history to check if specific words or phrases are frequently used. The analysis unit can also analyze the intent and background of statements from the speaker's social media activity. Furthermore, the analysis unit can consider the speaker's social media followers and friendships to analyze the influence of statements. This allows for a more accurate analysis of the intent and background of statements by analyzing the speaker's social media activity. Specific methods for analyzing social media activity include, but are not limited to, posting content, likes, and comment history. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the analysis unit can input the speaker's social media activity into a generative AI and have the generative AI perform the analysis.

[0043] The detection unit can improve detection accuracy by considering the frequency of remarks when detecting potentially harassing statements. For example, the detection unit may determine that a statement is likely to be harassment if certain words or phrases are used frequently. The detection unit can also determine whether certain words or phrases constitute harassment based on their frequency of use. Furthermore, the detection unit can improve detection accuracy if certain words or phrases are used repeatedly, taking into account the frequency of remarks. This improves detection accuracy and enables more accurate harassment detection by considering the frequency of remarks. Specific methods for measuring the frequency of remarks include, but are not limited to, the number of remarks or the intervals between remarks within a certain period. Some or all of the above processing in the detection unit may be performed using, for example, a generative AI, or without a generative AI. For example, the detection unit can input remark frequency data into a generative AI and have the generative AI perform the improvement of detection accuracy.

[0044] The detection unit can perform detection by considering the speaker's attribute information when detecting potentially harassing remarks. For example, the detection unit can consider the speaker's age and gender to determine whether a particular word or phrase constitutes harassment. The detection unit can also consider the speaker's job title and position to analyze the influence of the remarks. Furthermore, the detection unit can determine whether a particular word or phrase constitutes harassment based on the speaker's past remark history. This allows for a more accurate analysis of the influence and intent of remarks by considering the speaker's attribute information. Specific types of attribute information include, but are not limited to, age, gender, and occupation. Some or all of the above processing in the detection unit may be performed using, for example, a generative AI, or without a generative AI. For example, the detection unit can input the speaker's attribute information into a generative AI and have the generative AI perform the detection.

[0045] The detection unit can perform detection while considering the speaker's geographical location information when detecting potentially harassing remarks. For example, if the speaker is in a specific region, the detection unit will consider the culture and customs of that region when performing detection. The detection unit can also determine whether specific words or phrases are region-specific based on the speaker's geographical location information. Furthermore, the detection unit can analyze the background and intent of the remarks based on the speaker's geographical location information. This makes it possible to perform detection that reflects region-specific culture and customs by considering the speaker's geographical location information. Specific methods for obtaining geographical location information include, but are not limited to, GPS data and IP addresses. Some or all of the above processing in the detection unit may be performed using, for example, a generative AI, or without a generative AI. For example, the detection unit can input the speaker's geographical location information into a generative AI and have the generative AI perform the detection.

[0046] The detection unit can improve detection accuracy by referring to relevant literature of the speaker when detecting potentially harassing remarks. For example, the detection unit can determine whether specific words or phrases constitute harassment based on the speaker's past remark history. The detection unit can also determine whether specific words or phrases constitute harassment by referring to relevant literature of the speaker. Furthermore, the detection unit can determine whether specific words or phrases constitute harassment based on relevant literature of the speaker. By referring to relevant literature of the speaker, detection accuracy is improved, enabling more accurate harassment detection. Specific types of relevant literature include, but are not limited to, academic papers and industry reports. Some or all of the above processing in the detection unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the detection unit can input relevant literature of the speaker into a generative AI and have the generative AI perform the improvement of detection accuracy.

[0047] The notification unit can adjust the level of detail of a warning message based on the importance of the statement when sending a warning message. For example, the notification unit can send a detailed warning message for high-importance statements. It can also send a concise warning message for low-importance statements. Furthermore, the notification unit can adjust the level of detail of the message based on the importance of the statement. This ensures that a more appropriate warning message is sent by adjusting the level of detail based on the importance of the statement. Specific evaluation criteria for the importance of a statement include, but are not limited to, the influence and content of the statement. Some or all of the above processing in the notification unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the notification unit can input statement importance data into a generative AI and have the generative AI adjust the level of detail of the message.

[0048] The notification unit can apply different message algorithms depending on the category of the statement when sending a warning message. For example, the notification unit can apply different message algorithms depending on the type of harassment. It can also apply different message algorithms depending on the category of the statement. Furthermore, the notification unit can apply different message algorithms depending on the content of the statement. This ensures that more appropriate warning messages are sent by applying different message algorithms depending on the category of the statement. Specific classifications of statement categories include, but are not limited to, insult, discrimination, and sexual harassment. Some or all of the processing described above in the notification unit may be performed using, for example, a generative AI, or without a generative AI. For example, the notification unit can input statement category data into a generative AI and have the generative AI apply the message algorithm.

[0049] The notification unit can determine the priority of warning messages based on when the message was submitted. For example, if the message was made recently, the notification unit will prioritize sending the warning message. Conversely, if the message was made in the past, the notification unit can also send the warning message with a lower priority. Furthermore, the notification unit can also determine the priority of messages based on when the message was submitted. This ensures that warning messages are sent at a more appropriate time by prioritizing messages based on when they were submitted. Specific methods for measuring the submission time include, but are not limited to, the submission date and time or the time elapsed since submission. Some or all of the above processing in the notification unit may be performed using, for example, a generating AI, or not using a generating AI. For example, the notification unit can input the submission time data of the message into a generating AI and have the generating AI determine the priority of messages.

[0050] The notification unit can adjust the order of warning messages based on the relevance of the statements when sending warning messages. For example, the notification unit will prioritize sending warning messages if the statements are highly relevant. Conversely, the notification unit can also lower the priority of sending warning messages if the statements are less relevant. Furthermore, the notification unit can adjust the order of messages based on the relevance of the statements. This ensures that warning messages are sent in a more appropriate order by adjusting the order of messages based on their relevance. Specific criteria for evaluating relevance include, but are not limited to, the similarity of the statements or the relationship between the speakers. Some or all of the above processing in the notification unit may be performed using, for example, a generative AI, or without a generative AI. For example, the notification unit can input statement relevance data into a generative AI and have the generative AI perform the adjustment of the message order.

[0051] The registration unit can improve registration accuracy by considering the context of offensive remarks when registering them. For example, the registration unit can analyze the context before and after a remark to determine whether a particular word constitutes harassment. The registration unit can also analyze the flow of a conversation to determine whether a remark is taken as a joke. Furthermore, the registration unit can consider the relationship between the speaker and the listener to analyze the intent of the remark. By considering the context of the remark, registration accuracy is improved, enabling more accurate harassment detection. Specific methods for considering context include, but are not limited to, the content of the preceding and following remarks and related topics. Some or all of the processing described above in the registration unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the registration unit can input contextual data of the remarks into a generative AI and have the generative AI perform the improvement of registration accuracy.

[0052] The registration unit can prioritize registering highly relevant statements by considering the speaker's geographical location information when registering offensive statements. For example, if the speaker is in a specific region, the registration unit will register the statement considering the culture and customs of that region. The registration unit can also determine whether specific words or phrases are region-specific based on the speaker's geographical location information. Furthermore, the registration unit can analyze the background and intent of a statement based on the speaker's geographical location information and prioritize registering highly relevant statements. This makes it possible to register statements that reflect region-specific culture and customs by considering the speaker's geographical location information. Specific methods for obtaining geographical location information include, but are not limited to, GPS data and IP addresses. Some or all of the above processing in the registration unit may be performed using, for example, a generative AI, or without a generative AI. For example, the registration unit can input the speaker's geographical location information into a generative AI and have the generative AI register highly relevant statements.

[0053] The learning unit can optimize its learning algorithm by referring to past learning data during the learning process. For example, the learning unit can determine whether a particular word or phrase constitutes harassment based on past learning data. The learning unit can also optimize its learning algorithm based on past learning data to improve detection accuracy. Furthermore, the learning unit can analyze whether a particular word or phrase is used repeatedly based on past learning data. This allows the learning algorithm to be optimized and detection accuracy to be improved by referring to past learning data. Specific methods for referring to past learning data include, but are not limited to, the data retention period and data type. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input past learning data into a generative AI and have the generative AI perform the optimization of the learning algorithm.

[0054] The learning unit can weight the training data during training, taking into account the frequency of utterances. For example, the learning unit can set a higher weight for certain words or phrases if they are used frequently. The learning unit can also adjust the weighting of the training data based on the frequency of utterances. Furthermore, the learning unit can set a higher weight for certain words or phrases if they are used repeatedly, taking into account the frequency of utterances. This optimizes the weighting of the training data by considering the frequency of utterances, improving detection accuracy. Specific methods for measuring the frequency of utterances include, but are not limited to, the number of utterances within a certain period or the interval between utterances. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input the utterance frequency data into a generative AI and have the generative AI perform the weighting.

[0055] The learning unit can weight the training data based on the submission date of statements during training. For example, the learning unit can prioritize weighting recent statements as training data. The learning unit can also weight past statements as training data with normal weights. Furthermore, the learning unit can adjust the weighting of the training data based on the submission date of statements. This ensures that more important data is prioritized for training by weighting the training data based on the submission date of statements. Specific methods for measuring the submission date include, but are not limited to, the submission date and time or the time elapsed since submission. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input the submission date data of statements into a generative AI and have the generative AI perform the weighting.

[0056] The learning unit can improve its learning accuracy by referring to relevant literature on the speaker during the learning process. For example, the learning unit can determine whether certain words or phrases constitute harassment based on the speaker's past speaking history. The learning unit can also determine whether certain words or phrases constitute harassment by referring to relevant literature on the speaker. Furthermore, the learning unit can determine whether certain words or phrases constitute harassment based on relevant literature on the speaker. This improves learning accuracy and enables more accurate harassment detection by referring to relevant literature on the speaker. Specific types of relevant literature include, but are not limited to, academic papers and industry reports. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the learning unit can input relevant literature on the speaker into a generative AI and have the generative AI perform the improvement of learning accuracy.

[0057] The learning unit can optimize its learning algorithm by referring to the speaker's past speech history during training. For example, the learning unit can determine whether a particular word or phrase constitutes harassment based on the speaker's past speech history. The learning unit can also optimize the learning algorithm based on the speaker's past speech history to improve detection accuracy. Furthermore, the learning unit can analyze whether a particular word or phrase is used repeatedly based on the speaker's past speech history. This optimizes the learning algorithm and improves detection accuracy by referring to the speaker's past speech history. Specific methods for referring to the speech history include, but are not limited to, past speech content and frequency. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input the speaker's past speech history into a generative AI and have the generative AI perform the optimization of the learning algorithm.

[0058] The learning unit can weight the training data while considering the speaker's geographical location information. For example, if the speaker is in a specific region, the learning unit will weight the training data while considering the culture and customs of that region. The learning unit can also determine whether specific words or phrases are region-specific based on the speaker's geographical location information. Furthermore, the learning unit can analyze the background and intent of a statement based on the speaker's geographical location information and weight the training data accordingly. This makes it possible to perform learning that reflects region-specific culture and customs by considering the speaker's geographical location information. Specific methods for obtaining geographical location information include, but are not limited to, GPS data and IP addresses. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input the speaker's geographical location information into a generative AI and have the generative AI perform the weighting.

[0059] The learning unit can analyze the speaker's social media activity during training and incorporate relevant statements as training data. For example, the learning unit can analyze the speaker's social media posting history to check if specific words or phrases are frequently used. The learning unit can also analyze the intent and background of statements from the speaker's social media activity and incorporate this as training data. Furthermore, the learning unit can consider the speaker's social media followers and friendships to analyze the influence of statements and incorporate this as training data. This allows for more accurate learning of the intent and background of statements by analyzing the speaker's social media activity. Specific methods for analyzing social media activity include, but are not limited to, posting content, likes, and comment history. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the learning unit can input the speaker's social media activity into a generative AI and have the generative AI incorporate it as training data.

[0060] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0061] The analysis unit can improve the accuracy of its analysis by considering the context of the statements when analyzing conversation content. For example, it can analyze the context before and after a statement to determine whether a particular word constitutes harassment. It can also analyze the flow of the conversation to determine whether a statement is taken as a joke. Furthermore, it can consider the relationship between the speaker and the listener to analyze the intent of the statement. As a result, considering the context of the statements improves the accuracy of the analysis and enables more accurate harassment detection. Specific methods for considering context include, but are not limited to, the content of the preceding and following statements and related topics.

[0062] The analysis unit can analyze conversation content by referring to the speaker's past statements. For example, it can check whether the speaker has made similar statements in the past to determine the possibility of harassment. It can also analyze whether specific words or phrases are frequently used based on the speaker's past statements. Furthermore, it can analyze the intent and background of a statement based on the speaker's past statements. This allows for a more accurate analysis of the intent and background of a statement by referring to the speaker's past statements. Specific methods for referring to the statement history include, but are not limited to, past statements and their frequency.

[0063] The analysis unit can perform analysis of conversation content while considering the speaker's geographical location. For example, if the speaker is in a specific region, the analysis will take into account the culture and customs of that region. It can also determine whether specific words or phrases are unique to a particular region based on the speaker's geographical location. Furthermore, it can analyze the background and intent of a statement based on the speaker's geographical location. This makes it possible to perform analysis that reflects the culture and customs unique to a region by considering the speaker's geographical location. Specific methods for obtaining geographical location information include, but are not limited to, GPS data and IP addresses.

[0064] The detection unit can improve detection accuracy by considering the frequency of remarks when detecting potentially harassing statements. For example, if a particular word or phrase is used frequently, it may be judged as likely to be harassment. It can also determine whether a particular word or phrase constitutes harassment based on its frequency of use. Furthermore, by considering the frequency of remarks, detection accuracy can be improved if a particular word or phrase is used repeatedly. As a result, by considering the frequency of remarks, detection accuracy is improved, enabling more accurate harassment detection. Specific methods for measuring the frequency of remarks include, but are not limited to, the number of remarks made within a certain period or the intervals between remarks.

[0065] The notification unit can adjust the level of detail in a warning message based on the importance of the statement when sending a warning message. For example, a detailed warning message can be sent for high-importance statements. Conversely, a concise warning message can be sent for low-importance statements. Furthermore, the level of detail can be adjusted based on the importance of the statement. This ensures that more appropriate warning messages are sent by adjusting the level of detail based on the importance of the statement. Specific criteria for evaluating the importance of a statement include, but are not limited to, the impact and content of the statement.

[0066] The following briefly describes the processing flow for example form 1.

[0067] Step 1: The analysis unit analyzes the conversation content in real time. The analysis unit uses natural language processing technology, machine learning models, and rule-based systems to analyze the conversation content and determine whether specific words or phrases constitute harassment. Step 2: The detection unit detects potentially harassing remarks from the conversation content analyzed by the analysis unit. The detection unit uses natural language processing technology, machine learning models, and rule-based systems to detect whether specific words or phrases constitute harassment. Step 3: The notification unit sends a message to the speaker regarding the statement detected by the detection unit. The notification unit can automatically send a message to the speaker regarding the statement detected by the detection unit, warning them that "that statement may be harassment." It can also send the message via voice or text message. Step 4: The registration unit registers comments that the user finds offensive in everyday life. The registration unit allows the victim to register comments that they find offensive in the system, and these can be registered via voice or text message. Step 5: The learning unit learns from the utterances registered by the registration unit to improve detection accuracy. The learning unit improves detection accuracy by inputting the utterances registered by the registration unit into machine learning models, natural language processing technologies, and rule-based systems, and learning from them.

[0068] (Example of form 2) The harassment prevention system according to an embodiment of the present invention is a system that encourages perpetrators of harassment to recognize their own behavior and review their words and actions. The harassment prevention system detects potentially harassing remarks and alerts the speaker, and when an employee receives a remark they find offensive in their daily life, the content of that remark is registered, allowing the AI ​​to learn and improve detection accuracy. This contributes to creating a workplace environment where all employees can work with peace of mind. For example, the harassment prevention system uses AI to analyze conversation content in real time and detect potentially harassing remarks. For example, during a meeting or chat exchange, the AI ​​determines whether specific words or phrases constitute harassment. Next, the harassment prevention system automatically sends an alert message to the speaker regarding the detected remark. This gives the speaker an opportunity to review their words and actions. Furthermore, when an employee receives a remark they find offensive in their daily life, they register the content of that remark in the system. The registered remarks are learned by the AI ​​and used to improve future detection accuracy. For example, by repeatedly registering specific words or phrases, the AI ​​can determine that these words or phrases are highly likely to constitute harassment, enabling faster and more accurate detection. This mechanism encourages perpetrators of harassment to recognize their own behavior and reconsider their actions. Furthermore, by registering remarks that victims find offensive, the AI ​​learns and improves detection accuracy. This contributes to creating a workplace environment where all employees can work with peace of mind. For instance, if someone says "You're always late" during a meeting, the AI ​​will detect the remark and send a message to the speaker warning them that "That remark may be harassment." Also, if a victim registers with the system that "I found my boss's remark 'You're useless' offensive," the AI ​​learns from that remark and can quickly detect similar remarks in the future. In this way, the harassment prevention system encourages perpetrators of harassment to recognize their own behavior and reconsider their actions, while also providing a mechanism for the AI ​​to learn from and improve detection accuracy through the registration of remarks that victims find offensive. This contributes to creating a workplace environment where all employees can work with peace of mind.This allows harassment prevention systems to encourage perpetrators of harassment to recognize their actions and reconsider their behavior, thereby contributing to the improvement of the workplace environment.

[0069] The harassment prevention system according to this embodiment comprises an analysis unit, a detection unit, a notification unit, a registration unit, and a learning unit. The analysis unit analyzes conversation content in real time. The analysis unit uses AI to determine, for example, whether specific words or phrases constitute harassment in a meeting or chat exchange. The analysis unit can analyze conversation content using, for example, natural language processing technology to determine whether specific words or phrases constitute harassment. The analysis unit can also use a machine learning model to analyze conversation content and detect potentially harassing remarks. Furthermore, the analysis unit can use a rule-based system to determine whether specific words or phrases constitute harassment. The detection unit detects potentially harassing remarks from the conversation content analyzed by the analysis unit. The detection unit can, for example, use natural language processing technology to detect whether specific words or phrases constitute harassment from the conversation content analyzed by the analysis unit. The detection unit can also use a machine learning model to detect potentially harassing remarks from the conversation content analyzed by the analysis unit. Furthermore, the detection unit can use a rule-based system to detect whether specific words or phrases constitute harassment from the conversation content analyzed by the analysis unit. The notification unit sends a message to the speaker to alert them to the statement detected by the detection unit. For example, the notification unit can automatically send a message to the speaker informing them that "that statement may be harassment" in response to a statement detected by the detection unit. The notification unit can also send a voice message to the speaker in response to a statement detected by the detection unit. Furthermore, the notification unit can send a text message to the speaker in response to a statement detected by the detection unit. The registration unit registers statements that are found offensive in everyday life. For example, the registration unit can register statements that victims find offensive in the system. The registration unit can also register statements that victims find offensive in voice. Furthermore, the registration unit can register statements that victims find offensive in text messages.The learning unit learns the content of statements registered by the registration unit to improve detection accuracy. For example, the learning unit can improve detection accuracy by inputting the content of statements registered by the registration unit into a machine learning model and learning from it. The learning unit can also improve detection accuracy by learning the content of statements registered by the registration unit using natural language processing technology. Furthermore, the learning unit can improve detection accuracy by inputting the content of statements registered by the registration unit into a rule-based system and learning from it. As a result, the harassment prevention system according to this embodiment can encourage perpetrators of harassment to recognize their own behavior and review their words and actions, thereby contributing to the improvement of the workplace environment.

[0070] The analysis unit analyzes conversation content in real time. For example, in meetings or chat exchanges, the analysis unit uses AI to determine whether specific words or phrases constitute harassment. Specifically, the analysis unit uses natural language processing technology to analyze the context and emotions of the conversation and determine whether specific words or phrases constitute harassment. Natural language processing technology includes morphological analysis, grammatical analysis, and sentiment analysis, which are combined to analyze the conversation content in detail. For example, morphological analysis is used to break down words in a sentence, grammatical analysis is used to understand the structure of the sentence, and sentiment analysis is used to infer the speaker's emotions. Furthermore, the analysis unit uses machine learning models to detect potentially harassing remarks based on past data. Machine learning models have the ability to learn from large amounts of conversation data and identify patterns that constitute harassment. This allows the analysis unit to detect new signs of harassment. It can also use a rule-based system to determine whether specific words or phrases constitute harassment. The rule-based system evaluates the conversation content based on predefined rules and identifies potentially harassing remarks. This allows the analysis unit to analyze conversation content from multiple perspectives, enabling early detection of harassment.

[0071] The detection unit detects potentially harassing remarks from the conversation content analyzed by the analysis unit. For example, the detection unit can use natural language processing techniques to detect whether specific words or phrases constitute harassment from the conversation content analyzed by the analysis unit. Specifically, the detection unit identifies potentially harassing remarks based on the analysis results provided by the analysis unit. Natural language processing techniques include text classification, keyword extraction, and sentiment analysis, which are combined to detect the content of remarks in detail. For example, text classification is used to categorize remarks, keyword extraction is used to extract specific words or phrases, and sentiment analysis is used to evaluate the speaker's emotions. Furthermore, the detection unit can also use machine learning models to detect potentially harassing remarks from the conversation content analyzed by the analysis unit. Machine learning models have the ability to learn from large amounts of conversation data and identify patterns that constitute harassment. This allows the detection unit to detect new signs of harassment. Additionally, rule-based systems can be used to detect whether specific words or phrases constitute harassment from the conversation content analyzed by the analysis unit. The rule-based system evaluates conversation content based on predefined rules and identifies potentially harassing remarks. This allows the detection unit to detect conversation content from multiple angles, enabling early detection of harassment.

[0072] The notification unit sends a message to the speaker to alert them to the statement detected by the detection unit. Specifically, the notification unit can automatically send a message to the speaker informing them that "that statement may be harassment" in response to a statement detected by the detection unit. The notification unit can, for example, send a message to the speaker's device in real time to alert them. Message delivery methods include text messages, voice messages, and pop-up notifications. Text messages are sent directly to the speaker's device, allowing the speaker to see the message. Voice messages provide an audible alert from the speaker's device, allowing the speaker to respond immediately. Pop-up notifications appear in a pop-up format on the speaker's device, allowing the speaker to see the alert message. Furthermore, the notification unit can send alert messages not only to the speaker but also to managers and supervisors. This allows managers and supervisors to understand the situation and take appropriate action. The notification unit can also record the speaker's response and provide data for later analysis. This allows the notification unit to alert the speaker quickly and appropriately, preventing the recurrence of harassment.

[0073] The registration unit registers statements that users find offensive in their daily lives. Specifically, the registration unit allows victims to register statements they find offensive in the system. Registration methods include text input, voice input, and sending text messages. Victims can use text input to describe offensive statements in detail and register them in the system. When using voice input, victims can record offensive statements aloud and register them in the system. When using text messages, victims can send offensive statements to the system as text messages and register them. Furthermore, the registration unit classifies the statements registered by victims and stores them in a database for use by the analysis and detection units. The registered statements are used as training data for the analysis and detection units to improve the accuracy of harassment detection. This allows the registration unit to efficiently register statements that victims find offensive and improve the overall harassment detection capability of the system. In addition, the registration unit can anonymize statements registered by victims to protect their privacy. This allows victims to register offensive statements with peace of mind.

[0074] The learning unit improves detection accuracy by learning from the utterances registered by the registration unit. Specifically, the learning unit can improve detection accuracy by inputting the utterances registered by the registration unit into a machine learning model and learning from it. The machine learning model has the ability to learn from large amounts of data and identify patterns that constitute harassment. This allows the learning unit to detect new signs of harassment. The learning unit can also improve detection accuracy by learning from the utterances registered by the registration unit using natural language processing techniques. Natural language processing techniques include morphological analysis, grammatical analysis, and sentiment analysis, and these are combined to analyze and learn from the utterances in detail. Furthermore, the learning unit can improve detection accuracy by inputting the utterances registered by the registration unit into a rule-based system and learning from it. The rule-based system evaluates the utterances based on predefined rules and identifies utterances that may constitute harassment. This allows the learning unit to learn from utterances using a multifaceted approach and improve the accuracy of harassment detection. In addition, the learning unit can maintain detection accuracy by periodically retraining the model based on the latest data. This allows the learning unit to always perform highly accurate harassment detection based on the latest information, improving the reliability and security of the entire system.

[0075] The analysis unit can use AI to determine whether specific words or phrases in meetings or chat interactions constitute harassment. For example, the analysis unit can analyze conversation content using natural language processing technology to determine whether specific words or phrases constitute harassment. Furthermore, the analysis unit can use machine learning models to determine whether specific words or phrases in meetings or chat interactions constitute harassment. In addition, the analysis unit can use rule-based systems to determine whether specific words or phrases constitute harassment. This enables real-time harassment detection by using AI to determine whether specific words or phrases in meetings or chat interactions constitute harassment. Specific words and phrases include, but are not limited to, discriminatory words and insulting phrases. AI-based determinations include, but are not limited to, machine learning models and rule-based systems. Some or all of the above-described processes in the analysis unit may be performed using, for example, generative AI, or without generative AI. For example, the analysis unit can input the conversation content into the generating AI and have the AI ​​determine whether specific words or phrases constitute harassment.

[0076] The notification unit can automatically send a message to the speaker to alert them to detected remarks. For example, the notification unit can automatically send a message to the speaker stating, "That remark may be harassment." The notification unit can also send a voice message to the speaker to alert them to detected remarks. Furthermore, the notification unit can send a text message to the speaker to alert them to detected remarks. By automatically sending a message to the speaker to alert them to detected remarks, the notification unit can provide the speaker with an opportunity to review their words and actions. Automatic sending includes, but is not limited to, immediate sending or sending when certain conditions are met. Some or all of the above processing in the notification unit may be performed using, for example, a generation AI, or without a generation AI. For example, the notification unit can input the detected remarks into a generation AI and have the generation AI generate a warning message.

[0077] The registration unit can register statements that victims find offensive to the system. For example, the registration unit can register statements that victims find offensive to the system. The registration unit can also register statements that victims find offensive to the system via voice. Furthermore, the registration unit can register statements that victims find offensive to the system via text message. By registering statements that victims find offensive to the system, the AI ​​can learn and improve its detection accuracy. Offensive statements include, but are not limited to, specific words, phrases, or contexts. Some or all of the processing described above in the registration unit may be performed, for example, using a generative AI, or without using a generative AI. For example, the registration unit can input statements that victims find offensive to a generative AI and have the generative AI register the statements.

[0078] The learning unit can learn from registered statements and use this information to improve future detection accuracy. For example, the learning unit can improve detection accuracy by inputting registered statements into a machine learning model and having it learn. Alternatively, the learning unit can improve detection accuracy by learning from registered statements using natural language processing techniques. Furthermore, the learning unit can improve detection accuracy by inputting registered statements into a rule-based system and having it learn. This allows for more accurate harassment detection by learning from registered statements and improving future detection accuracy. Detection accuracy includes, but is not limited to, accuracy, recall, and F-score. Some or all of the above-described processes in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input registered statements into a generative AI and have the generative AI perform the learning.

[0079] The analysis unit can estimate the user's emotions and adjust the method of analyzing the conversation content based on the estimated user emotions. For example, if the user is angry, the AI ​​will analyze the tone and wording of the speech with particular care. Similarly, if the user is sad, the AI ​​can focus on analyzing speech containing emotional expressions. Furthermore, if the user is relaxed, the AI ​​can apply normal analysis methods and capture a broader context of the speech. This allows for more appropriate analysis by adjusting the conversation content analysis method based on the user's emotions. User emotions include, but are not limited to, facial recognition, speech analysis, and text analysis. Adjusting the analysis method includes, but is not limited to, changing the analysis algorithm or adjusting analysis parameters. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the analysis unit may be performed using, for example, a generative AI, or without using a generative AI. For example, the analysis unit can input user emotion data into a generative AI and have the generative AI perform adjustments to the analysis method.

[0080] The analysis unit can improve the accuracy of its analysis by considering the context of the statements when analyzing conversation content. For example, the analysis unit can analyze the context before and after a statement to determine whether a particular word constitutes harassment. It can also analyze the flow of the conversation to determine whether a statement is taken as a joke. Furthermore, the analysis unit can consider the relationship between the speaker and the listener to analyze the intent of the statement. By considering the context of the statements, the analysis accuracy is improved, enabling more accurate harassment detection. Specific methods for considering context include, but are not limited to, the content of the preceding and following statements and related topics. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input the conversation content into a generative AI and have the generative AI perform an analysis that considers the context.

[0081] The analysis unit can perform analysis by referring to the speaker's past statement history when analyzing conversation content. For example, the analysis unit can check whether the speaker has made similar statements in the past to determine the possibility of harassment. The analysis unit can also analyze whether specific words or phrases are frequently used from the speaker's past statement history. Furthermore, the analysis unit can analyze the intent and background of a statement based on the speaker's past statement history. This allows for a more accurate analysis of the intent and background of a statement by referring to the speaker's past statement history. Specific methods for referring to the statement history include, but are not limited to, past statement content and frequency. Some or all of the above processing in the analysis unit may be performed using, for example, a generating AI, or without a generating AI. For example, the analysis unit can input the speaker's past statement history into a generating AI and have the generating AI perform the analysis.

[0082] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated user emotions. For example, if the user is nervous, the analysis unit can provide a simple and highly visible display method. If the user is relaxed, the analysis unit can also provide a display method that includes detailed information. Furthermore, if the user is in a hurry, the analysis unit can provide a display method that gets straight to the point. By adjusting the display method of the analysis results based on the user's emotions, it becomes possible to provide a display that is easy for the user to understand. User emotions include, but are not limited to, facial recognition, voice analysis, and text analysis. Adjusting the display method includes, but are not limited to, changing the display format or filtering the displayed content. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, for example, but is not limited to these examples. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input user emotion data into the generating AI and have the generating AI adjust the display method.

[0083] The analysis unit can perform analysis of conversation content while considering the speaker's geographical location information. For example, if the speaker is in a specific region, the analysis unit will consider the culture and customs of that region during the analysis. The analysis unit can also determine whether specific words or phrases are region-specific based on the speaker's geographical location information. Furthermore, the analysis unit can analyze the background and intent of the speech based on the speaker's geographical location information. This makes it possible to perform analysis that reflects region-specific culture and customs by considering the speaker's geographical location information. Specific methods for obtaining geographical location information include, but are not limited to, GPS data and IP addresses. Some or all of the above-described processes in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input the speaker's geographical location information into a generative AI and have the generative AI perform the analysis.

[0084] The analysis unit can analyze the speaker's social media activity and analyze relevant statements when analyzing conversation content. For example, the analysis unit can analyze the speaker's social media posting history to check if specific words or phrases are frequently used. The analysis unit can also analyze the intent and background of statements from the speaker's social media activity. Furthermore, the analysis unit can consider the speaker's social media followers and friendships to analyze the influence of statements. This allows for a more accurate analysis of the intent and background of statements by analyzing the speaker's social media activity. Specific methods for analyzing social media activity include, but are not limited to, posting content, likes, and comment history. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the analysis unit can input the speaker's social media activity into a generative AI and have the generative AI perform the analysis.

[0085] The detection unit can estimate the user's emotions and adjust the detection criteria for potentially harassing remarks based on the estimated user emotions. For example, if the user is angry, the AI ​​will pay particular attention to the tone and wording of the remarks. Similarly, if the user is sad, the AI ​​can focus on detecting remarks containing emotional expressions. Furthermore, if the user is relaxed, the AI ​​can apply normal detection criteria and capture a broader context of the remarks. This allows for more accurate harassment detection by adjusting the detection criteria based on the user's emotions. User emotions include, but are not limited to, facial recognition, voice analysis, and text analysis. Adjusting the detection criteria includes, but are not limited to, changing thresholds or adjusting the detection algorithm. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the detection unit may be performed using, for example, a generating AI, or without using a generating AI. For example, the detection unit can input user emotion data into a generating AI and have the generating AI adjust the detection criteria.

[0086] The detection unit can improve detection accuracy by considering the frequency of remarks when detecting potentially harassing statements. For example, the detection unit may determine that a statement is likely to be harassment if certain words or phrases are used frequently. The detection unit can also determine whether certain words or phrases constitute harassment based on their frequency of use. Furthermore, the detection unit can improve detection accuracy if certain words or phrases are used repeatedly, taking into account the frequency of remarks. This improves detection accuracy and enables more accurate harassment detection by considering the frequency of remarks. Specific methods for measuring the frequency of remarks include, but are not limited to, the number of remarks or the intervals between remarks within a certain period. Some or all of the above processing in the detection unit may be performed using, for example, a generative AI, or without a generative AI. For example, the detection unit can input remark frequency data into a generative AI and have the generative AI perform the improvement of detection accuracy.

[0087] The detection unit can perform detection by considering the speaker's attribute information when detecting potentially harassing remarks. For example, the detection unit can consider the speaker's age and gender to determine whether a particular word or phrase constitutes harassment. The detection unit can also consider the speaker's job title and position to analyze the influence of the remarks. Furthermore, the detection unit can determine whether a particular word or phrase constitutes harassment based on the speaker's past remark history. This allows for a more accurate analysis of the influence and intent of remarks by considering the speaker's attribute information. Specific types of attribute information include, but are not limited to, age, gender, and occupation. Some or all of the above processing in the detection unit may be performed using, for example, a generative AI, or without a generative AI. For example, the detection unit can input the speaker's attribute information into a generative AI and have the generative AI perform the detection.

[0088] The detection unit can estimate the user's emotions and adjust the display method of the detection results based on the estimated user emotions. For example, if the user is tense, the detection unit provides a simple and highly visible display method. If the user is relaxed, the detection unit can also provide a display method that includes detailed information. Furthermore, if the user is in a hurry, the detection unit can provide a display method that gets straight to the point. By adjusting the display method of the detection results based on the user's emotions, it becomes possible to provide a display that is easy for the user to understand. User emotions include, but are not limited to, facial recognition, voice analysis, and text analysis. Adjusting the display method includes, but are not limited to, changing the display format or filtering the display content. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, for example, but is not limited to these examples. Some or all of the above processing in the detection unit may be performed using, for example, a generative AI, or without a generative AI. For example, the detection unit can input user emotion data into the generating AI and have the generating AI adjust the display method.

[0089] The detection unit can perform detection while considering the speaker's geographical location information when detecting potentially harassing remarks. For example, if the speaker is in a specific region, the detection unit will consider the culture and customs of that region when performing detection. The detection unit can also determine whether specific words or phrases are region-specific based on the speaker's geographical location information. Furthermore, the detection unit can analyze the background and intent of the remarks based on the speaker's geographical location information. This makes it possible to perform detection that reflects region-specific culture and customs by considering the speaker's geographical location information. Specific methods for obtaining geographical location information include, but are not limited to, GPS data and IP addresses. Some or all of the above processing in the detection unit may be performed using, for example, a generative AI, or without a generative AI. For example, the detection unit can input the speaker's geographical location information into a generative AI and have the generative AI perform the detection.

[0090] The detection unit can improve detection accuracy by referring to relevant literature of the speaker when detecting potentially harassing remarks. For example, the detection unit can determine whether specific words or phrases constitute harassment based on the speaker's past remark history. The detection unit can also determine whether specific words or phrases constitute harassment by referring to relevant literature of the speaker. Furthermore, the detection unit can determine whether specific words or phrases constitute harassment based on relevant literature of the speaker. By referring to relevant literature of the speaker, detection accuracy is improved, enabling more accurate harassment detection. Specific types of relevant literature include, but are not limited to, academic papers and industry reports. Some or all of the above processing in the detection unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the detection unit can input relevant literature of the speaker into a generative AI and have the generative AI perform the improvement of detection accuracy.

[0091] The notification unit can estimate the user's emotions and adjust the way the warning message is delivered based on the estimated emotions. For example, if the user is tense, the notification unit can send a warning message in a calm tone. If the user is relaxed, the notification unit can also send a warning message in a friendly tone. Furthermore, if the user is in a hurry, the notification unit can send a concise and quick warning message. By adjusting the way the warning message is delivered based on the user's emotions, a more appropriate message is delivered. User emotions include, but are not limited to, facial recognition, speech analysis, and text analysis. Adjustments to the delivery method include, but are not limited to, the tone and wording of the message. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) and multimodal generation AI. Some or all of the above processing in the notification unit may be performed using, for example, generative AI, or without generative AI. For example, the notification unit can input user emotion data into a generating AI and have the generating AI adjust the way the message is expressed.

[0092] The notification unit can adjust the level of detail of a warning message based on the importance of the statement when sending a warning message. For example, the notification unit can send a detailed warning message for high-importance statements. It can also send a concise warning message for low-importance statements. Furthermore, the notification unit can adjust the level of detail of the message based on the importance of the statement. This ensures that a more appropriate warning message is sent by adjusting the level of detail based on the importance of the statement. Specific evaluation criteria for the importance of a statement include, but are not limited to, the influence and content of the statement. Some or all of the above processing in the notification unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the notification unit can input statement importance data into a generative AI and have the generative AI adjust the level of detail of the message.

[0093] The notification unit can apply different message algorithms depending on the category of the statement when sending a warning message. For example, the notification unit can apply different message algorithms depending on the type of harassment. It can also apply different message algorithms depending on the category of the statement. Furthermore, the notification unit can apply different message algorithms depending on the content of the statement. This ensures that more appropriate warning messages are sent by applying different message algorithms depending on the category of the statement. Specific classifications of statement categories include, but are not limited to, insult, discrimination, and sexual harassment. Some or all of the processing described above in the notification unit may be performed using, for example, a generative AI, or without a generative AI. For example, the notification unit can input statement category data into a generative AI and have the generative AI apply the message algorithm.

[0094] The notification unit can estimate the user's emotions and adjust the length of the warning message based on the estimated emotions. For example, if the user is tense, the notification unit can send a short, concise warning message. It can also send a more detailed warning message if the user is relaxed. Furthermore, if the user is in a hurry, the notification unit can send a brief, quick warning message. By adjusting the length of the warning message based on the user's emotions, a more appropriate warning message is delivered. User emotions include, but are not limited to, facial recognition, speech analysis, and text analysis. Adjusting the message length includes, but are not limited to, the number of characters or paragraphs. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the notification unit may be performed using, for example, generative AI, or without generative AI. For example, the notification unit can input user emotion data into a generating AI and have the AI ​​adjust the length of the message.

[0095] The notification unit can determine the priority of warning messages based on when the message was submitted. For example, if the message was made recently, the notification unit will prioritize sending the warning message. Conversely, if the message was made in the past, the notification unit can also send the warning message with a lower priority. Furthermore, the notification unit can also determine the priority of messages based on when the message was submitted. This ensures that warning messages are sent at a more appropriate time by prioritizing messages based on when they were submitted. Specific methods for measuring the submission time include, but are not limited to, the submission date and time or the time elapsed since submission. Some or all of the above processing in the notification unit may be performed using, for example, a generating AI, or not using a generating AI. For example, the notification unit can input the submission time data of the message into a generating AI and have the generating AI determine the priority of messages.

[0096] The notification unit can adjust the order of warning messages based on the relevance of the statements when sending warning messages. For example, the notification unit will prioritize sending warning messages if the statements are highly relevant. Conversely, the notification unit can also lower the priority of sending warning messages if the statements are less relevant. Furthermore, the notification unit can adjust the order of messages based on the relevance of the statements. This ensures that warning messages are sent in a more appropriate order by adjusting the order of messages based on their relevance. Specific criteria for evaluating relevance include, but are not limited to, the similarity of the statements or the relationship between the speakers. Some or all of the above processing in the notification unit may be performed using, for example, a generative AI, or without a generative AI. For example, the notification unit can input statement relevance data into a generative AI and have the generative AI perform the adjustment of the message order.

[0097] The registration unit can estimate the user's emotions and adjust the registration method for offensive remarks based on the estimated user emotions. For example, if the user is angry, the registration unit can enable quick registration of remarks through a simple interface. If the user is sad, the registration unit can also provide detailed input options to express their emotions. Furthermore, if the user is relaxed, the registration unit can provide a normal registration method that captures a broader context of the remarks. This allows for more appropriate registration by adjusting the registration method for offensive remarks based on the user's emotions. User emotions include, but are not limited to, facial recognition, speech analysis, and text analysis. Adjusting the registration method includes, but is not limited to, changing the registration procedure or adding registration items. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the registration unit may be performed using, for example, generative AI, or without generative AI. For example, the registration unit can input user emotion data into a generating AI and have the generating AI adjust the registration method.

[0098] The registration unit can improve registration accuracy by considering the context of offensive remarks when registering them. For example, the registration unit can analyze the context before and after a remark to determine whether a particular word constitutes harassment. The registration unit can also analyze the flow of a conversation to determine whether a remark is taken as a joke. Furthermore, the registration unit can consider the relationship between the speaker and the listener to analyze the intent of the remark. By considering the context of the remark, registration accuracy is improved, enabling more accurate harassment detection. Specific methods for considering context include, but are not limited to, the content of the preceding and following remarks and related topics. Some or all of the processing described above in the registration unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the registration unit can input contextual data of the remarks into a generative AI and have the generative AI perform the improvement of registration accuracy.

[0099] The registration unit can estimate the user's emotions and determine the priority of statements to register based on the estimated user emotions. For example, if the user is showing strong discomfort, the registration unit will prioritize registering that statement. If the user is showing mild discomfort, the registration unit can register it with normal priority. Furthermore, the registration unit can adjust the registration priority of statements based on the intensity of the user's emotions. This ensures that more important statements are registered preferentially by determining the priority of statements to register based on the user's emotions. User emotions include, but are not limited to, facial recognition, speech analysis, and text analysis. Prioritization includes, but are not limited to, importance and urgency. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) and multimodal generation AI. Some or all of the above processing in the registration unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the registration unit can input user emotion data into a generating AI and have the generating AI determine priorities.

[0100] The registration unit can prioritize registering highly relevant statements by considering the speaker's geographical location information when registering offensive statements. For example, if the speaker is in a specific region, the registration unit will register the statement considering the culture and customs of that region. The registration unit can also determine whether specific words or phrases are region-specific based on the speaker's geographical location information. Furthermore, the registration unit can analyze the background and intent of a statement based on the speaker's geographical location information and prioritize registering highly relevant statements. This makes it possible to register statements that reflect region-specific culture and customs by considering the speaker's geographical location information. Specific methods for obtaining geographical location information include, but are not limited to, GPS data and IP addresses. Some or all of the above processing in the registration unit may be performed using, for example, a generative AI, or without a generative AI. For example, the registration unit can input the speaker's geographical location information into a generative AI and have the generative AI register highly relevant statements.

[0101] The learning unit can estimate the user's emotions and select training data based on the estimated user emotions. For example, the learning unit can prioritize selecting statements that indicate strong user discomfort as training data. It can also select statements that indicate mild user discomfort as training data with normal priority. Furthermore, the learning unit can adjust the selection of training data based on the intensity of the user's emotions. This ensures that more important data is prioritized for training by selecting training data based on the user's emotions. User emotions include, but are not limited to, facial recognition, speech analysis, and text analysis. Selection of training data includes, but are not limited to, data importance and data relevance. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) and multimodal generation AI. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input user emotion data into the generating AI and have the generating AI select the training data.

[0102] The learning unit can optimize its learning algorithm by referring to past learning data during the learning process. For example, the learning unit can determine whether a particular word or phrase constitutes harassment based on past learning data. The learning unit can also optimize its learning algorithm based on past learning data to improve detection accuracy. Furthermore, the learning unit can analyze whether a particular word or phrase is used repeatedly based on past learning data. This allows the learning algorithm to be optimized and detection accuracy to be improved by referring to past learning data. Specific methods for referring to past learning data include, but are not limited to, the data retention period and data type. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input past learning data into a generative AI and have the generative AI perform the optimization of the learning algorithm.

[0103] The learning unit can weight the training data during training, taking into account the frequency of utterances. For example, the learning unit can set a higher weight for certain words or phrases if they are used frequently. The learning unit can also adjust the weighting of the training data based on the frequency of utterances. Furthermore, the learning unit can set a higher weight for certain words or phrases if they are used repeatedly, taking into account the frequency of utterances. This optimizes the weighting of the training data by considering the frequency of utterances, improving detection accuracy. Specific methods for measuring the frequency of utterances include, but are not limited to, the number of utterances within a certain period or the interval between utterances. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input the utterance frequency data into a generative AI and have the generative AI perform the weighting.

[0104] The learning unit can estimate the user's emotions and adjust the learning frequency based on the estimated emotions. For example, if the user is showing strong discomfort, the learning unit will increase the learning frequency. Conversely, if the user is showing mild discomfort, the learning unit can perform learning at the normal frequency. Furthermore, the learning unit can also adjust the learning frequency based on the intensity of the user's emotions. This allows more important data to be learned preferentially by adjusting the learning frequency based on the user's emotions. User emotions include, but are not limited to, facial recognition, speech analysis, and text analysis. Adjustments to the learning frequency include, but are not limited to, the learning interval and the number of learning iterations. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) and multimodal generation AI. Some or all of the above-described processing in the learning unit may be performed using, for example, generative AI, or without generative AI. For example, the learning unit can input user emotion data into the generating AI and have the generating AI adjust the frequency of learning.

[0105] The learning unit can weight the training data based on the submission date of statements during training. For example, the learning unit can prioritize weighting recent statements as training data. The learning unit can also weight past statements as training data with normal weights. Furthermore, the learning unit can adjust the weighting of the training data based on the submission date of statements. This ensures that more important data is prioritized for training by weighting the training data based on the submission date of statements. Specific methods for measuring the submission date include, but are not limited to, the submission date and time or the time elapsed since submission. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input the submission date data of statements into a generative AI and have the generative AI perform the weighting.

[0106] The learning unit can improve its learning accuracy by referring to relevant literature on the speaker during the learning process. For example, the learning unit can determine whether certain words or phrases constitute harassment based on the speaker's past speaking history. The learning unit can also determine whether certain words or phrases constitute harassment by referring to relevant literature on the speaker. Furthermore, the learning unit can determine whether certain words or phrases constitute harassment based on relevant literature on the speaker. This improves learning accuracy and enables more accurate harassment detection by referring to relevant literature on the speaker. Specific types of relevant literature include, but are not limited to, academic papers and industry reports. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the learning unit can input relevant literature on the speaker into a generative AI and have the generative AI perform the improvement of learning accuracy.

[0107] The learning unit can optimize its learning algorithm by referring to the speaker's past speech history during training. For example, the learning unit can determine whether a particular word or phrase constitutes harassment based on the speaker's past speech history. The learning unit can also optimize the learning algorithm based on the speaker's past speech history to improve detection accuracy. Furthermore, the learning unit can analyze whether a particular word or phrase is used repeatedly based on the speaker's past speech history. This optimizes the learning algorithm and improves detection accuracy by referring to the speaker's past speech history. Specific methods for referring to the speech history include, but are not limited to, past speech content and frequency. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input the speaker's past speech history into a generative AI and have the generative AI perform the optimization of the learning algorithm.

[0108] The learning unit can weight the training data while considering the speaker's geographical location information. For example, if the speaker is in a specific region, the learning unit will weight the training data while considering the culture and customs of that region. The learning unit can also determine whether specific words or phrases are region-specific based on the speaker's geographical location information. Furthermore, the learning unit can analyze the background and intent of a statement based on the speaker's geographical location information and weight the training data accordingly. This makes it possible to perform learning that reflects region-specific culture and customs by considering the speaker's geographical location information. Specific methods for obtaining geographical location information include, but are not limited to, GPS data and IP addresses. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or without a generative AI. For example, the learning unit can input the speaker's geographical location information into a generative AI and have the generative AI perform the weighting.

[0109] The learning unit can analyze the speaker's social media activity during training and incorporate relevant statements as training data. For example, the learning unit can analyze the speaker's social media posting history to check if specific words or phrases are frequently used. The learning unit can also analyze the intent and background of statements from the speaker's social media activity and incorporate this as training data. Furthermore, the learning unit can consider the speaker's social media followers and friendships to analyze the influence of statements and incorporate this as training data. This allows for more accurate learning of the intent and background of statements by analyzing the speaker's social media activity. Specific methods for analyzing social media activity include, but are not limited to, posting content, likes, and comment history. Some or all of the above processing in the learning unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the learning unit can input the speaker's social media activity into a generative AI and have the generative AI incorporate it as training data.

[0110] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0111] The analysis unit can estimate the user's emotions and adjust the method of analyzing the conversation content based on the estimated user emotions. For example, if the user is angry, the AI ​​will analyze the tone and wording of the speech with particular care. If the user is sad, the AI ​​can also focus on analyzing speech that includes emotional expressions. Furthermore, if the user is relaxed, the AI ​​can apply the normal analysis method and capture the broad context of the speech. This allows for more appropriate analysis by adjusting the method of analyzing the conversation content based on the user's emotions. User emotions include, but are not limited to, facial recognition, speech analysis, and text analysis. Adjusting the analysis method includes, but are not limited to, changing the analysis algorithm or adjusting the analysis parameters.

[0112] The analysis unit can improve the accuracy of its analysis by considering the context of the statements when analyzing conversation content. For example, it can analyze the context before and after a statement to determine whether a particular word constitutes harassment. It can also analyze the flow of the conversation to determine whether a statement is taken as a joke. Furthermore, it can consider the relationship between the speaker and the listener to analyze the intent of the statement. As a result, considering the context of the statements improves the accuracy of the analysis and enables more accurate harassment detection. Specific methods for considering context include, but are not limited to, the content of the preceding and following statements and related topics.

[0113] The analysis unit can analyze conversation content by referring to the speaker's past statements. For example, it can check whether the speaker has made similar statements in the past to determine the possibility of harassment. It can also analyze whether specific words or phrases are frequently used based on the speaker's past statements. Furthermore, it can analyze the intent and background of a statement based on the speaker's past statements. This allows for a more accurate analysis of the intent and background of a statement by referring to the speaker's past statements. Specific methods for referring to the statement history include, but are not limited to, past statements and their frequency.

[0114] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated user emotions. For example, if the user is nervous, it can provide a simple and highly visible display method. If the user is relaxed, it can provide a display method that includes detailed information. Furthermore, if the user is in a hurry, it can provide a display method that gets straight to the point. In this way, by adjusting the display method of the analysis results based on the user's emotions, it becomes possible to provide a display that is easy for the user to understand. User emotions include, but are not limited to, facial recognition, voice analysis, and text analysis. Adjustments to the display method include, but are not limited to, changing the display format or filtering the displayed content.

[0115] The analysis unit can perform analysis of conversation content while considering the speaker's geographical location. For example, if the speaker is in a specific region, the analysis will take into account the culture and customs of that region. It can also determine whether specific words or phrases are unique to a particular region based on the speaker's geographical location. Furthermore, it can analyze the background and intent of a statement based on the speaker's geographical location. This makes it possible to perform analysis that reflects the culture and customs unique to a region by considering the speaker's geographical location. Specific methods for obtaining geographical location information include, but are not limited to, GPS data and IP addresses.

[0116] The detection unit can estimate the user's emotions and adjust the detection criteria for potentially harassing remarks based on the estimated user emotions. For example, if the user is angry, the AI ​​will pay particular attention to the tone and wording of the remarks. If the user is sad, the AI ​​can also focus on detecting remarks that contain emotional expressions. Furthermore, if the user is relaxed, the AI ​​can apply normal detection criteria and capture a broader context of the remarks. This allows for more appropriate harassment detection by adjusting the detection criteria based on the user's emotions. User emotions include, but are not limited to, facial recognition, voice analysis, and text analysis. Adjusting the detection criteria includes, but are not limited to, changing thresholds or adjusting the detection algorithm.

[0117] The detection unit can improve detection accuracy by considering the frequency of remarks when detecting potentially harassing statements. For example, if a particular word or phrase is used frequently, it may be judged as likely to be harassment. It can also determine whether a particular word or phrase constitutes harassment based on its frequency of use. Furthermore, by considering the frequency of remarks, detection accuracy can be improved if a particular word or phrase is used repeatedly. As a result, by considering the frequency of remarks, detection accuracy is improved, enabling more accurate harassment detection. Specific methods for measuring the frequency of remarks include, but are not limited to, the number of remarks made within a certain period or the intervals between remarks.

[0118] The notification unit can estimate the user's emotions and adjust the way the warning message is delivered based on those emotions. For example, if the user is tense, the warning message can be sent in a calm tone. If the user is relaxed, the warning message can be sent in a friendly tone. Furthermore, if the user is in a hurry, a concise and quick warning message can be sent. By adjusting the way the warning message is delivered based on the user's emotions, a more appropriate message is delivered. User emotions include, but are not limited to, facial recognition, speech analysis, and text analysis. Adjustments to the delivery method include, but are not limited to, the tone and wording of the message.

[0119] The notification unit can adjust the level of detail in a warning message based on the importance of the statement when sending a warning message. For example, a detailed warning message can be sent for high-importance statements. Conversely, a concise warning message can be sent for low-importance statements. Furthermore, the level of detail can be adjusted based on the importance of the statement. This ensures that more appropriate warning messages are sent by adjusting the level of detail based on the importance of the statement. Specific criteria for evaluating the importance of a statement include, but are not limited to, the impact and content of the statement.

[0120] The registration unit can estimate the user's emotions and adjust the registration method for offensive remarks based on the estimated emotions. For example, if the user is angry, it can provide a simple interface for quick registration of remarks. If the user is sad, it can provide detailed input options to express their emotions. Furthermore, if the user is relaxed, it can provide a normal registration method that captures a broader context of the remarks. This allows for more appropriate registration by adjusting the registration method for offensive remarks based on the user's emotions. User emotions include, but are not limited to, facial recognition, speech analysis, and text analysis. Adjustments to the registration method include, but are not limited to, changing the registration procedure or adding registration items.

[0121] The following briefly describes the processing flow for example form 2.

[0122] Step 1: The analysis unit analyzes the conversation content in real time. The analysis unit uses natural language processing technology, machine learning models, and rule-based systems to analyze the conversation content and determine whether specific words or phrases constitute harassment. Step 2: The detection unit detects potentially harassing remarks from the conversation content analyzed by the analysis unit. The detection unit uses natural language processing technology, machine learning models, and rule-based systems to detect whether specific words or phrases constitute harassment. Step 3: The notification unit sends a message to the speaker regarding the statement detected by the detection unit. The notification unit can automatically send a message to the speaker regarding the statement detected by the detection unit, warning them that "that statement may be harassment." It can also send the message via voice or text message. Step 4: The registration unit registers comments that the user finds offensive in everyday life. The registration unit allows the victim to register comments that they find offensive in the system, and these can be registered via voice or text message. Step 5: The learning unit learns from the utterances registered by the registration unit to improve detection accuracy. The learning unit improves detection accuracy by inputting the utterances registered by the registration unit into machine learning models, natural language processing technologies, and rule-based systems, and learning from them.

[0123] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0124] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0125] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0126] Each of the multiple elements described above, including the analysis unit, detection unit, notification unit, registration unit, and learning unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the analysis unit is implemented by the processor 46 of the smart device 14 and analyzes the conversation content in real time. The detection unit is implemented by the identification processing unit 290 of the data processing unit 12 and detects potentially harassing remarks from the analyzed conversation content. The notification unit is implemented by the control unit 46A of the smart device 14 and sends a message to the speaker to draw their attention. The registration unit is implemented by the receiving device 38 of the smart device 14 and registers remarks that the victim finds offensive. The learning unit is implemented by the identification processing unit 290 of the data processing unit 12 and learns the registered remarks to improve detection accuracy. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0127] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0128] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0129] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0130] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0131] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0132] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0133] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0134] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0135] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0136] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0137] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0138] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0139] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0140] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0141] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0142] Each of the multiple elements described above, including the analysis unit, detection unit, notification unit, registration unit, and learning unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the analysis unit is implemented by the processor 46 of the smart glasses 214 and analyzes the conversation content in real time. The detection unit is implemented by the identification processing unit 290 of the data processing unit 12 and detects potentially harassing remarks from the analyzed conversation content. The notification unit is implemented by the control unit 46A of the smart glasses 214 and sends a message to the speaker to draw their attention. The registration unit is implemented by the microphone 238 of the smart glasses 214 and registers remarks that the victim finds offensive. The learning unit is implemented by the identification processing unit 290 of the data processing unit 12 and learns the registered remarks to improve detection accuracy. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0143] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0144] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0145] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0146] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0147] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0148] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0149] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0150] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0151] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0152] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0153] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0154] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0155] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0156] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0157] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0158] Each of the multiple elements described above, including the analysis unit, detection unit, notification unit, registration unit, and learning unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the analysis unit is implemented by the processor 46 of the headset terminal 314 and analyzes the conversation content in real time. The detection unit is implemented by the identification processing unit 290 of the data processing unit 12 and detects potentially harassing remarks from the analyzed conversation content. The notification unit is implemented by the control unit 46A of the headset terminal 314 and sends a message to the speaker to draw their attention. The registration unit is implemented by the microphone 238 of the headset terminal 314 and registers remarks that the victim finds offensive. The learning unit is implemented by the identification processing unit 290 of the data processing unit 12 and learns the registered remarks to improve detection accuracy. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0159] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0160] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0161] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0162] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0163] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0164] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0165] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0166] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0167] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0168] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0169] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0170] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0171] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0172] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0173] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0174] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0175] Each of the multiple elements described above, including the analysis unit, detection unit, notification unit, registration unit, and learning unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the analysis unit is implemented by the processor 46 of the robot 414 and analyzes the conversation content in real time. The detection unit is implemented by the identification processing unit 290 of the data processing unit 12 and detects potentially harassing remarks from the analyzed conversation content. The notification unit is implemented by the control unit 46A of the robot 414 and sends a message to the speaker to draw their attention. The registration unit is implemented by the microphone 238 of the robot 414 and registers remarks that the victim finds offensive. The learning unit is implemented by the identification processing unit 290 of the data processing unit 12 and learns the registered remarks to improve detection accuracy. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0176] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0177] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0178] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0179] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0180] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0181] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0182] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0183] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0184] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0185] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0186] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0187] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0188] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0189] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0190] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0191] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0192] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0193] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0194] (Note 1) An analysis unit that analyzes the conversation content in real time, A detection unit detects potentially harassing remarks from the conversation content analyzed by the aforementioned analysis unit, A notification unit sends a message to the speaker to alert them to the statement detected by the detection unit, A registration section for registering comments that you find offensive in everyday life, The system includes a learning unit that learns the content of speech registered by the registration unit and improves detection accuracy. A system characterized by the following features. (Note 2) The aforementioned analysis unit, The AI ​​will determine whether certain words or phrases in meetings or chat conversations constitute harassment. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned notification unit, The system automatically sends a message to the speaker to alert them to any detected statements. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned registration unit is Register any comments that the victim finds offensive in the system. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned learning unit, The registered content of statements will be used to learn from and improve the accuracy of future detections. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned analysis unit, It estimates the user's emotions and adjusts the conversation analysis method based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned analysis unit, When analyzing conversation content, consider the context of the statements to improve analysis accuracy. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned analysis unit, When analyzing conversation content, the analysis is performed by referring to the speaker's past statement history. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned analysis unit, It estimates the user's emotions and adjusts how the analysis results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned analysis unit, When analyzing conversation content, the geographical location information of the speakers is taken into consideration. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned analysis unit, When analyzing conversation content, the social media activity of the speaker is analyzed, and relevant statements are identified. The system described in Appendix 1, characterized by the features described herein. (Note 12) The detection unit is The system estimates the user's emotions and adjusts the detection criteria for potentially harassing remarks based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The detection unit is When detecting potentially harassing remarks, the detection accuracy is improved by considering the frequency of the remarks. The system described in Appendix 1, characterized by the features described herein. (Note 14) The detection unit is When detecting potentially harassing remarks, the system takes into account the speaker's attribute information. The system described in Appendix 1, characterized by the features described herein. (Note 15) The detection unit is It estimates the user's emotions and adjusts how the detection results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The detection unit is When detecting potentially harassing remarks, the system takes the speaker's geographical location into consideration. The system described in Appendix 1, characterized by the features described herein. (Note 17) The detection unit is When detecting potentially harassing remarks, we improve detection accuracy by referring to relevant literature about the speaker. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned notification unit, The system estimates the user's emotions and adjusts the way warning messages are presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned notification unit, When sending a warning message, adjust the level of detail based on the importance of the statement. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned notification unit, When sending a warning message, different message algorithms are applied depending on the category of the message. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned notification unit, The system estimates the user's emotions and adjusts the length of the warning message based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned notification unit, When sending a warning message, the message priority is determined based on when the comment was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned notification unit, When sending warning messages, the order of messages will be adjusted based on the relevance of the statements. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned registration unit is The system estimates the user's emotions and adjusts how offensive comments are registered based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned registration unit is When registering offensive remarks, the registration accuracy will be improved by considering the context of the remarks. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned registration unit is The system estimates the user's emotions and determines the priority of posts to register based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned registration unit is When registering offensive comments, the system prioritizes registering comments that are highly relevant, taking into account the speaker's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned learning unit, The system estimates the user's emotions and selects training data based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned learning unit, During training, the learning algorithm is optimized by referring to past training data. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned learning unit, During training, the training data is weighted considering the frequency of statements. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned learning unit, It estimates the user's emotions and adjusts the learning frequency based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 32) The aforementioned learning unit, During training, the training data is weighted based on when the comments were submitted. The system described in Appendix 1, characterized by the features described herein. (Note 33) The aforementioned learning unit, During training, we improve learning accuracy by referring to the speaker's relevant literature. The system described in Appendix 1, characterized by the features described herein. (Note 34) The aforementioned learning unit, During training, the learning algorithm is optimized by referring to the speaker's past speech history. The system described in Appendix 1, characterized by the features described herein. (Note 35) The aforementioned learning unit, During training, the training data is weighted considering the geographical location of the speaker. The system described in Appendix 1, characterized by the features described herein. (Note 36) The aforementioned learning unit, During training, the social media activity of the speaker is analyzed, and relevant statements are incorporated as training data. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0195] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. An analysis unit that analyzes the conversation content in real time, A detection unit detects potentially harassing remarks from the conversation content analyzed by the aforementioned analysis unit, A notification unit sends a message to the speaker to alert them to the statement detected by the detection unit, A registration section for registering comments that you find offensive in everyday life, The system includes a learning unit that learns the content of speech registered by the registration unit and improves detection accuracy. A system characterized by the following features.

2. The aforementioned analysis unit, The AI ​​will determine whether certain words or phrases in meetings or chat conversations constitute harassment. The system according to feature 1.

3. The aforementioned notification unit, The system automatically sends a message to the speaker to alert them to any detected statements. The system according to feature 1.

4. The aforementioned registration unit is Register any comments that the victim finds offensive in the system. The system according to feature 1.

5. The aforementioned learning unit, The registered content of statements will be used to learn from and improve the accuracy of future detections. The system according to feature 1.

6. The aforementioned analysis unit, It estimates the user's emotions and adjusts the conversation analysis method based on the estimated user emotions. The system according to feature 1.

7. The aforementioned analysis unit, When analyzing conversation content, consider the context of the statements to improve analysis accuracy. The system according to feature 1.

8. The aforementioned analysis unit, When analyzing conversation content, the analysis is performed by referring to the speaker's past statement history. The system according to feature 1.

9. The aforementioned analysis unit, It estimates the user's emotions and adjusts how the analysis results are displayed based on the estimated emotions. The system according to feature 1.

10. The aforementioned analysis unit, When analyzing conversation content, the geographical location information of the speakers is taken into consideration. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A