Information processing system, information processing method, and program

The information processing system addresses the challenge of online defamation by using a control unit to analyze messages for inappropriate content and manage their display, enhancing content moderation and user safety through AI-powered analysis.

WO2025109795A1PCT designated stage expired Publication Date: 2025-05-30SASANO KENTO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/024654
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-24
Filing Date
2024-07-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing technologies face challenges in effectively handling and mitigating online defamation, as determining inappropriate or problematic statements is subjective and requires advanced methods to manage online content.

Method used

An information processing system that includes a control unit capable of acquiring messages from social networking services, determining whether they contain inappropriate or problematic statements, and executing processes to hide or display messages based on these determinations, utilizing natural language models for analysis.

Benefits of technology

The system effectively prevents inappropriate messages from being displayed, enhances content moderation by using AI for accurate detection, and improves user experience by maintaining a safer online environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024024654_30052025_PF_FP_ABST
    Figure JP2024024654_30052025_PF_FP_ABST
Patent Text Reader

Abstract

According to one embodiment of the present invention, provided is an information processing system. The information processing system has at least one control unit. The control unit acquires a message posted from a user, determines whether or not the acquired message includes an inappropriate or problematic statement, and executes processing for displaying the message that does not include the inappropriate or problematic statement.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, information processing method and program

[0001] The present invention relates to an information processing system, an information processing method, and a program.

[0002] Patent Document 1 discloses a labeling technology related to a learning system that detects fraudulent behavior such as posting slanderous comments against others.

[0003] Patent No. 7302107

[0004] How to deal with posts that slander others is a subjective issue, and new technologies are needed to address these issues and deal with slander on the Internet.

[0005] According to one aspect of the present invention, there is provided an information processing system including at least one control unit that executes a process of acquiring a message posted on a social networking service, determining whether the acquired message contains inappropriate or problematic content, and displaying the message that does not contain the inappropriate or problematic content.

[0006] 1 is a diagram illustrating an example of the system configuration of an information processing system. FIG. 2 is a diagram illustrating an example of the hardware configuration of a client device. FIG. 3 is a diagram illustrating an example of the hardware configuration of a server device. FIG. 4 is a flowchart illustrating an example of information processing in a client device. FIG. 5 is an example of an instruction statement requesting output of a confirmation result as to whether a message contains inappropriate or problematic discourse. FIG. 6 is a diagram illustrating an example of a screen for selecting a category of messages to hide. FIG. 7 is a diagram illustrating an example of a selection screen for hiding messages when the messages violate the policies of a company that provides AI such as ChatGPT.

[0007] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described below with reference to the accompanying drawings. Various features shown in the following embodiments can be combined with each other.

[0008] In particular, in this specification, the term "unit" may include, for example, a combination of hardware resources implemented by a circuit in the broad sense and software information processing that can be specifically realized by these hardware resources. Furthermore, while various types of information are handled in the first embodiment, communication and calculation can be performed on a circuit in the broad sense, regardless of whether the information is represented by a high or low signal value as a binary bit collection consisting of 0 or 1, a physical numerical value of a signal value, or a quantum superposition.

[0009] Furthermore, a circuit in the broad sense is a circuit realized by at least an appropriate combination of a circuit, circuitry, a processor, a memory, etc. That is, it includes an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)), etc.

[0010] In addition, the program for realizing the software appearing in the embodiments may be implemented in a manner that allows it to be downloaded from a server, the program may be executed on a cloud computer, or it may be stored on a non-volatile or volatile non-transitory storage medium and distributed.

[0011] <First Embodiment> 1. System Configuration Fig. 1 is a diagram showing an example of the system configuration of an information processing system 1000. As shown in Fig. 1, the information processing system 1000 includes, as a system configuration, a client device 110, a server device 100, and an AI system 120. The client device 110 and the server device 100 are connected to each other so as to be able to communicate with each other via a network 150. Furthermore, the server device 100 and the AI ​​system 120 are connected to each other so as to be able to communicate with each other via the network 150.

[0012] In response to a request from the client device 110, the server device 100 transmits to the client device 110 a program that provides an extended function to a specific web browser of the client device 110. The client device 110 is a device operated by a user, logs in to an SNS (social networking service; for example, X (formerly Twitter) or Facebook; not shown), posts messages on the SNS, and displays messages posted by other users. The AI ​​system 120 is a system that provides large-scale natural language model (LLM) functions. In this specification, the AI ​​system 120 is described as an example of an API (Application Programming Interface) server for ChatGPT provided by OpenAI, Inc. However, this does not limit the embodiment. For example, the AI ​​system 120 that provides the API may be operated independently, or models such as OpenAI's GPT (GPT3.5, GPT4, etc.), Meta's Llama2, and Google's PaLM2 may be managed on the server device 100 to perform inference directly.

[0013] 1, for the sake of simplicity, only one client device 110 is shown in the information processing system 1000. However, the information processing system 1000 may include as many client devices 110 as there are users.

[0014] 1 shows a PC (Personal Computer) as an example of the client device 110, the client device 110 is not limited to a PC (Personal Computer) and may be a smartphone or a tablet computer. In other words, the client device 110 may be any device that has a predetermined web browser installed and is capable of posting and displaying messages on an SNS.

[0015] Here, the claimed information processing system may be composed of multiple devices or may be composed of a single device. When the claimed information processing system is composed of a single device, an example of that device is client device 110.

[0016] 2. Hardware Configuration (1) Hardware Configuration of Client Device 110 Figure 2 is a diagram showing an example of the hardware configuration of the client device 110. The client device 110 includes, as its hardware configuration, a control unit 210, a storage unit 220, an input unit 230, an output unit 240, and a communication unit 250. The control unit 210 is a central processing unit (CPU) or the like, and controls the entire client device 110. The storage unit 220 is any one of a hard disk drive (HDD), a read-only memory (ROM), a random access memory (RAM), a solid state drive (SSD), or any combination thereof, and stores programs and data used when the control unit 210 executes processing based on the programs. The control unit 210 executes a program stored in the storage unit 220 to realize each process and each function of the client device 110, which will be described later. For simplicity's sake, the following description will be given assuming that the control unit 210 executes the processes described herein. The program stored in the storage unit 220 may be provided by the server device 100. The client device 110 accesses the server device 100, and the program provided by the server device 100 may be stored in the storage unit 220, for example. For example, the client device 110 sends a request to the server device 100, the server device 100 sends HTML including a link to a program written in JavaScript, the client device 110 accesses the URL indicated by the link, and the server device 100 responds with a program corresponding to the URL, thereby enabling the client device 110 to acquire the program. The input unit 230 is a keyboard and / or a mouse, etc., and inputs user operations to the control unit 210. The input unit 230 may further include a microphone and a speaker, etc., and input the user's voice to the control unit 210. The output unit 240 may include a display, etc., and displays the results of information processing by the control unit 210 and / or data acquired from other devices.The communication unit 250 is a device such as a network interface card (NIC), which connects the client device 110 to the network 150 and controls communication with other devices.

[0017] (2) Hardware Configuration of Server Device 100 FIG. 3 is a diagram illustrating an example of the hardware configuration of the server device 100. The server device 100 includes, as its hardware configuration, a control unit 310, a storage unit 320, and a communication unit 330. The control unit 310 is a CPU or the like, and controls the entire server device 100. The storage unit 320 is any one of an HDD, ROM, RAM, SSD, etc., or any combination thereof, and stores programs and data used when the control unit 310 executes processes based on the programs. The control unit 310 executes processes based on the programs stored in the storage unit 320, thereby realizing the functions of the server device 100. The communication unit 330 is a NIC or the like, and connects the server device 100 to the network 150 and manages communication with other devices. In this embodiment, the data used by the control unit 310 when executing processing based on a program is described as being stored in the memory unit 320, but the data may also be stored in the memory unit of another device with which the server device 100 can communicate.

[0018] 3. Information Processing The information processing of this embodiment will now be described. If a message posted to an SNS contains inappropriate or problematic content, the control unit 210 performs processing to prevent the message from being displayed. Note that messages posted to an SNS include not only posts that are publicly available to an unspecified number of users or posts that are privately available to a number of specific users, but also direct messages (messages exchanged between one user and another).

[0019] (1) Message Acquisition Process The control unit 210 acquires messages posted to the SNS. The control unit 210 may acquire messages before they are posted to the SNS, or may acquire messages after they have been posted to the SNS. The control unit 210 may accept messages from users and post them to the SNS, or may monitor the SNS to detect posted messages. The control unit 210 may also accept notifications of newly posted messages from the SNS. The control unit 210 may acquire only messages from the user operating the client device 110, or may acquire all messages that the user can view.

[0020] (2) Message Hiding Process When a message is acquired after it has been posted on the SNS, the control unit 210 can execute a process of hiding the newly posted message on the SNS. Note that the hiding process may be omitted. Also, when a message is acquired before it has been posted on the SNS, the hiding process may be omitted. Also, the control unit 210 may execute a process of hiding the newly posted message on the SNS only when a setting to turn on a predetermined operation mode is accepted. Note that the control unit 210 can hide the newly posted message by setting to hide new posts on the SNS.

[0021] (3) Process for Determining Whether a Message Contains Inappropriate or Problematic Content The control unit 210 determines whether a message contains inappropriate or problematic content.

[0022] The control unit 210 can determine, for example, whether a message contains a preset inappropriate or problematic keyword or phrase.

[0023] The control unit 210 can also use the LLM to determine whether a message contains inappropriate or problematic speech. For example, the control unit 210 can determine whether a message contains inappropriate or problematic speech by calling a Moderation API provided by OpenAI, Inc. The control unit 210 can obtain a response using the LLM as to whether a message contains inappropriate or problematic speech by sending a request including a message to an endpoint of the Moderation API of the AI ​​system 120. For example, if the return value "flagged" from the Moderation API is true, the control unit 210 can determine that the message contains inappropriate or problematic speech.

[0024] Here, the control unit 210 can accept a category designation via the screen. A category is a classification related to inappropriateness or problematic content, such as sexual, hate, harassment, violence, etc., and indicates the content of a SNS message. The response from the Moderation API includes a flag value indicating whether the message falls into each category. If the flag value corresponding to the category designated by the user is true, the control unit 210 may determine that the message contains inappropriate or problematic content (if all flag values ​​corresponding to the categories designated by the user are false, the control unit 210 may determine that the message does not contain inappropriate or problematic content, even if the flag values ​​corresponding to other categories are true).

[0025] Furthermore, for example, the control unit 210 may infer whether a message contains inappropriate or problematic speech by providing a prompt to the LLM in response to the message. The prompt may include the acquired message and a command statement instructing the LLM to predict whether the message or multiple character strings contained in the message contain inappropriate or problematic speech. The prompt may also include a definition of "inappropriate or problematic." The control unit 210 may send the prompt to an endpoint of an API provided by the AI ​​system 120 (a so-called ChatGPT API endpoint), and receive a response obtained by providing the prompt to the LLM in the AI ​​system 120.

[0026] (4) Message Display Process The control unit 210 performs a process of displaying messages that do not contain inappropriate or problematic statements on the SNS.

[0027] When the control unit 210 obtains a message before it is posted on the SNS, if it determines that the message contains inappropriate or problematic language, it will not post the message on the SNS, and if the message does not contain inappropriate or problematic language, it will post the message on the SNS.

[0028] The control unit 210 acquires a message after it has been newly registered on the SNS, and when the message is hidden, if it determines that the message contains inappropriate or problematic discourse, it can perform a process of keeping the message hidden, and changing the message from hidden to displayed if it determines that the message does not contain inappropriate or problematic discourse.

[0029] When the control unit 210 determines that a message posted on an SNS contains inappropriate or problematic statements, the control unit 210 can process the message to delete the message.

[0030] (5) Operation Mode The control unit 210 can accept a setting to turn on a predetermined operation mode via a screen. The screen can be a screen that displays a timeline of a social networking service. The control unit 210 can perform at least one of the following processes only when the predetermined operation mode is turned on: (1) message acquisition processing, (2) message hiding processing, (3) message inappropriateness or problem determination processing, and (4) message display processing.

[0031] (6) Specific Example FIG. 4 is a flowchart showing an example of information processing in the client device 110.

[0032] The first step is to connect to the X (formerly Twitter) API and temporarily hide newly posted tweets initially. New tweets are hidden from users until they are evaluated via the Moderation API, which is based on the OpenAI API. Next, individual processing of tweets begins. Each hidden tweet is processed individually for evaluation by the Moderation API. This API analyzes the content of the tweet to determine whether it contains inappropriate or problematic content.

[0033] If the response from the Moderation API indicates that the tweet contains problematic content, the tweet will remain hidden. If the tweet is deemed to be non-problematic, it will be made public and available for viewing by users.

[0034] Finally, the occurrence of new tweets is monitored, and if a new tweet is detected, the process returns to the initial filtering phase and the above process is repeated.

[0035] FIG. 5 shows an example of a prompt requesting output of the results of checking whether a message contains inappropriate or problematic language.

[0036] The prompt shown in Figure 5 has the following structure: (1) clearly explain the task; (2) define what is inappropriate or problematic ("toxic" in the context); (3) present the sentence to be evaluated; and (4) ask the model to make a judgment based on the explanation and definition provided.

[0037] Fig. 6 is a diagram showing an example of a screen for selecting a category related to a message to be hidden. When a predetermined operation is performed on the screen of the user's SNS account, etc., the control unit 210 displays a screen such as that shown in Fig. 6. The same applies to Fig. 7, which will be described later. The category is each category included in the return value of the Moderation API.

[0038] 7 is a diagram showing an example of a selection screen for hiding a message if it violates the policy of a company that provides AI such as ChatGPT (if the return value "flagged" of the Moderation API is true). Note that if all the boxes in FIGS. 6 and 7 are unchecked, the process of hiding messages that contain inappropriate or problematic statements can be prevented.

[0039] According to this embodiment, messages containing inappropriate or problematic statements can be prevented from being displayed on SNS. Furthermore, messages containing inappropriate or problematic statements can be deleted from SNS. Furthermore, messages containing inappropriate or problematic statements can be prevented from being posted on SNS.

[0040] Using 500 known inappropriate or problematic (toxic) texts and 500 known inappropriate or problematic (not toxic) texts, the results of the Moderation API (first model) were compared with the results of querying ChatGPT using prompts (second model). The accuracy rate was improved to 0.64 for the first model and 0.68 for the second model. The precision rate was 1.00 for the first model and 0.96 for the improved model. The recall rate was improved to 0.31 for the first model and 0.42 for the second model. The F1 score was improved to 0.47 for the first model and 0.58 for the second model.

[0041] In this embodiment, messages posted on SNS are assumed, but the present invention is not limited to this and can be applied to any message posted by a user. For example, it is possible to determine whether comments posted to videos on a video sharing site contain inappropriate or problematic statements, and post only those comments that do not contain inappropriate or problematic statements, or to hide posted comments and display those that do not contain inappropriate or problematic statements.

[0042] It can also be applied to messages such as emails and chats, and before sending an email or chat message, it can be determined whether the message contains inappropriate or problematic language, and a message that does not contain inappropriate or problematic language can be sent. It can also be configured to hide messages received via email or chat, and display messages that do not contain inappropriate or problematic language.

[0043] <Additional Notes> The invention may be provided in the following aspects. [Item 1] An information processing system having at least one control unit, the control unit acquiring a message posted on a social networking service, determining whether the acquired message contains inappropriate or problematic speech, and executing a process to display the message that does not contain the inappropriate or problematic speech. [Item 2] The information processing system according to item 1, wherein the control unit acquires the message after it has been posted on the social networking service but before it is displayed, and determines whether the acquired message contains the inappropriate or problematic speech. [Item 3] The information processing system according to item 1, wherein the control unit executes a process to hide the message posted on the social networking service, and changes the message that does not contain the inappropriate or problematic speech from hidden to displayed. [Item 4] The information processing system according to item 1, wherein the control unit generates an instruction statement requesting output of a confirmation result as to whether the message contains inappropriate or problematic discourse, inputs the instruction statement to a predetermined large-scale natural language model, and determines whether the message contains inappropriate or problematic discourse based on output from the large-scale natural language model. [Item 5] The information processing system according to item 4, wherein the control unit inputs the instruction statement to the large-scale language model, including the message, a definition of inappropriate or problematic, and an instruction to determine whether text included in the message is inappropriate or problematic. [Item 6] An information processing method comprising: acquiring a message posted on a social networking service; determining whether the acquired message contains inappropriate or problematic discourse; and executing a process to display the message including the inappropriate or problematic discourse.[Item 7] A program for causing a computer to execute the steps of: acquiring a message posted on a social networking service; determining whether the acquired message contains inappropriate or problematic speech; and executing a process to display the message containing the inappropriate or problematic speech. For example, the program may be provided as a computer-readable non-transitory storage medium that stores the program.

[0044] Finally, while various embodiments of the present invention have been described, these are presented by way of example only and are not intended to limit the scope of the invention. The novel embodiments may be embodied in various other forms, and various omissions, substitutions, and modifications may be made without departing from the spirit of the invention. The embodiments and their modifications are intended to be included within the scope and spirit of the invention, as well as within the scope of the inventions and their equivalents as defined in the appended claims.

[0045] For example, in this embodiment, when a message is posted on an SNS, it is determined whether the message contains inappropriate or problematic statements. However, it may also be determined for messages posted on the SNS in the past. In this case, when displaying a list of messages posted on the SNS in the past, it may be determined whether each message is inappropriate or problematic, and only inappropriate or problematic messages may be displayed. Inappropriate or problematic messages may be deleted from the SNS or set to not be displayed. This makes it possible, for example, to prevent minors from viewing inappropriate or problematic messages.

[0046] In this embodiment, when a prompt is used to inquire whether an LLM is inappropriate or problematic, the prompt simply asks whether it is inappropriate or problematic. However, it may also be possible to inquire whether the LLM falls into each category determined by the Moderation API (especially a category specified by the user). In this case, the prompt should specifically state the definition of each category.

[0047] 100: Server device 110: Client device 120: AI system 150: Network 210: Control unit 220: Storage unit 230: Input unit 240: Output unit 250: Communication unit 310: Control unit 320: Storage unit 330: Communication unit 1000: Information processing system

Claims

1. An information processing system having at least one control unit that executes a process of acquiring messages posted by users, determining whether the acquired messages contain inappropriate or problematic speech, and displaying the messages that do not contain the inappropriate or problematic speech.

2. An information processing system as described in claim 1, wherein the control unit acquires the message after it is posted by the user and before it is displayed, and determines whether the acquired message contains the inappropriate or problematic discourse.

3. An information processing system as described in claim 1, wherein the control unit executes a process of hiding the messages posted by the users, and changes the messages that do not contain the inappropriate or problematic discourse from hidden to displayed.

4. An information processing system as described in claim 1, wherein the control unit generates an instruction statement requesting output of a confirmation result as to whether the message contains inappropriate or problematic discourse, inputs the instruction statement into a predetermined large-scale natural language model, and determines whether the message contains inappropriate or problematic discourse based on the output from the large-scale language model.

5. An information processing system as described in claim 4, wherein the control unit inputs the instruction sentence including the message, a definition of being inappropriate or problematic, and an instruction to determine whether the text contained in the message is inappropriate or problematic into the large-scale language model.

6. An information processing method comprising: acquiring a message posted on a social networking service; determining whether the acquired message contains inappropriate or problematic speech; and executing a process of displaying the message including the inappropriate or problematic speech.

7. A program for causing a computer to execute the steps of: acquiring a message posted by a user; determining whether the acquired message contains inappropriate or problematic speech; and executing a process to display the message containing the inappropriate or problematic speech.

Citation Information

Patent Citations

  • Contributed data evaluation device

    JP2006268304A