Information processing apparatus, information processing method, and information processing program
The Human-in-the-Loop reinforcement learning approach enhances fraud detection systems by using evaluator feedback to retrain models, addressing the limitations of conventional methods and improving detection accuracy and adaptability.
Patent Information
- Application Number
- JP2024081728
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-12-03
AI Technical Summary
Conventional fraud detection systems struggle with improving the detection performance of learning models due to challenges in generating accurate labeled data, distinguishing fraudulent from legitimate behavior, and the evolving nature of fraudulent activities, leading to decreased model performance and trust issues.
An information processing system incorporating Human-in-the-Loop (HITL) reinforcement learning, where evaluators provide feedback on suspected fraudulent listings, allowing the learning model to be retrained using their evaluation results, thereby enhancing the model's accuracy and adaptability.
The system significantly improves fraud detection performance by leveraging human expertise to update the learning model, ensuring higher accuracy and responsiveness to new fraudulent patterns, while maintaining transparency and trust.
Smart Images

Figure 2025175557000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]
[0002] In recent years, fraud detection systems have become increasingly important due to the increasing number of fraudulent listings that violate terms of service on online services that allow individuals to buy and sell goods (e.g., auctions and flea markets). Traditionally, supervised learning-based approaches to fraud detection have been widely used. However, creating labeled data, which is essential for training supervised learning models, can be difficult. For example, in the context of fraud detection, it can be difficult to accurately distinguish between fraudulent and legitimate behavior. In such cases, inaccurate labels can significantly degrade the model's performance.
[0003] Furthermore, a technology for training a learning model using human advice is known (see Non-Patent Document 1 below). This technology makes it possible to provide a system that combines a learning model with a machine learning concept called Human-in-the-Loop (HITL), which can reflect human advice. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] David Abel, John Salvatier, Andreas Stuhlmueller, Owain Evans, “Agent-Agnostic Human-in-the-Loop Reinforcement Learning”, [online], [Retrieved April 22, 2024], Internet<https: / / ar5iv.labs.arxiv.org / html / 1701.04079> Summary of the Invention [Problem to be solved by the invention]
[0005] However, conventional techniques have not been able to promote improvements in the detection performance of learning models for detecting fraudulent items for auction, for example.
[0006] The present application has been made in consideration of the above, and aims to promote improvement in the detection performance of a learning model for detecting fraud in items for sale. [Means for solving the problem]
[0007] The information processing device of the present application is characterized by having an acquisition unit that acquires the evaluation results of evaluators regarding fraud against items for sale selected by a learning model, and a learning unit that re-trains the learning model using the evaluation results acquired by the acquisition unit as training data. [Effects of the Invention]
[0008] According to one aspect of the embodiment, it is possible to promote improvement in the detection performance of a learning model for detecting fraud in items for sale. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of an information processing system according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of information processing according to the embodiment. [Figure 3] FIG. 3 is a diagram showing an example of the results of a comparative experiment between the case where the method according to the embodiment is used and the case where a conventional method is used. [Figure 4] FIG. 4 is a diagram illustrating an example of the configuration of a terminal device according to the embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of the configuration of a listing information providing server according to the embodiment. [Figure 6] FIG. 6 is a diagram illustrating an example of a selling information storage unit according to the embodiment. [Figure 7]FIG. 7 is a diagram illustrating an example of the configuration of an information processing device according to the embodiment. [Figure 8] FIG. 8 is a diagram illustrating an example of a learning model storage unit according to the embodiment. [Figure 9] FIG. 9 is a diagram illustrating an example of the evaluator information storage unit according to the embodiment. [Figure 10] FIG. 10 is a flowchart illustrating an example of a procedure for information processing according to the embodiment. [Figure 11] FIG. 11 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of the information processing device. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, an information processing device, an information processing method, and an information processing program according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the information processing device, the information processing method, and the information processing program according to the present application are not limited to these embodiments. Furthermore, the same components in the following embodiments will be denoted by the same reference numerals, and duplicated descriptions will be omitted.
[0011] (Embodiment) [1. Information Processing System Configuration] An information processing system 1 shown in Fig. 1 will be described. As shown in Fig. 1, the information processing system 1 includes a terminal device 10, an information processing device 100, and a listing information providing server 200. The terminal device 10, the information processing device 100, and the listing information providing server 200 are connected to each other via a predetermined communication network (network N) so as to be able to communicate with each other via wired or wireless communication. Fig. 1 is a diagram showing an example of the configuration of the information processing system 1 according to an embodiment.
[0012] The terminal device 10 is an information processing device used by an evaluator (e.g., an examiner) who performs evaluations for fraud detection. The evaluator who uses the terminal device 10 has a predetermined relationship (e.g., an employee) with an operator of a service that allows online transactions between individuals, and detects fraudulent listings that violate the terms of service on such a service. The evaluator who uses the terminal device 10, for example, evaluates whether a listing violates the terms of service. For example, the evaluator who uses the terminal device 10 detects fraudulent listings from a certain group of detection targets. For example, the evaluator who uses the terminal device 10 evaluates whether each detection target included in a certain detection target group is a fraudulent listing and labels the detection target as fraudulent or not. The evaluator who uses the terminal device 10 provides the evaluation result to the information processing device 100 by performing such labeling. Furthermore, for example, the evaluator who uses the terminal device 10 may have a predetermined specialty for fraud detection and may be selected according to the category of fraudulent behavior or the category of detection target. The terminal device 10 may be any device that can implement the processing described in the embodiment. The terminal device 10 may also be a smartphone, a tablet terminal, a notebook PC, a desktop PC, a mobile phone, a PDA, etc. Fig. 2 shows a case where the terminal device 10 is a smartphone.
[0013] The terminal device 10 is, for example, a smart device such as a smartphone or tablet, and is a portable terminal device that can communicate with any server device via a wireless communication network such as 4G to 5G (Generations) or LTE (Long Term Evolution). The terminal device 10 may have a screen such as a liquid crystal display with a touch panel function, and may accept various operations on display data such as advertisements, such as tapping, sliding, and scrolling, performed by an evaluator using a finger or a stylus. In FIG. 2, the terminal device 10 is used by an evaluator U1.
[0014] The information processing device 100 is an information processing device intended to promote improvement in the detection performance of a learning model for detecting fraud in items for sale, and may be any device capable of implementing the processes described in the embodiments. For example, the information processing device 100 is an information processing device for providing a fraud detection system based on reinforcement learning incorporating Human-in-the-Loop (HITL), and aims to improve the accuracy of fraud detection by utilizing the knowledge of evaluators to perform continuous learning and improvement, while also building a system that is flexible enough to quickly respond to new fraudulent activities. The information processing device 100 is realized, for example, by a server device or cloud system of an operator of a service that enables person-to-person buying and selling over the Internet.
[0015] The listing information providing server 200 is an information processing device intended to provide listing information (which may include any information related to listings, such as information about the seller himself or herself or a list of items to be listed) on a service that allows for buying and selling between individuals over the Internet, and may be any device that can realize the processing in the embodiment. The listing information providing server 200 is realized by a server device or a cloud system of an operator that operates such a service.
[0016] Although FIG. 1 shows a case where the information processing device 100 and the selling information providing server 200 are separate devices, the information processing device 100 and the selling information providing server 200 may be integrated.
[0017] [2. An example of information processing] In recent years, fraud detection systems have become increasingly important due to the increasing number of fraudulent listings that violate terms of service on services that allow online transactions between individuals. Traditionally, supervised learning-based approaches to fraud detection have been widely used. However, generating labeled data, which is essential for training supervised learning models, can be challenging. For example, it can be difficult to accurately distinguish between fraudulent and legitimate behavior in the context of fraud detection. In such cases, inaccurate labeling can significantly degrade the model's performance.
[0018] In addition, assigning accurate labels to fraud detection is often difficult. Fraudulent behavior is diverse and constantly evolving, so generating labels based on past data can lead to missing new fraudulent patterns. Furthermore, because fraudulent behavior occurs less frequently than normal behavior, it is often difficult to build a balanced training dataset.
[0019] Furthermore, while automation using machine learning models holds great promise for fraud detection systems, the predictive values output by the learning models can be difficult to explain. Especially with complex models, it can be difficult for humans to understand how a particular output was obtained and interpret the results. This lack of explainability can lead to a decrease in trust and increase the opacity of human intervention in final decision-making.
[0020] This application has been made in light of the above, and proposes a fraud detection system based on reinforcement learning that incorporates Human-in-the-Loop (HITL), with the aim of improving the accuracy of fraud detection by utilizing the knowledge of evaluators to carry out continuous learning and improvement, and building a system that is flexible enough to respond quickly to new fraudulent acts.
[0021] 2 is a diagram illustrating an example of information processing of the information processing system according to the embodiment. The information processing device 100 acquires listing information from the listing information providing server 200 (step S101). For example, the information processing device 100 acquires a list of listings that are being offered for sale on a predetermined service. At this time, for example, the information processing device 100 may acquire a list of listings that have been newly put up for sale within a predetermined period. In FIG. 2, the information processing device 100 acquires a list L1.
[0022] The information processing device 100 uses the learning model M1 to select a listing target suspected of fraud from the list L1 (step S102). Specifically, the information processing device 100 inputs the listing target information from the list L1 into the learning model M1, causing the learning model M1 to select a listing target. When the information processing device 100 selects a listing target, it transmits selection information regarding the selection of the listing target to the terminal device 10, thereby providing the selection information to the evaluator U1 (step S103). When the terminal device 10 receives the selection information from the information processing device 100, it displays the selection information. Note that product information of the selected listing target may be displayed, or a list of the selected listing targets may be displayed.
[0023] The evaluator U1 is a judge who evaluates whether an item corresponds to a fraudulent listing and labels the fraudulent listing as fraudulent. The evaluator U1 evaluates whether each listing target selected by the information processing device 100 corresponds to a fraudulent listing (step S11). The terminal device 10 transmits the evaluation result of the evaluator U1 to the information processing device 100. The information processing device 100 then acquires the evaluation result of the evaluator U1 transmitted from the terminal device 10 (step S104). The evaluation result includes the number of fraudulent items evaluated as corresponding to fraudulent listings. For example, if the number of listing targets selected by the learning model M1 is 100 and the number of fraudulent items in the evaluation result is 50, information indicating that the number of fraudulent items is 50 is included.
[0024] Furthermore, the evaluator U1 requests the listing information providing server 200 to delete the listing of the listing that he / she has evaluated as a fraudulent listing, as a listing that is suspected of being fraudulent (step S12). When the listing information providing server 200 receives the listing deletion request from the terminal device 10, it deletes the listing that has been requested to be deleted from the listings. At this time, the listing information providing server 200 may notify the seller of the deleted listing that the listing has been deleted. The notification may include information indicating the type of fraudulent activity.
[0025] When the information processing device 100 acquires the evaluation result of the evaluator U1, it retrains the learning model M1 using the acquired evaluation result as training data (step S105). For example, when the information processing device 100 reselects items to be sold using the learning model M1, it retrains the learning model M1 so that the number of fraudulent items is greater than when the learning model M1 before retraining was used, i.e., so that the items selected by the learning model M1 include more fraudulent items. Also, for example, the information processing device 100 identifies and retrains feature quantities (vectors) common to the items to be sold that have been evaluated as fraudulent. For example, if the items to be sold that have been evaluated as fraudulent have specific drug ingredients or efficacy or the like described in the product section, the information processing device 100 retrains the learning model M1 based on the description of such drug ingredients or efficacy or the like.
[0026] Below, a specific example will be described using experimental results. For example, assume that there are 10,000 users in total, of which 1,000 are fraudulent users. Each user has 10 features that range from 0 to 1, and users for which feature values 1 to 9 are 1 are considered fraudulent users. The task of selecting 1,000 users from the total users is performed 1,000 times (1,000 episodes). Then, the method of selecting 1,000 users is compared between a conventional method of randomly selecting users and a method according to an embodiment using reinforcement learning. Fraudulent users are excluded from the selected 1,000 users, and new fraudulent users are added to make up for the excluded users. Therefore, the total number of users and the number of fraudulent users remain unchanged. The comparison is made based on the number of fraudulent users included among the selected 1,000 users. Regarding the reward for reinforcement learning, the number of fraudulent users is used as the reward. FIG. 3 is a diagram showing an example of the results of an experiment comparing a case where the method according to an embodiment is used with a case where the conventional method is used. In FIG. 3, the vertical axis represents the moving average, and the horizontal axis represents the time series. Graph G1 shows the relationship between the moving average and time series when using the conventional method and randomly selecting some users from all users. Since 1,000 users are randomly selected from 10,000 users, 10% of them are fraudulent users, so the expected number of fraudulent users is 100. Therefore, the moving average of graph G1 is constant at approximately 100. On the other hand, graph group GN shows the relationship between the moving average and time series when using the method according to the embodiment. The results show that using the method according to the embodiment results in a higher number of fraudulent users among the selected users than when randomly selecting some users from all users. In other words, it shows that fraud detection performance is better.
[0027] (Information processing variation 1: Using evaluation results from multiple evaluators) In the above embodiment, a case has been described in which a single evaluator, evaluator U1, evaluates the selection target using the learning model M1, and the learning model M1 is retrained using the evaluation results of the single evaluator, evaluator U1. However, the number of evaluators who evaluate the selection target using the learning model M1 is not particularly limited. Although not shown in FIG. 2, evaluations may be received from multiple evaluators, such as evaluator U2 and evaluator U3, and the learning model M1 may be retrained using the evaluation results of the multiple evaluators. For example, the learning model M1 may be retrained using an average of the evaluations of the multiple evaluators or the most significant evaluation among the evaluations of the multiple evaluators as the evaluation result.
[0028] For example, if evaluator U1 evaluates the listing F1 as "fraudulent," evaluator U2 evaluates it as "not fraudulent," and evaluator U3 evaluates it as "fraudulent," the majority of evaluations are "fraudulent," and the learning model M1 may be retrained with the listing F1 as a fraudulent target. On the other hand, if evaluator U1 evaluates the listing F1 as "fraudulent," evaluator U2 evaluates it as "not fraudulent," and evaluator U3 evaluates it as "not fraudulent," the majority of evaluations are "not fraudulent," and the listing F1 may be excluded from the list of fraudulent targets and the learning model M1 may be retrained. Furthermore, for example, the learning model M1 may be retrained for each evaluator of multiple evaluators, or for each group of evaluators. For example, the learning model M1 may be retrained for each group of evaluators who specialize in specific fraudulent activities.
[0029] Furthermore, the learning model M1 may be retrained using only the “number of fraudulent items” without acquiring information indicating whether or not each auction item is fraudulent, such as whether it is “fraudulent” or “not fraudulent.” That is, the learning model M1 may be retrained using information such as “how many of the items evaluated by the evaluator are fraudulent” rather than “which items are fraudulent” or “this item is fraudulent for this reason.” For example, the information processing device 100 may acquire evaluation results from multiple evaluators for a predetermined group of evaluation items and retrain the learning model M1 using the average number of fraudulent items. In this case, the information processing device 100 may retrain the learning model M1 by calculating an average value that takes into account, for example, the expertise and weighting of the evaluators’ evaluations. For example, the information processing device 100 may retrain the learning model M1 by weighting the evaluation of an evaluator with higher expertise more heavily than that of an evaluator with lower expertise. Furthermore, when multiple evaluations of "fraudulent" are available, such as "100% fraudulent" and "80% fraudulent," the learning model M1 may be retrained by calculating the average value taking into account the weighting of each evaluation.
[0030] (Information processing variation 2: Using evaluation results by professional evaluators) In the above embodiment, the learning model M1 is retrained without taking into account the evaluator U1's expertise (the specialty of the evaluator U1). However, the learning model M1 may be retrained taking into account the evaluator U1's expertise. For example, specialized evaluators may be pre-defined for each type of fraudulent activity, and a highly specialized evaluator may be identified for each type of fraudulent activity. The identified evaluator may then evaluate the fraudulent activity, and the learning model M1 may be retrained based on the evaluation results. For example, if the auction item F1 selected by the learning model M1 is related to medicine, the information processing device 100 may identify an evaluator who specializes in medicine and request him or her to evaluate the auction item F1. Furthermore, the learning model M1 may be a learning model generated for each type of fraudulent activity. The information processing device 100 may use the evaluation results of an evaluator who specializes in medicine to retrain a learning model for detecting fraudulent activity involving drugs.
[0031] (Information processing variation 3: Using multiple learning models for each type of fraudulent behavior) In the above embodiment, when the information processing device 100 uses multiple learning models that differ for each fraudulent act, the information processing device 100 may perform re-learning for each of the multiple learning models. For example, the information processing device 100 may acquire selection targets for each of the multiple learning models and perform re-learning for each of the learning models using evaluation results for each of the acquired selection targets. For example, the information processing device 100 may perform re-learning for each of the learning models using evaluation results for auction items selected in common by the multiple learning models.
[0032] (Information processing variation 4: Used in conjunction with the current review system) In the above embodiment, the information processing device 100 may pass the targets extracted by Human-in-the-Loop (HITL) reinforcement learning through a machine learning model that is already in use to review the degree of fraud before providing them to the evaluator. This allows the information processing device 100 to evaluate only targets that have passed the current review system.
[0033] (Information Processing Variation 5: Creating Sample Code) In the above embodiment, for example, sample code for executing a program may be created using a generation AI. For example, sample code may be created that assumes that a detection target group has a predetermined number of integer-valued features and causes a learning model to select a predetermined number from the detection target group. By evaluating this predetermined number of selections, it is possible to confirm how the reward changes. Then, by considering the total value of the features, the values of specific features, etc., it becomes possible to use a learning model that takes into account the total value of the features, the values of specific features, etc. Furthermore, for example, evaluation may be performed after a person with extensive knowledge of reinforcement learning has confirmed whether the machine learning is operating correctly. Furthermore, for example, sample code may be created based on simple assumptions, or sample code may be created by assuming in detail the actual operating environment.
[0034] [3. Configuration of terminal device] Next, the configuration of the terminal device 10 according to the embodiment will be described with reference to Fig. 4. Fig. 4 is a diagram showing an example of the configuration of the terminal device 10 according to the embodiment. As shown in Fig. 4, the terminal device 10 has a communication unit 11, an input unit 12, an output unit 13, and a control unit 14.
[0035] (Communications Department 11) The communication unit 11 is realized by, for example, a network interface card (NIC), etc. The communication unit 11 is connected to a predetermined network N by wire or wirelessly, and transmits and receives information to and from the information processing device 100, etc., via the predetermined network N.
[0036] (Input section 12) The input unit 12 accepts various operations from an evaluator. In FIG. 2, the input unit 12 accepts various operations from an evaluator U1. For example, the input unit 12 may accept various operations from an evaluator via a display screen using a touch panel function. The input unit 12 may also accept various operations from buttons provided on the terminal device 10 or a keyboard or mouse connected to the terminal device 10.
[0037] (Output section 13) The output unit 13 is a display screen of a tablet terminal or the like realized by, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display, and is a display device for displaying various information. For example, the output unit 13 displays information transmitted from the information processing device 100. For example, the output unit 13 displays selection information such as a list of selection targets based on the learning model M1 and product information of the selection targets.
[0038] (Control unit 14) The control unit 14 is, for example, a controller, and is realized by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or the like executing various programs stored in a storage device inside the terminal device 10 using a RAM (Random Access Memory) as a work area. For example, these various programs include application programs installed in the terminal device 10. For example, these various programs include application programs that display information transmitted from the information processing device 100. The control unit 14 is also realized by an integrated circuit, such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0039] As shown in FIG. 4, the control unit 14 has a receiving unit 141 and a transmitting unit 142, and realizes or executes the information processing operations described below.
[0040] (Receiving unit 141) The receiving unit 141 receives, for example, information transmitted from the information processing device 100. For example, the receiving unit 141 receives selection information such as a list of selection targets based on the learning model M1 and product information of the selection targets.
[0041] (Transmitter 142) The transmission unit 142 transmits, for example, the evaluation result received from the evaluator. For example, the transmission unit 142 transmits the evaluation result as to whether or not the selected item by the learning model M1 corresponds to a fraudulent listing made by the evaluator.
[0042] [4. Configuration of the listing information server] Next, a configuration of the listing information providing server 200 according to the embodiment will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the configuration of the listing information providing server 200 according to the embodiment. As shown in Fig. 5, the listing information providing server 200 includes a communication unit 210, a storage unit 220, and a control unit 230. The listing information providing server 200 may include an input unit (e.g., a keyboard or a mouse) that receives various operations from an administrator of the listing information providing server 200, and a display unit (e.g., a liquid crystal display) that displays various information.
[0043] (Communication unit 210) The communication unit 210 is realized by, for example, a NIC etc. The communication unit 210 is connected to a network N by wire or wirelessly, and transmits and receives information to and from the terminal device 10 etc. via the network N.
[0044] (Storage unit 220) The storage unit 220 is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. As shown in FIG.
[0045] The listing information storage unit 221 stores listing information of listing targets that are listed on a predetermined service. FIG. 6 shows an example of the listing information storage unit 221 according to the embodiment. The information stored in the listing information storage unit 221 is used, for example, to select a selection target by the learning model M1. As shown in FIG. 6, the listing information storage unit 221 has items such as "listing target ID," "seller ID," and "listing information."
[0046] "Item ID" indicates identification information for identifying the item to be sold. "Seller ID" indicates identification information for identifying the seller. "Item information" indicates the selling information. In the example shown in FIG. 6, conceptual information such as "item information #1" and "item information #2" is stored in "item information," but in reality, product information and image data of the item to be sold may be stored. For example, a uniform resource locator (URL) where the image data is located, or a file path name indicating the storage location may be stored.
[0047] (control unit 230) The control unit 230 is a controller, and is realized by, for example, a CPU or an MPU executing various programs stored in a storage device inside the commodity information providing server 200 using a RAM as a work area. The control unit 230 is also realized by, for example, an integrated circuit such as an ASIC or an FPGA.
[0048] 5, control unit 230 has acquisition unit 231 and provision unit 232, and realizes or executes the information processing action described below. Note that the internal configuration of control unit 230 is not limited to the configuration shown in FIG. 5, and may be any other configuration as long as it performs the information processing described below.
[0049] (Acquisition part 231) The acquisition unit 231 acquires various types of information from the storage unit 220. Furthermore, the acquisition unit 231 stores the acquired various types of information in the storage unit 220.
[0050] The acquisition unit 231 acquires various pieces of information from an external information processing device. The acquisition unit 231 acquires various pieces of information from other information processing devices such as the information processing device 100.
[0051] The acquisition unit 231 acquires, for example, a request to send auction information, such as a request to send a list of items being put up for sale on a predetermined service.
[0052] (Providing Department 232) The providing unit 232 provides auction information in accordance with the transmission request acquired by the acquiring unit 231. For example, the providing unit 232 provides a list of items auctioned on a predetermined service in accordance with the transmission request acquired by the acquiring unit 231.
[0053] 5. Configuration of Information Processing Device Next, the configuration of the information processing device 100 according to the embodiment will be described with reference to Fig. 7. Fig. 7 is a diagram showing an example of the configuration of the information processing device 100 according to the embodiment. As shown in Fig. 7, the information processing device 100 has a communication unit 110, a storage unit 120, and a control unit 130. Note that the information processing device 100 may also have an input unit (e.g., a keyboard or a mouse) that accepts various operations from an administrator of the information processing device 100, and a display unit (e.g., a liquid crystal display) that displays various information.
[0054] (Communication unit 110) The communication unit 110 is realized by, for example, a NIC etc. The communication unit 110 is connected to a network N by wire or wirelessly, and transmits and receives information to and from the terminal device 10 etc. via the network N.
[0055] (Storage unit 120) The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. As shown in FIG. 7 , the storage unit 120 includes a learning model storage unit 121 and an evaluator information storage unit 122.
[0056] The learning model storage unit 121 stores information related to the learning model. Here, FIG. 8 shows an example of the learning model storage unit 121 according to the embodiment. The information stored in the learning model storage unit 121 is used, for example, to generate a learning model. As shown in FIG. 8, the learning model storage unit 121 has items such as "learning model ID" and "learning model."
[0057] "Learning model ID" indicates identification information for identifying a learning model. "Learning model" indicates the training data of the learning model. In the example shown in Figure 8, conceptual information such as "Learning model #1" and "Learning model #2" is stored in "Learning model," but in reality, the feature values of the items being listed that have been evaluated as fraudulent are stored.
[0058] The evaluator information storage unit 122 stores evaluator information. FIG. 9 shows an example of the evaluator information storage unit 122 according to the embodiment. The information stored in the evaluator information storage unit 122 is used, for example, to select an evaluator for evaluating the selection target. As shown in FIG. 9, the evaluator information storage unit 122 has items such as "evaluator ID," "expertise," and "evaluation history."
[0059] "Evaluator ID" indicates identification information for identifying the evaluator. "Expertise" indicates the evaluator's expertise. "Evaluation history" indicates the evaluator's evaluation history (e.g., "fraudulent" or "not fraudulent"). For example, information combining the evaluator's evaluation and the item being sold may be stored.
[0060] (control unit 130) The control unit 130 is a controller, and is realized by, for example, a CPU or an MPU executing various programs stored in a storage device inside the information processing device 100 using RAM as a work area. The control unit 130 is also realized by, for example, an integrated circuit such as an ASIC or an FPGA.
[0061] 7, control unit 130 has an acquisition unit 131, a selection unit 132, a learning unit 133, and a determination unit 134, and realizes or executes the information processing action described below. Note that the internal configuration of control unit 130 is not limited to the configuration shown in FIG. 7, and other configurations may be used as long as they perform the information processing described below.
[0062] (Acquisition part 131) The acquiring unit 131 acquires various pieces of information from the storage unit 120. The acquiring unit 131 also stores the acquired various pieces of information in the storage unit 120.
[0063] The acquisition unit 131 acquires various pieces of information from an external information processing device. The acquisition unit 131 acquires various pieces of information from other information processing devices such as the terminal device 10.
[0064] The acquisition unit 131 acquires, for example, listing information. For example, the acquisition unit 131 acquires listing information of listing targets that are listed on a predetermined service. The acquisition unit 131 also acquires, for example, selection information. For example, the acquisition unit 131 acquires selection information of selection targets selected by a learning model from among the listing targets that are listed on a predetermined service. The acquisition unit 131 also acquires, for example, evaluation results. For example, the acquisition unit 131 acquires evaluation results in which an evaluator evaluates each selection target selected by the learning model. For example, the acquisition unit 131 acquires the number of fraudulent targets that are evaluated as fraudulent among the selection targets selected by the learning model.
[0065] (Selection unit 132) The selection unit 132 selects, for example, selection information. For example, the selection unit 132 uses a learning model to select a listing target suspected of fraud. For example, the selection unit 132 selects a listing target suspected of fraud from among listing targets listed on a predetermined service. For example, the selection unit 132 inputs listing target information into the learning model, causing the learning model to select a listing target suspected of fraud. Then, the selection unit 132 provides, for example, the selection information to an evaluator. For example, the selection unit 132 provides the selection information according to the evaluator's expertise. For example, the selection unit 132 identifies an evaluator with high expertise in specific fraudulent activities related to the selected target and provides the selection information.
[0066] (Learning Section 133) The learning unit 133, for example, retrains the learning model. For example, the learning unit 133 retrains the learning model that selects selection targets to be evaluated for fraud detection requested of the evaluator. For example, the learning unit 133 retrains the learning model using the evaluation results acquired by the acquisition unit 131 as training data. For example, the learning unit 133 retrains the learning model using the number of fraud targets evaluated by the evaluator as fraudulent as the evaluation result. For example, the learning unit 133 retrains the learning model using feature quantities common to fraud targets evaluated by the evaluator as fraudulent as the evaluation result.
[0067] The learning unit 133 retrains the learning model so that the number of fraudulent targets included in the selection targets becomes larger than when the learning model before retraining is used, for example. For example, when the number of fraudulent targets included in the selection targets is "50", the learning unit 133 retrains the learning model so that the selection targets are selected so that the number of fraudulent targets included in the newly selected selection targets exceeds "50".
[0068] For example, the learning unit 133 retrains the learning model corresponding to each fraudulent act. For example, if the fraudulent act is related to a violation of the Pharmaceutical Affairs Act, the learning unit 133 retrains the learning model that selects a listing item suspected of violating the Pharmaceutical Affairs Act.
[0069] The learning unit 133 retrains the learning model using, for example, evaluation results from specialized evaluators predetermined for each type of fraudulent activity. For example, if the fraudulent activity relates to a violation of the Pharmaceutical Affairs Act, the learning unit 133 retrains the learning model using evaluation results from evaluators who specialize in detecting fraudulent activity that violates the Pharmaceutical Affairs Act. For example, the learning unit 133 retrains the learning model that selects listings suspected of violating the Pharmaceutical Affairs Act using evaluation results from evaluators who specialize in detecting fraudulent activity that violates the Pharmaceutical Affairs Act.
[0070] The learning unit 133, for example, uses a plurality of learning models to perform re-learning for each learning model. For example, the learning unit 133 uses a plurality of learning models that differ for each type of fraudulent activity to perform re-learning for each learning model. For example, the learning unit 133 performs re-learning for each learning model using evaluation results for auction items that are commonly selected by the plurality of learning models.
[0071] The learning unit 133, for example, resets the learning model and retrains it. For example, when the accuracy of fraud detection decreases, the learning unit 133 resets the learning model and retrains it to track new fraud. For example, when the learning unit 133 compares the number of fraud targets included in newly selected selection targets with the number of fraud targets when the learning model before retraining is used, and if it is determined that the accuracy of fraud detection has not improved, the learning unit 133 resets the learning model and retrains it to track new fraud. In addition, there may be cases where the accuracy of fraud detection decreases due to learning using incorrect information. For this reason, for example, the learning unit 133 may reset the learning model and retrain it to redo learning when learning using incorrect information. For example, when the accuracy of fraud detection decreases, the learning unit 133 may reset the learning model and retrain it to redo learning using incorrect information. For example, the learning unit 133 may compare the number of fraudulent targets included in newly selected selection targets with the number of fraudulent targets when the learning model before re-learning is used, and if it is determined that the accuracy of fraud detection has not improved, the learning unit 133 may reset the learning model once and re-learn it in order to start learning again with incorrect information. Also, for example, the learning unit 133 may reset the learning model once and re-learn it in order to start learning again when, due to learning with incorrect information, the learning unit 133 has fallen into a local solution that extracts targets that deviate from the actual screening situation or when the progress of learning has stagnated.
[0072] (Judgment unit 134) The determination unit 134 determines, for example, whether the accuracy of fraud detection has improved as a result of the re-learning. For example, the determination unit 134 compares the number of fraud targets included in the newly selected selection targets with the number of fraud targets when the learning model before the re-learning is used to determine whether the accuracy of fraud detection has improved. For example, the determination unit 134 determines whether the increase in the number of fraud targets exceeds a predetermined increase based on the comparison result, and if the increase in the number of fraud targets exceeds the predetermined increase, determines that the accuracy of fraud detection has improved, and if the increase in the number of fraud targets does not exceed the predetermined increase, determines that the accuracy of fraud detection has not improved.
[0073] [6. Information Processing Flow] Next, the procedure of information processing by the information processing system 1 according to the embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing the procedure of information processing by the information processing system according to the embodiment.
[0074] As shown in FIG. 10, the information processing device 100 acquires items for sale on a predetermined service (step S201). The information processing device 100 selects items for sale that are suspected of being fraudulent (step S202). The information processing device 100 provides information about the selected items to the corresponding evaluators according to the selected items (step S203). The information processing device 100 acquires the number of fraudulent items among the selected items that have been evaluated by the evaluators as being fraudulent (step S204). The information processing device 100 trains a learning model using the number of fraudulent items as training data (step S205).
[0075] [7. Effects] As described above, the information processing device 100 according to the embodiment includes the acquisition unit 131 and the learning unit 133. The acquisition unit 131 acquires the evaluator's evaluation results regarding fraudulent acts against the auction items selected by the learning model. The learning unit 133 retrains the learning model using the evaluation results acquired by the acquisition unit 131 as training data.
[0076] As a result, the information processing device 100 according to the embodiment can provide a highly accurate learning model for fraud detection that combines, for example, the machine learning concept of the learning model and the machine learning concept of human-in-the-loop.
[0077] The acquisition unit 131 also acquires, as an evaluation result, the number of fraudulent items that the evaluator has evaluated as fraudulent among the items for sale.
[0078] As a result, the information processing device 100 according to the embodiment can perform reinforcement learning using the number of cheating targets as a reward, for example.
[0079] Furthermore, the acquisition unit 131 acquires, as an evaluation result, the number of fraud targets that have been evaluated as fraud targets by the evaluator as to whether or not they are fraud targets.
[0080] As a result, the information processing device 100 according to the embodiment can perform reinforcement learning in which the number of cheating targets is used as a reward based on an evaluation of whether or not a cheating target is present.
[0081] Furthermore, the learning unit 133 retrains the learning model so that the number of fraudulent objects included in the auction items becomes larger than when the learning model before retraining is used.
[0082] As a result, the information processing device 100 according to the embodiment can provide a highly accurate learning model for fraud detection, for example, by using human-in-the-loop machine learning to reinforce the detection results of the learning model. Also, for example, the evaluator can determine that the majority of the selected items are likely to be counterfeit products, which can improve the work efficiency of the evaluator.
[0083] Furthermore, the learning unit 133 causes the learning model to learn feature amounts common to fraudulent items included in the items for auction sale.
[0084] As a result, the information processing device 100 according to the embodiment can, for example, enable reinforcement learning in which a feature amount common to fraud targets is set as a notable feature amount. Also, for example, the information processing device 100 according to the embodiment can enable reinforcement learning for each pattern of fraudulent behavior based on a notable feature amount.
[0085] Furthermore, the learning unit 133 retrains the learning model for each misconduct using the evaluation results from specialized evaluators predetermined for each misconduct.
[0086] This enables the information processing device 100 according to the embodiment to perform reinforcement learning based on the evaluation results of experts, for example, and therefore can provide a highly accurate learning model for fraud detection.
[0087] Furthermore, the learning unit 133 retrains each learning model using the evaluation results for the auction items selected in common by a plurality of learning models that differ for each type of fraudulent activity.
[0088] As a result, the information processing device 100 according to the embodiment can, for example, filter the selection targets to be evaluated based on multiple learning models, thereby providing a highly accurate learning model for fraud detection.
[0089] Furthermore, if the learning unit 133 determines that the accuracy of fraud detection has not improved by comparing the number of fraudulent objects contained in the items being auctioned with the number of fraudulent objects when the learning model before re-learning is used, it resets the learning model to track new fraud.
[0090] As a result, the information processing device 100 according to the embodiment can perform reinforcement learning to track new fraud even if the trend of fraud changes, for example, and can therefore provide a highly accurate learning model for fraud detection.
[0091] [8. Hardware Configuration] The information processing device 100 according to the embodiment described above is realized, for example, by a computer 1000 configured as shown in Fig. 11. Fig. 11 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device 100. The computer 1000 has a CPU 1100, a RAM 1200, a ROM 1300, an HDD 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.
[0092] The CPU 1100 operates and controls each unit based on programs stored in the ROM 1300 or the HDD 1400. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 starts up, programs that depend on the hardware of the computer 1000, and the like.
[0093] The HDD 1400 stores programs executed by the CPU 1100, data used by these programs, etc. The communication interface 1500 acquires data from other devices via a predetermined communication network and sends it to the CPU 1100, and transmits data generated by the CPU 1100 to other devices via the predetermined communication network.
[0094] The CPU 1100 controls output devices such as a display and a printer, and input devices such as a keyboard and a mouse, via the input / output interface 1600. The CPU 1100 acquires data from the input devices via the input / output interface 1600. The CPU 1100 also outputs generated data to the output devices via the input / output interface 1600.
[0095] Media interface 1700 reads a program or data stored in recording medium 1800 and provides it to CPU 1100 via RAM 1200. CPU 1100 loads the program or data from recording medium 1800 onto RAM 1200 via media interface 1700 and executes the loaded program. Recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.
[0096] For example, when the computer 1000 functions as the information processing device 100 according to the embodiment, the CPU 1100 of the computer 1000 executes programs loaded onto the RAM 1200 to realize the functions of the control unit 130. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800, but as another example, the CPU 1100 may obtain these programs from another device via a predetermined communication network.
[0097] [9. Other] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.
[0098] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0099] Furthermore, the above-described embodiments can be combined as appropriate within the scope of not causing any contradiction in the processing content.
[0100] Although some of the embodiments of the present application have been described in detail above with reference to the drawings, these are merely examples, and the present invention can be implemented in other forms that include the embodiments described in the Disclosure of the Invention section and that have undergone various modifications and improvements based on the knowledge of those skilled in the art.
[0101] Furthermore, the above-mentioned "section, module, unit" can be read as "means" or "circuit," etc. For example, an acquisition unit can be read as an acquisition means or an acquisition circuit. [Explanation of symbols]
[0102] 1. Information Processing Systems 10 Terminal Equipment 11 Communications Department 12 Input section 13 Output section 14 Control Unit 100 Information processing device 110 Communications Department 120 Storage section 121 Learning model memory unit 122 Evaluator information storage unit 130 control section 131 Acquisition Department 132 Selection section 133 Learning Department 134 Judgment section 141 Receiving unit 142 Transmitter 200 Listing information server 210 Communications Department 220 Storage section 221 Exhibition information storage section 230 Control Unit 231 Acquisition Department 232 Provision Department N Network
Claims
1. an acquisition unit that acquires an evaluation result of an evaluator regarding fraud for the auction item selected by the learning model; a learning unit that re-learns the learning model using the evaluation results acquired by the acquisition unit as training data; An information processing device comprising:
2. The acquisition unit The number of fraudulent items that the evaluator has evaluated as fraudulent among the items for sale is obtained as the evaluation result.
2. The information processing apparatus according to claim 1, wherein:
3. The acquisition unit The number of fraud targets evaluated as fraud targets by the evaluator is obtained as the evaluation result.
2. The information processing apparatus according to claim 1, wherein:
4. The learning unit The learning model is retrained so that the number of fraudulent objects included in the auction items is greater than that in the case where the learning model before the retraining is used.
2. The information processing apparatus according to claim 1, wherein:
5. The learning unit The learning model is trained on a feature amount common to fraudulent objects included in the auction items.
2. The information processing apparatus according to claim 1, wherein:
6. The learning unit The learning model for each misconduct is retrained using the evaluation results by a specialized evaluator predetermined for each misconduct.
2. The information processing apparatus according to claim 1, wherein:
7. The learning unit The learning model is retrained for each of the learning models using the evaluation results for the auction items commonly selected by the plurality of learning models that differ for each fraudulent act.
2. The information processing apparatus according to claim 1, wherein:
8. The learning unit If it is determined that the accuracy of fraud detection has not improved by comparing the number of fraudulent objects included in the auction items with the number of fraudulent objects when the learning model before the re-learning is used, the learning model is reset to track new fraud.
2. The information processing apparatus according to claim 1, wherein:
9. 1. A computer-implemented information processing method, comprising: an acquisition step of acquiring an evaluation result of an evaluator regarding fraud for the auction item selected by the learning model; a learning step of re-learning the learning model using the evaluation results acquired in the acquisition step as training data; An information processing method comprising:
10. An acquisition step of acquiring an evaluation result of an evaluator regarding fraud for the listing object selected by the learning model; a learning step of re-learning the learning model using the evaluation results acquired by the acquisition step as training data; An information processing program characterized by causing a computer to execute the above.