system

JP2026085722APending Publication Date: 2026-05-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-11-13
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Existing automated learning systems face challenges such as bias in learning data, increased costs due to outsourcing, lack of consistency in feedback, and security concerns including unauthorized access, which hinder the effectiveness and accuracy of image recognition.

Method used

A system that collects feedback using user image recognition behavior, aggregates and analyzes this data to improve the training dataset, and implements security measures to distinguish between humans and automated programs, ensuring data reliability and preventing unauthorized access.

Benefits of technology

This approach enhances the accuracy of automated generation technology efficiently and cost-effectively by improving the learning model's accuracy and security, while reducing misidentification and unauthorized access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026085722000001_ABST
    Figure 2026085722000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of presenting image authentication in order to collect image selection feedback from users, A means of aggregating and analyzing the collected feedback data, A means of improving the training dataset of an automated learning system using the analyzed data, Means for implementing security measures to distinguish between humans and automated programs, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] It is required to reduce the bias included in the learning data of the automatic generation technology and realize a more accurate learning model. However, in the conventional feedback process, due to problems such as increased cost due to outsourcing, bias in feedback, and lack of consistency, the effect cannot be fully exerted. Also, from the perspective of security, there is a situation where unauthorized access is a concern. Therefore, it is required to solve these problems efficiently and effectively.

Means for Solving the Problems

[0005] This invention provides a system that collects feedback using a user's image recognition behavior. This system presents image recognition to the user, collects feedback from the user, and uses that data to improve the training dataset of an automated learning system. The ability to aggregate and analyze feedback data improves the accuracy of image recognition. Furthermore, security measures are in place to distinguish between humans and automated programs, ensuring data reliability while protecting against unauthorized access. This makes it possible to improve the accuracy of automated generation technology efficiently and at low cost.

[0006] "Image authentication" is an authentication method that demonstrates human judgment by having a user select an image that matches specific criteria from a set of images presented.

[0007] "Feedback data" refers to information based on images selected by the user through image recognition, and is used to improve the automated learning system.

[0008] An "automatic learning system" is a program or device that learns independently based on collected data and improves its decision-making ability.

[0009] A "training dataset" is a collection of information used by an automated learning system to acquire knowledge through the learning process.

[0010] "Security measures" are means of identifying and preventing automated programs and unauthorized access, and protecting the safety of systems and data. [Brief explanation of the drawing]

[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0013] First, let's explain the terminology used in the following explanation.

[0014] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0015] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0016] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0017] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0019] [First Embodiment]

[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0032] The present invention provides a feedback collection system utilizing image authentication, which includes a process for providing feedback through image authentication when a user accesses a specific information system. This system is realized through the cooperation of three parties: a server, a terminal, and a user.

[0033] The server first generates a randomly selected set of image verification images for the website the user has accessed. This set of images includes images that the user must select according to specific criteria. For example, the user might be instructed to "select all images of cats." The server then sends these instructions and image information to the user's device.

[0034] The device displays the received image authentication to the user. The user selects the correct image from the displayed images according to the instructions. For example, if instructed to select images of cats, the user clicks on all the images of cats. Once the user has completed their selection, the device sends the selection data back to the server.

[0035] The server analyzes the selection data collected from users and compiles data on which images were selected and how many times. Based on this compiled data, the server updates and improves the training dataset for the automated learning system. For example, if other animals that are easily mistaken for cats are repeatedly selected, this information is recorded as training data so that the AI ​​can correct the misidentification.

[0036] Furthermore, the server uses security algorithms to check for any unauthorized access during this process. If suspicious patterns are detected, it can issue a warning and, if necessary, request additional authentication.

[0037] One concrete example of this system is its use as user authentication on a shopping platform. When a user logs in, they are presented with a CAPTCHA-style image verification, through which they provide feedback to the system. This feedback is used to improve the accuracy of visual product recognition. This approach makes it possible to efficiently collect data using user behavior and correct AI bias.

[0038] The following describes the processing flow.

[0039] Step 1:

[0040] The server generates a set of images for CAPTCHA for the digital platform accessed by the user. The image set is selected based on specific criteria and is prepared along with instructions to prompt the user to make a selection.

[0041] Step 2:

[0042] The server sends the generated image set and instructions to the terminal. The images are sent in an appropriate format to ensure clear display and rapid authentication.

[0043] Step 3:

[0044] The terminal presents the received image set and instructions to the user. These are visually clear and easy to understand on the web page the user accesses.

[0045] Step 4:

[0046] The user views the images displayed on the device screen and selects the correct image according to the instructions. For example, if the instruction is "Select all the cat images," the user will click on the cat images.

[0047] Step 5:

[0048] The device sends information about the image selections made by the user to the server. This information includes the IDs and selection order of the selected images.

[0049] Step 6:

[0050] The server analyzes the received selection information and compiles data on which images were selected correctly and which were selected incorrectly. Based on these results, feedback data is generated.

[0051] Step 7:

[0052] The server uses the collected feedback data to update the training dataset of the automated learning system and improve the model's bias. This improves the AI's recognition accuracy.

[0053] Step 8:

[0054] The server analyzes user behavior through the CAPTCHA process and executes security algorithms. This verifies that there are no unnatural selection patterns or unauthorized access attempts, thus ensuring security.

[0055] (Example 1)

[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0057] Conventional image recognition systems have faced challenges such as the potential for misidentification and reduced system learning efficiency due to insufficient utilization of user feedback information. Furthermore, detecting unauthorized access was difficult, potentially leading to security vulnerabilities.

[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0059] In this invention, the server includes means for generating a randomly selected set of image authentications to present to the user; means for transmitting the generated image authentications and instructions to a terminal; means for collecting the user's image selection actions via the terminal and transmitting them to the server; means for analyzing and aggregating the collected feedback information; means for updating and optimizing the data set of the automatic learning device based on the analysis and aggregation results; and means for performing security management to detect abnormalities in behavioral patterns and identify unauthorized access. This makes it possible to effectively utilize user feedback to reduce misidentification, improve the learning efficiency of the system, and enable the detection of unauthorized access and strengthen security.

[0060] A "user" refers to a person who operates a device in order to solve a problem provided through image authentication.

[0061] "Image verification" refers to a collection of randomly selected images in which a user is instructed to choose a specific object.

[0062] "Terminal" refers to a device or equipment that provides image authentication to a user and transmits the user's input to a server.

[0063] A "server" refers to a central processing unit that generates image authentication data, sends instructions, analyzes feedback information, and updates the automated learning device.

[0064] "Feedback information" refers to data that includes the results of the user's selections through image authentication.

[0065] An "automatic learning device" refers to an artificial intelligence system that uses feedback information to learn and improve its recognition capabilities.

[0066] "Unauthorized access" refers to any act that disrupts the normal functioning of a system or intrudes into the system without permission.

[0067] "Security management" refers to the process of detecting unauthorized access and taking appropriate countermeasures to ensure the security of a system.

[0068] This invention relates to a system for providing feedback through image authentication when a user accesses a specific information system. The following describes embodiments for carrying out this invention.

[0069] The server generates a randomly selected set of image authentication images for the information system accessed by the user. This process utilizes an image processing library to prevent misrecognition by combining random images. For example, OpenCV could be chosen as the library.

[0070] The generated image authentication data and accompanying instructions (such as "Select all the cat images") are sent from the server to the terminal. This data transmission is secured using the HTTPS protocol. The terminal's role is to present the image authentication data received from the server to the user. On the terminal, a web browser uses HTML and JavaScript (registered trademark) to construct a user interface, allowing the user to select images.

[0071] The user selects the correct image from those displayed on the device, following the instructions. This operation is often performed using common input devices such as a mouse or touchscreen. After the user completes their selection, the device sends the selection result back to the server.

[0072] The server analyzes feedback information collected from users and compiles data on which images were selected and how many times. This analysis involves an automated learning system, and the dataset is updated and optimized based on the analysis results. Images with many misidentifications are subjected to more detailed training to improve the accuracy of the automated learning system. Machine learning libraries such as TENSORFLOW® and PyTorch are used in this process.

[0073] Furthermore, the server implements security measures to identify unauthorized access. For this purpose, security algorithms are in place, and additional authentication procedures may be required if unusual patterns are detected.

[0074] A concrete example of its use is user authentication on a shopping platform. When users log in, they are provided with CAPTCHA-style image verification, and this feedback is used to improve the AI ​​model. An example of a prompt might be, "Please provide a step-by-step detailed explanation of the process of an image verification system that identifies cats."

[0075] This system leverages user feedback to improve AI accuracy and enhance security, thereby achieving stable user authentication.

[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0077] Step 1:

[0078] The server generates a randomly selected set of image authentications based on the user's request for the information system they accessed. This step uses an image processing library to randomly extract images from a database. The input is a specific authentication request from the user, and the output is a set of images and instructions (e.g., "Select all images of cats").

[0079] Step 2:

[0080] The server sends the generated image authentication and instructions to the terminal. Here, the data is transmitted securely using the HTTPS protocol. The input is the image set and instructions generated in step 1, and the output is the data converted into a format viewable by the terminal.

[0081] Step 3:

[0082] The terminal displays images and instructions received from the server to the user. Using HTML and JavaScript on a web browser, the images are laid out in a user-friendly manner. Input consists of image data and instructions sent from the server. Output is a visual presentation to the user.

[0083] Step 4:

[0084] The user selects the correct image from those displayed on the terminal, following the instructions. The user makes the selection using an input device (e.g., mouse, touchscreen). The input is the instructions and image set displayed by the terminal, and the output is the image selection result by the user.

[0085] Step 5:

[0086] The terminal collects the user's selection results and sends them to the server. During this process, data such as the ID of the selected image and the selection time are formatted and securely transmitted to the server. The input is information about the image selected by the user, and the output is the data sent to the server.

[0087] Step 6:

[0088] The server analyzes the received data and aggregates the selection frequency of each image. An AI model is used to identify misrecognition trends and optimize the training dataset. The input is the user's selection data, and the output is the improved training dataset.

[0089] Step 7:

[0090] The server verifies that no unauthorized access has occurred during the selection process. If the security algorithm detects any unusual patterns or behavior, it issues a warning and requests additional authentication if necessary. Input is the operation history of all users, and output is logs and alerts for unauthorized detection.

[0091] (Application Example 1)

[0092] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0093] Current electronic transactions face challenges such as increasing unauthorized access, limitations in AI learning accuracy, and a lack of enhanced security. For example, there is a lack of efficient means to collect user feedback, and concerns remain about improving AI recognition accuracy. Furthermore, there is a need for methods to maintain high usability while achieving a high level of security.

[0094] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0095] In this invention, the server includes means for presenting visual authentication to collect visual selection feedback from the user, means for aggregating and processing the collected feedback information, and means for improving the learning information set of the automated learning system using the processed information. This enables efficient collection of user feedback and improvement of AI learning accuracy while maintaining a high level of security in electronic transaction settings.

[0096] "Visual selection feedback" refers to information resulting from a user's selection of visual information according to specific criteria.

[0097] "Visual authentication" is an authentication method that uses the visual recognition of an object to verify user input.

[0098] "Feedback information" is data generated based on user choices and actions, and is used to improve the system.

[0099] An "automatic learning system" is a system that has a learning algorithm that automatically improves the accuracy of recognition and judgment based on collected data.

[0100] "Unauthorized access" refers to unauthorized access to a system by someone without the proper authorization, and is considered a security threat.

[0101] "Electronic trading" refers to a form of transaction in which goods and services are bought and sold via a network.

[0102] "Security" refers to the protective measures and technologies used to safeguard information and systems from unauthorized access and data breaches.

[0103] "Usability" is a concept that refers to the ease of use and satisfaction a user experiences when using a system.

[0104] In an embodiment of this invention, a server plays a primary role. The server generates a randomly selected set of visual authentications to efficiently collect visual selection feedback from the user. This authentication set includes instructions and images that comply with specific conditions of the transaction, and is transmitted to the user's terminal.

[0105] The user's device has an interface that presents the received visual authentication to the user. Through this interface, the user selects the correct visual information according to specified conditions. Once the selection is complete, the device sends the selection result back to the server.

[0106] The server aggregates the submitted selection results and improves the automated learning system's training information set while performing processing based on the generated AI model. Specifically, it updates the AI's visual recognition algorithm based on the collected feedback information to improve recognition accuracy. In this process, in addition to the usual recognition tasks, unauthorized access detection is also performed simultaneously from a security perspective. If unauthorized access is detected, the terminal will be required to grant additional authorization.

[0107] Furthermore, this system provides a high level of security in the electronic transaction process, creating an environment where users can trade with peace of mind. For example, when a user purchases goods through online shopping, a visual authentication system is implemented before payment, and the user is instructed to "select all the fruit images."

[0108] An example of a prompt might be, "Please tell me how to use this image recognition system to prevent unauthorized access while improving the AI's visual recognition capabilities." This makes it easier to develop solutions for specific problems.

[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0110] Step 1:

[0111] When the server receives a request necessary for user authentication, it initiates the visual authentication process. During this process, the server generates a randomly selected set of visual authentication credentials. It receives user ID and access status as input, creates the visual authentication credentials based on this information, and prepares to present them to the user.

[0112] Step 2:

[0113] The server sends the generated visual authentication set to the user's device. This includes image data and recognition conditions (e.g., "Select all images of apples"). The output is then ready for the device to present authentication information to the user.

[0114] Step 3:

[0115] The terminal displays the received visual authentication set to the user. The user makes the appropriate selection from the presented images according to the given conditions. The terminal receives image data and instructions sent from the server as input and prepares the image information selected by the user as output.

[0116] Step 4:

[0117] The user selects images as instructed on the device. For example, if the instruction is "Select all images of apples," the user clicks on the apple images to complete the selection. The selected images are then saved on the device.

[0118] Step 5:

[0119] The terminal sends the user's selection results back to the server. It receives the user's selection data as input and prepares to send it to the server as output. No data processing is performed during this process, but reliable data transmission is required.

[0120] Step 6:

[0121] The server aggregates and analyzes the selection data sent from the terminals. Here, a generative AI model is used to process the data and improve the training information set of the automated learning system to enhance visual recognition accuracy. The server takes the selection data as input and maintains the improved training information set as output.

[0122] Step 7:

[0123] The server checks for unauthorized access based on the aggregated data. It receives selected data as input, checks for any abnormal selection patterns, and requests additional authentication if necessary. As output, it decides whether to send an instruction to the user for additional authentication.

[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0125] The system of the present invention is implemented by combining image recognition, feedback collection, improvement of training data for an automated learning system, and an emotion engine. When a user accesses the system, the server first generates a set of image recognition images. These image recognition images include instructions to prompt selections that meet specific conditions.

[0126] The device displays a set of image authentication data received from the server, allowing the user to select an image according to the instructions. As the user selects an image, the device uses its built-in emotion engine to analyze the user's facial expressions. The emotion engine estimates the user's emotions based on their facial features and collects this data.

[0127] After the selection is complete, the device sends the image selection data and sentiment data to the server. The server aggregates this data and incorporates it into the training dataset of the automated learning system. Based on the feedback data and the user's sentiment data, the automated learning system's decisions become more accurate and flexible.

[0128] As a concrete example, consider a case where the system of the present invention is used in an online learning platform. When a learner answers a specific quiz, CAPTCHA-style image authentication is performed, and user sentiment data is collected during this process. This sentiment data is used to appropriately adjust the difficulty level of the quiz. If the learner finds it difficult, the overall flexibility of the system is increased so that a simpler set of images is provided for the next authentication.

[0129] The introduction of this system will not only increase the amount of information obtained from user feedback, but will also make the image recognition process itself more interactive and personalized. By utilizing emotion recognition, the quality of the user experience will improve, contributing to improved AI performance.

[0130] The following describes the processing flow.

[0131] Step 1:

[0132] When a user accesses the server, it generates a corresponding set of image recognition data. This set prompts the user to make a selection based on specific recognition criteria and includes any necessary instructions.

[0133] Step 2:

[0134] The server sends the generated image set and authentication instructions to the terminal. The images are formatted to be easily visually verifiable by the user.

[0135] Step 3:

[0136] The terminal displays the image set and instructions received from the server on the user screen. This prepares the user to make a selection.

[0137] Step 4:

[0138] The user looks at the images displayed on the device and selects the correct image according to the instructions. For example, if the instructions say, "Select all the images of dogs," the user clicks on the images of dogs.

[0139] Step 5:

[0140] While the user is selecting an image, the device activates its built-in emotion engine to analyze the user's facial expressions and estimate their emotional state.

[0141] Step 6:

[0142] Once the user has made their selection, the device sends the selected image data and analyzed sentiment data to the server.

[0143] Step 7:

[0144] The server receives the transmitted image selection information and sentiment data. Based on this, it aggregates feedback data and incorporates it into the training dataset of the automated learning system.

[0145] Step 8:

[0146] The server uses this feedback data to update the automated learning system, correcting misrecognitions and improving model biases.

[0147] Step 9:

[0148] The server uses the collected sentiment data to assess the user's stress and anxiety levels and adjust the difficulty of the next image recognition test as needed. This ensures a comfortable user experience.

[0149] (Example 2)

[0150] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0151] While conventional image recognition systems function as security measures, they suffer from a uniform user experience and a lack of flexibility in responding to the individual needs and emotions of users. Furthermore, the ineffective use of feedback data limits the improvement of the learning system's accuracy. This presents a challenge, particularly in learning support settings, where it is difficult to provide users with the most suitable content.

[0152] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0153] In this invention, the server includes means for generating a visual dataset to be presented for collecting image selection feedback from the user, means for analyzing emotional data from the user's facial expressions using a terminal device, and means for aggregating the image selection data and emotional data and analyzing them to improve the learning dataset of the automated learning device. This makes it possible to flexibly adjust the service based on the user's emotions and improve the accuracy of the learning system.

[0154] A "visual dataset" is a set of images generated according to specific conditions in order to collect image selection feedback from users.

[0155] "Emotional data" refers to data that indicates the type and intensity of emotions estimated from the user's facial expression information analyzed by the terminal device.

[0156] An "automatic learning device" is a system that can autonomously learn and improve its accuracy based on collected data.

[0157] "Feedback data" refers to data provided based on user intentions and choices, and is used to improve the system.

[0158] "Flexible adjustment" means adaptively changing the settings and outputs of a service or system based on user feedback and sentiment data.

[0159] This invention improves the adaptability and accuracy of a system based on image selection feedback and emotion data acquired through user interaction. The following hardware and software are used to implement the invention.

[0160] The server generates a visual dataset. This involves the process of creating a series of images that are presented as image authentication when a user accesses the system. There are no specific restrictions on the server platform used; any platform capable of efficiently generating and distributing image data is acceptable. In this process, the server uses image generation software to generate a set of images based on a specific algorithm.

[0161] The terminal displays images received from the server to the user, supporting the user's selection process. The terminal has an emotion engine installed that analyzes the user's facial expressions in real time, and this engine uses software that utilizes image recognition technology. Specifically, it uses a camera and CPU resources.

[0162] When a user engages in image recognition and makes a selection, the facial expression data is analyzed by an emotion engine. Emotions such as joy or confusion are estimated from the user's facial expressions and collected as emotion data. This enables data collection based on the user's reactions.

[0163] One concrete example is its use in online learning platforms. When learners take a quiz, their emotions at that moment can be analyzed through image recognition, and the difficulty of the quiz can be adjusted based on the estimated emotions. As a result, it becomes possible to provide a flexible learning experience tailored to the user.

[0164] An example of a prompt message is, "How can I use image recognition and emotion recognition together to adjust the difficulty of quizzes based on a user's learning progress on an online learning platform?" This demonstrates one method of improving the user experience through the use of generative AI models.

[0165] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0166] Step 1:

[0167] The server generates a visual dataset when a user accesses the system. Inputs include user profile information and existing data within the system. Based on this information, image generation software is used to generate a set of images containing specific selection instructions, which are then output as an image authentication set for use in authentication. Specifically, an algorithm is used to randomly or conditionally select images.

[0168] Step 2:

[0169] The terminal displays the image authentication set received from the server to the user. The input here is the image data sent from the server. This image is displayed on the screen, and screen interaction elements are prepared so that the user can select an image. The output is the configuration of the interaction interface related to the displayed image. Specifically, it generates a UI that displays the images side by side to show the appropriate choice.

[0170] Step 3:

[0171] The user selects an image from the images displayed on the device according to the instructions. The input consists of the displayed image and its instructions. The user selects one or more images based on this and confirms their selection. At this time, the user's selection information is output.

[0172] Step 4:

[0173] The device analyzes the user's facial expressions in real time using its built-in emotion engine. The input is a facial image of the user captured by the device's camera. This is processed by analysis software to estimate emotions from the facial expressions. As a result of the analysis, estimated emotion data is output. Specifically, it generates numerical emotion parameters based on eye movements and changes in mouth shape using an AI model.

[0174] Step 5:

[0175] The terminal sends the user's image selection data and emotion data to the server. The input consists of the image selection information obtained in step 3 and the emotion data obtained in step 4. This information is configured as a data packet and transferred to the server. The transmitted information becomes the output.

[0176] Step 6:

[0177] The server aggregates the received image selection data and sentiment data, and updates the dataset of the automated learning device. It uses data sent from the terminal as input. Using this information, it updates the training data based on the AI ​​algorithm and outputs the results. A specific example is analyzing the reaction trends of each user and reflecting this in the generation of subsequent datasets.

[0178] (Application Example 2)

[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0180] Conventional factory automation systems have limited interaction with workers, making it difficult to respond flexibly to individual work situations and workers' emotions. This has resulted in increased worker burden without optimized work efficiency. To solve this, it is necessary to incorporate dynamic task adjustments based on the work environment and feedback based on the worker's state.

[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0182] In this invention, the server includes means for providing a function to present image authentication and collect feedback from the user based on candidate selection; means for aggregating and analyzing the collected feedback information; means for using the analyzed information to enhance the learning dataset of the automatic learning function; and means for collecting image information and sentiment information from information terminals in the work environment and adjusting work conditions. This enables dynamic work adjustment according to the worker's state.

[0183] "Image verification" is a method of verifying a user's identity by having them select a specific image.

[0184] "Candidate selection" refers to the process of choosing the appropriate option from the choices presented to the user.

[0185] "Feedback information" refers to information obtained based on user choices and responses.

[0186] "Aggregation" is the process of combining and organizing multiple data sets.

[0187] "Analysis" is the process of examining collected data in detail to derive results and trends.

[0188] "Automatic learning function" refers to a function that allows a machine to learn and improve independently using the data it has collected.

[0189] A "training dataset" is a collection of data used to train a model in machine learning.

[0190] An "information terminal" is an electronic device used to collect and process information.

[0191] "Image information" refers to information that includes visual data.

[0192] "Emotional information" refers to emotional data inferred from a user's facial expressions and reactions.

[0193] "Adjusting working conditions" means appropriately changing tasks and methods according to the work environment and the condition of the workers.

[0194] This invention is a system for flexibly adjusting the operation of robots in a factory work environment. The system uses image recognition and emotion recognition to collect and analyze worker feedback information. This makes it possible to adjust working conditions according to the worker's state. A specific embodiment is shown below.

[0195] The server prompts the user to select a specific image via an image recognition system. Once the user selects an image, the terminal generates feedback information based on that selection. Simultaneously, a camera on the information terminal captures the user's facial expression, and the collected image information is analyzed using emotion recognition technology. This emotion recognition is performed using machine learning libraries and emotion engines.

[0196] The collected feedback and sentiment information is aggregated on a server. The server utilizes standard database systems and data analysis tools for data aggregation and analysis. The analyzed data is saved as a training dataset enhanced by automated learning functions, which is then used to refine future work. Specifically, data processing is performed using Python machine learning libraries and streaming frameworks.

[0197] For example, if a worker is deemed to have difficulty concentrating during a particular assembly task, the next work instructions will be simplified and the work speed adjusted. This reduces the burden on the worker and improves efficiency.

[0198] Specific examples of prompts for a generated AI model include: "Design an assistant function for a robot performing assembly work in a factory that can determine if the worker is tired and adjust the work if they are. Specifically, consider a system that combines image recognition and emotion recognition."

[0199] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0200] Step 1:

[0201] The server presents the user with image authentication. At this time, it generates candidate images and sends them to the terminal along with instructions prompting the user to make a specific selection. The input consists of the user's authentication information and the image selection options, while the output consists of instructions and a set of images.

[0202] Step 2:

[0203] The terminal displays a set of images received from the server to the user. When the user selects an image, the user's selected image information is returned to the terminal as input, and this selection data is output.

[0204] Step 3:

[0205] The device captures the user's facial expression during selection using its camera and analyzes it with its built-in emotion engine. The input is the captured image information, and the device performs facial expression analysis using emotion recognition technology, generating user emotion data as output.

[0206] Step 4:

[0207] Once the user completes their image selection, the device sends the image selection data and sentiment data to the server. The input is the selection data and sentiment data, and the output is this dataset. Data transmission takes place over the network.

[0208] Step 5:

[0209] The server aggregates the received data into a database and performs aggregation and analysis. It receives image selection data and sentiment data as input data, aggregates it using the database system, and outputs the analysis results using data analysis tools. The analysis results are then used as a training dataset for future development.

[0210] Step 6:

[0211] The server improves its automatic learning function based on the analysis results and incorporates these improvements into the next work adjustment. The input is the analyzed data, and the output is the improved training dataset. Specifically, the model is readjusted using machine learning libraries.

[0212] This entire process enables dynamic adjustment of working conditions according to the user's state, providing a more efficient work environment.

[0213] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0214] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0215] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0216] [Second Embodiment]

[0217] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0218] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0219] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0220] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0221] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0222] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0223] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0224] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0225] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0226] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0227] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0228] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0229] The present invention provides a feedback collection system utilizing image authentication, which includes a process for providing feedback through image authentication when a user accesses a specific information system. This system is realized through the cooperation of three parties: a server, a terminal, and a user.

[0230] The server first generates a randomly selected set of image verification images for the website the user has accessed. This set of images includes images that the user must select according to specific criteria. For example, the user might be instructed to "select all images of cats." The server then sends these instructions and image information to the user's device.

[0231] The device displays the received image authentication to the user. The user selects the correct image from the displayed images according to the instructions. For example, if instructed to select images of cats, the user clicks on all the images of cats. Once the user has completed their selection, the device sends the selection data back to the server.

[0232] The server analyzes the selection data collected from users and compiles data on which images were selected and how many times. Based on this compiled data, the server updates and improves the training dataset for the automated learning system. For example, if other animals that are easily mistaken for cats are repeatedly selected, this information is recorded as training data so that the AI ​​can correct the misidentification.

[0233] Furthermore, the server uses security algorithms to check for any unauthorized access during this process. If suspicious patterns are detected, it can issue a warning and, if necessary, request additional authentication.

[0234] One concrete example of this system is its use as user authentication on a shopping platform. When a user logs in, they are presented with a CAPTCHA-style image verification, through which they provide feedback to the system. This feedback is used to improve the accuracy of visual product recognition. This approach makes it possible to efficiently collect data using user behavior and correct AI bias.

[0235] The following describes the processing flow.

[0236] Step 1:

[0237] The server generates a set of images for CAPTCHA for the digital platform accessed by the user. The image set is selected based on specific criteria and is prepared along with instructions to prompt the user to make a selection.

[0238] Step 2:

[0239] The server sends the generated image set and instructions to the terminal. The images are sent in an appropriate format to ensure clear display and rapid authentication.

[0240] Step 3:

[0241] The terminal presents the received image set and instructions to the user. These are visually clear and easy to understand on the web page the user accesses.

[0242] Step 4:

[0243] The user views the images displayed on the device screen and selects the correct image according to the instructions. For example, if the instruction is "Select all the cat images," the user will click on the cat images.

[0244] Step 5:

[0245] The device sends information about the image selections made by the user to the server. This information includes the IDs and selection order of the selected images.

[0246] Step 6:

[0247] The server analyzes the received selection information and compiles data on which images were selected correctly and which were selected incorrectly. Based on these results, feedback data is generated.

[0248] Step 7:

[0249] The server uses the collected feedback data to update the training dataset of the automated learning system and improve the model's bias. This improves the AI's recognition accuracy.

[0250] Step 8:

[0251] The server analyzes user behavior through the CAPTCHA process and executes security algorithms. This verifies that there are no unnatural selection patterns or unauthorized access attempts, thus ensuring security.

[0252] (Example 1)

[0253] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0254] Conventional image recognition systems have faced challenges such as the potential for misidentification and reduced system learning efficiency due to insufficient utilization of user feedback information. Furthermore, detecting unauthorized access was difficult, potentially leading to security vulnerabilities.

[0255] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0256] In this invention, the server includes means for generating a randomly selected set of image authentications to present to the user; means for transmitting the generated image authentications and instructions to a terminal; means for collecting the user's image selection actions via the terminal and transmitting them to the server; means for analyzing and aggregating the collected feedback information; means for updating and optimizing the data set of the automatic learning device based on the analysis and aggregation results; and means for performing security management to detect abnormalities in behavioral patterns and identify unauthorized access. This makes it possible to effectively utilize user feedback to reduce misidentification, improve the learning efficiency of the system, and enable the detection of unauthorized access and strengthen security.

[0257] A "user" refers to a person who operates a device in order to solve a problem provided through image authentication.

[0258] "Image verification" refers to a collection of randomly selected images in which a user is instructed to choose a specific object.

[0259] "Terminal" refers to a device or equipment that provides image authentication to a user and transmits the user's input to a server.

[0260] A "server" refers to a central processing unit that generates image authentication data, sends instructions, analyzes feedback information, and updates the automated learning device.

[0261] "Feedback information" refers to data that includes the results of the user's selections through image authentication.

[0262] An "automatic learning device" refers to an artificial intelligence system that uses feedback information to learn and improve its recognition capabilities.

[0263] "Unauthorized access" refers to any act that disrupts the normal functioning of a system or intrudes into the system without permission.

[0264] "Security management" refers to the process of detecting unauthorized access and taking appropriate countermeasures to ensure the security of a system.

[0265] This invention relates to a system for providing feedback through image authentication when a user accesses a specific information system. The following describes embodiments for carrying out this invention.

[0266] The server generates a randomly selected set of image authentication images for the information system accessed by the user. This process utilizes an image processing library to prevent misrecognition by combining random images. For example, OpenCV could be chosen as the library.

[0267] The generated image authentication data and accompanying instructions (such as "Select all the cat images") are sent from the server to the terminal. The HTTPS protocol is used to ensure the security of this data transmission. The terminal's role is to present the image authentication data received from the server to the user. On the terminal, a web browser uses HTML and JavaScript to construct a user interface, allowing the user to select images.

[0268] The user selects the correct image from those displayed on the device, following the instructions. This operation is often performed using common input devices such as a mouse or touchscreen. After the user completes their selection, the device sends the selection result back to the server.

[0269] The server analyzes feedback information collected from users and compiles data on which images were selected and how many times. This analysis involves an automated learning system, and the dataset is updated and optimized based on the analysis results. Images with many misidentifications are subjected to more detailed training to improve the accuracy of the automated learning system. Machine learning libraries such as TensorFlow and PyTorch are used in this process.

[0270] Furthermore, the server implements security measures to identify unauthorized access. For this purpose, security algorithms are in place, and additional authentication procedures may be required if unusual patterns are detected.

[0271] A concrete example of its use is user authentication on a shopping platform. When users log in, they are provided with CAPTCHA-style image verification, and this feedback is used to improve the AI ​​model. An example of a prompt might be, "Please provide a step-by-step detailed explanation of the process of an image verification system that identifies cats."

[0272] This system leverages user feedback to improve AI accuracy and enhance security, thereby achieving stable user authentication.

[0273] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0274] Step 1:

[0275] The server generates a randomly selected set of image authentications based on the user's request for the information system they accessed. This step uses an image processing library to randomly extract images from a database. The input is a specific authentication request from the user, and the output is a set of images and instructions (e.g., "Select all images of cats").

[0276] Step 2:

[0277] The server sends the generated image authentication and instructions to the terminal. Here, the data is transmitted securely using the HTTPS protocol. The input is the image set and instructions generated in step 1, and the output is the data converted into a format viewable by the terminal.

[0278] Step 3:

[0279] The terminal displays images and instructions received from the server to the user. Using HTML and JavaScript on a web browser, the images are laid out in a user-friendly manner. Input consists of image data and instructions sent from the server. Output is a visual presentation to the user.

[0280] Step 4:

[0281] The user selects the correct image from the images displayed on the terminal according to the instructions. The user makes the selection using an input device (e.g., mouse, touch screen). The input is the instructions and image set displayed by the terminal, and the output is the selection result of the image by the user.

[0282] Step 5:

[0283] The terminal collects the user's selection result and sends it to the server. At this time, data such as the ID of the selected image and the selection time is formatted and safely sent to the server. The input is the information of the image selected by the user, and the output is the data sent to the server.

[0284] Step 6:

[0285] The server analyzes the received data and aggregates the selection frequency of each image. It uses an AI model to identify the tendency of misrecognition and optimize the learning dataset. The input is the user's selection data, and the output is the improved learning dataset.

[0286] Step 7:

[0287] The server checks whether there has been any unauthorized access during the selection process. If an abnormal pattern or behavior is detected by the security algorithm, a warning is issued and additional authentication is requested if necessary. The input is the operation history of all users, and the output is the log and alert for unauthorized detection.

[0288] (Application Example 1)

[0289] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0290] Current electronic transactions face challenges such as increasing unauthorized access, limitations in AI learning accuracy, and a lack of enhanced security. For example, there is a lack of efficient means to collect user feedback, and concerns remain about improving AI recognition accuracy. Furthermore, there is a need for methods to maintain high usability while achieving a high level of security.

[0291] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0292] In this invention, the server includes means for presenting visual authentication to collect visual selection feedback from the user, means for aggregating and processing the collected feedback information, and means for improving the learning information set of the automated learning system using the processed information. This enables efficient collection of user feedback and improvement of AI learning accuracy while maintaining a high level of security in electronic transaction settings.

[0293] "Visual selection feedback" refers to information resulting from a user's selection of visual information according to specific criteria.

[0294] "Visual authentication" is an authentication method that uses the visual recognition of an object to verify user input.

[0295] "Feedback information" is data generated based on user choices and actions, and is used to improve the system.

[0296] An "automatic learning system" is a system that has a learning algorithm that automatically improves the accuracy of recognition and judgment based on collected data.

[0297] "Unauthorized access" refers to unauthorized access to a system by someone without the proper authorization, and is considered a security threat.

[0298] "Electronic trading" refers to a form of transaction in which goods and services are bought and sold via a network.

[0299] "Security" refers to the protective measures and technologies used to safeguard information and systems from unauthorized access and data breaches.

[0300] "Usability" is a concept that refers to the ease of use and satisfaction a user experiences when using a system.

[0301] In an embodiment of this invention, a server plays a primary role. The server generates a randomly selected set of visual authentications to efficiently collect visual selection feedback from the user. This authentication set includes instructions and images that comply with specific conditions of the transaction, and is transmitted to the user's terminal.

[0302] The user's device has an interface that presents the received visual authentication to the user. Through this interface, the user selects the correct visual information according to specified conditions. Once the selection is complete, the device sends the selection result back to the server.

[0303] The server aggregates the submitted selection results and improves the automated learning system's training information set while performing processing based on the generated AI model. Specifically, it updates the AI's visual recognition algorithm based on the collected feedback information to improve recognition accuracy. In this process, in addition to the usual recognition tasks, unauthorized access detection is also performed simultaneously from a security perspective. If unauthorized access is detected, the terminal will be required to grant additional authorization.

[0304] Furthermore, this system provides a high level of security in the electronic transaction process, creating an environment where users can trade with peace of mind. For example, when a user purchases goods through online shopping, a visual authentication system is implemented before payment, and the user is instructed to "select all the fruit images."

[0305] As an example of a prompt sentence, "Please teach me a method to improve the visual recognition ability of AI while preventing unauthorized access using this image authentication system." can be considered. This facilitates the development of solutions for specific problems.

[0306] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0307] Step 1:

[0308] When the server receives a request necessary for user authentication, it starts the visual authentication process. In this process, the server generates a randomly selected visual authentication set. It receives the user ID and access status as input, creates a visual authentication set based on them, and prepares to present it to the user.

[0309] Step 2:

[0310] The server sends the generated visual authentication set to the user's terminal. Here, it includes image data and recognition conditions (e.g., "Please select all images of apples"). As output, it prepares the terminal so that it can present authentication information to the user.

[0311] Step 3:

[0312] The terminal displays the received visual authentication set to the user. The user makes an appropriate selection according to the conditions from the presented images. It receives the image data and instructions sent from the server as input, and as output, prepares the image information selected by the user.

[0313] Step 4:

[0314] The user selects images as instructed by the terminal. For example, if the instruction is "Please select all images of apples", click on the images of apples to complete the selection. The selected result is held by the terminal.

[0315] Step 5:

[0316] The terminal sends the user's selection results back to the server. It receives the user's selection data as input and prepares to send it to the server as output. No data processing is performed during this process, but reliable data transmission is required.

[0317] Step 6:

[0318] The server aggregates and analyzes the selection data sent from the terminals. Here, a generative AI model is used to process the data and improve the training information set of the automated learning system to enhance visual recognition accuracy. The server takes the selection data as input and maintains the improved training information set as output.

[0319] Step 7:

[0320] The server checks for unauthorized access based on the aggregated data. It receives selected data as input, checks for any abnormal selection patterns, and requests additional authentication if necessary. As output, it decides whether to send an instruction to the user for additional authentication.

[0321] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0322] The system of the present invention is implemented by combining image recognition, feedback collection, improvement of training data for an automated learning system, and an emotion engine. When a user accesses the system, the server first generates a set of image recognition images. These image recognition images include instructions to prompt selections that meet specific conditions.

[0323] The device displays a set of image authentication data received from the server, allowing the user to select an image according to the instructions. As the user selects an image, the device uses its built-in emotion engine to analyze the user's facial expressions. The emotion engine estimates the user's emotions based on their facial features and collects this data.

[0324] After the selection is complete, the device sends the image selection data and sentiment data to the server. The server aggregates this data and incorporates it into the training dataset of the automated learning system. Based on the feedback data and the user's sentiment data, the automated learning system's decisions become more accurate and flexible.

[0325] As a concrete example, consider a case where the system of the present invention is used in an online learning platform. When a learner answers a specific quiz, CAPTCHA-style image authentication is performed, and user sentiment data is collected during this process. This sentiment data is used to appropriately adjust the difficulty level of the quiz. If the learner finds it difficult, the overall flexibility of the system is increased so that a simpler set of images is provided for the next authentication.

[0326] The introduction of this system will not only increase the amount of information obtained from user feedback, but will also make the image recognition process itself more interactive and personalized. By utilizing emotion recognition, the quality of the user experience will improve, contributing to improved AI performance.

[0327] The following describes the processing flow.

[0328] Step 1:

[0329] When a user accesses the server, it generates a corresponding set of image recognition data. This set prompts the user to make a selection based on specific recognition criteria and includes any necessary instructions.

[0330] Step 2:

[0331] The server sends the generated image set and authentication instructions to the terminal. The images are formatted to be easily visually verifiable by the user.

[0332] Step 3:

[0333] The terminal displays the image set and instructions received from the server on the user screen. This prepares the user to make a selection.

[0334] Step 4:

[0335] The user looks at the images displayed on the device and selects the correct image according to the instructions. For example, if the instructions say, "Select all the images of dogs," the user clicks on the images of dogs.

[0336] Step 5:

[0337] While the user is selecting an image, the device activates its built-in emotion engine to analyze the user's facial expressions and estimate their emotional state.

[0338] Step 6:

[0339] Once the user has made their selection, the device sends the selected image data and analyzed sentiment data to the server.

[0340] Step 7:

[0341] The server receives the transmitted image selection information and sentiment data. Based on this, it aggregates feedback data and incorporates it into the training dataset of the automated learning system.

[0342] Step 8:

[0343] The server uses this feedback data to update the automated learning system, correcting misrecognitions and improving model biases.

[0344] Step 9:

[0345] The server uses the collected sentiment data to assess the user's stress and anxiety levels and adjust the difficulty of the next image recognition test as needed. This ensures a comfortable user experience.

[0346] (Example 2)

[0347] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0348] While conventional image recognition systems function as security measures, they suffer from a uniform user experience and a lack of flexibility in responding to the individual needs and emotions of users. Furthermore, the ineffective use of feedback data limits the improvement of the learning system's accuracy. This presents a challenge, particularly in learning support settings, where it is difficult to provide users with the most suitable content.

[0349] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0350] In this invention, the server includes means for generating a visual dataset to be presented for collecting image selection feedback from the user, means for analyzing emotional data from the user's facial expressions using a terminal device, and means for aggregating the image selection data and emotional data and analyzing them to improve the learning dataset of the automated learning device. This makes it possible to flexibly adjust the service based on the user's emotions and improve the accuracy of the learning system.

[0351] A "visual dataset" is a set of images generated according to specific conditions in order to collect image selection feedback from users.

[0352] "Emotional data" refers to data that indicates the type and intensity of emotions estimated from the user's facial expression information analyzed by the terminal device.

[0353] An "automatic learning device" is a system that can autonomously learn and improve its accuracy based on collected data.

[0354] "Feedback data" refers to data provided based on user intentions and choices, and is used to improve the system.

[0355] "Flexible adjustment" means adaptively changing the settings and outputs of a service or system based on user feedback and sentiment data.

[0356] This invention improves the adaptability and accuracy of a system based on image selection feedback and emotion data acquired through user interaction. The following hardware and software are used to implement the invention.

[0357] The server generates a visual dataset. This involves the process of creating a series of images that are presented as image authentication when a user accesses the system. There are no specific restrictions on the server platform used; any platform capable of efficiently generating and distributing image data is acceptable. In this process, the server uses image generation software to generate a set of images based on a specific algorithm.

[0358] The terminal displays images received from the server to the user, supporting the user's selection process. The terminal has an emotion engine installed that analyzes the user's facial expressions in real time, and this engine uses software that utilizes image recognition technology. Specifically, it uses a camera and CPU resources.

[0359] When a user engages in image recognition and makes a selection, the facial expression data is analyzed by an emotion engine. Emotions such as joy or confusion are estimated from the user's facial expressions and collected as emotion data. This enables data collection based on the user's reactions.

[0360] One concrete example is its use in online learning platforms. When learners take a quiz, their emotions at that moment can be analyzed through image recognition, and the difficulty of the quiz can be adjusted based on the estimated emotions. As a result, it becomes possible to provide a flexible learning experience tailored to the user.

[0361] An example of a prompt message is, "How can I use image recognition and emotion recognition together to adjust the difficulty of quizzes based on a user's learning progress on an online learning platform?" This demonstrates one method of improving the user experience through the use of generative AI models.

[0362] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0363] Step 1:

[0364] The server generates a visual dataset when a user accesses the system. Inputs include user profile information and existing data within the system. Based on this information, image generation software is used to generate a set of images containing specific selection instructions, which are then output as an image authentication set for use in authentication. Specifically, an algorithm is used to randomly or conditionally select images.

[0365] Step 2:

[0366] The terminal displays the image authentication set received from the server to the user. The input here is the image data sent from the server. This image is displayed on the screen, and screen interaction elements are prepared so that the user can select an image. The output is the configuration of the interaction interface related to the displayed image. Specifically, it generates a UI that displays the images side by side to show the appropriate choice.

[0367] Step 3:

[0368] The user selects an image from the images displayed on the device according to the instructions. The input consists of the displayed image and its instructions. The user selects one or more images based on this and confirms their selection. At this time, the user's selection information is output.

[0369] Step 4:

[0370] The device analyzes the user's facial expressions in real time using its built-in emotion engine. The input is a facial image of the user captured by the device's camera. This is processed by analysis software to estimate emotions from the facial expressions. As a result of the analysis, estimated emotion data is output. Specifically, it generates numerical emotion parameters based on eye movements and changes in mouth shape using an AI model.

[0371] Step 5:

[0372] The terminal sends the user's image selection data and emotion data to the server. The input consists of the image selection information obtained in step 3 and the emotion data obtained in step 4. This information is configured as a data packet and transferred to the server. The transmitted information becomes the output.

[0373] Step 6:

[0374] The server aggregates the received image selection data and sentiment data, and updates the dataset of the automated learning device. It uses data sent from the terminal as input. Using this information, it updates the training data based on the AI ​​algorithm and outputs the results. A specific example is analyzing the reaction trends of each user and reflecting this in the generation of subsequent datasets.

[0375] (Application Example 2)

[0376] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0377] Conventional factory automation systems have limited interaction with workers, making it difficult to respond flexibly to individual work situations and workers' emotions. This has resulted in increased worker burden without optimized work efficiency. To solve this, it is necessary to incorporate dynamic task adjustments based on the work environment and feedback based on the worker's state.

[0378] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0379] In this invention, the server includes means for providing a function to present image authentication and collect feedback from the user based on candidate selection; means for aggregating and analyzing the collected feedback information; means for using the analyzed information to enhance the learning dataset of the automatic learning function; and means for collecting image information and sentiment information from information terminals in the work environment and adjusting work conditions. This enables dynamic work adjustment according to the worker's state.

[0380] "Image verification" is a method of verifying a user's identity by having them select a specific image.

[0381] "Candidate selection" refers to the process of choosing the appropriate option from the choices presented to the user.

[0382] "Feedback information" refers to information obtained based on user choices and responses.

[0383] "Aggregation" is the process of combining and organizing multiple data sets.

[0384] "Analysis" is the process of examining collected data in detail to derive results and trends.

[0385] "Automatic learning function" refers to a function that allows a machine to learn and improve independently using the data it has collected.

[0386] A "training dataset" is a collection of data used to train a model in machine learning.

[0387] An "information terminal" is an electronic device used to collect and process information.

[0388] "Image information" refers to information that includes visual data.

[0389] "Emotional information" refers to emotional data inferred from a user's facial expressions and reactions.

[0390] "Adjusting working conditions" means appropriately changing tasks and methods according to the work environment and the condition of the workers.

[0391] This invention is a system for flexibly adjusting the operation of robots in a factory work environment. The system uses image recognition and emotion recognition to collect and analyze worker feedback information. This makes it possible to adjust working conditions according to the worker's state. A specific embodiment is shown below.

[0392] The server prompts the user to select a specific image via an image recognition system. Once the user selects an image, the terminal generates feedback information based on that selection. Simultaneously, a camera on the information terminal captures the user's facial expression, and the collected image information is analyzed using emotion recognition technology. This emotion recognition is performed using machine learning libraries and emotion engines.

[0393] The collected feedback and sentiment information is aggregated on a server. The server utilizes standard database systems and data analysis tools for data aggregation and analysis. The analyzed data is saved as a training dataset enhanced by automated learning functions, which is then used to refine future work. Specifically, data processing is performed using Python machine learning libraries and streaming frameworks.

[0394] For example, if a worker is deemed to have difficulty concentrating during a particular assembly task, the next work instructions will be simplified and the work speed adjusted. This reduces the burden on the worker and improves efficiency.

[0395] Specific examples of prompts for a generated AI model include: "Design an assistant function for a robot performing assembly work in a factory that can determine if the worker is tired and adjust the work if they are. Specifically, consider a system that combines image recognition and emotion recognition."

[0396] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0397] Step 1:

[0398] The server presents the user with image authentication. At this time, it generates candidate images and sends them to the terminal along with instructions prompting the user to make a specific selection. The input consists of the user's authentication information and the image selection options, while the output consists of instructions and a set of images.

[0399] Step 2:

[0400] The terminal displays a set of images received from the server to the user. When the user selects an image, the user's selected image information is returned to the terminal as input, and this selection data is output.

[0401] Step 3:

[0402] The device captures the user's facial expression during selection using its camera and analyzes it with its built-in emotion engine. The input is the captured image information, and the device performs facial expression analysis using emotion recognition technology, generating user emotion data as output.

[0403] Step 4:

[0404] Once the user completes their image selection, the device sends the image selection data and sentiment data to the server. The input is the selection data and sentiment data, and the output is this dataset. Data transmission takes place over the network.

[0405] Step 5:

[0406] The server aggregates the received data into a database and performs aggregation and analysis. It receives image selection data and sentiment data as input data, aggregates it using the database system, and outputs the analysis results using data analysis tools. The analysis results are then used as a training dataset for future development.

[0407] Step 6:

[0408] The server improves its automatic learning function based on the analysis results and incorporates these improvements into the next work adjustment. The input is the analyzed data, and the output is the improved training dataset. Specifically, the model is readjusted using machine learning libraries.

[0409] This entire process enables dynamic adjustment of working conditions according to the user's state, providing a more efficient work environment.

[0410] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0411] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0412] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0413] [Third Embodiment]

[0414] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0415] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0416] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0417] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0418] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0419] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0420] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0421] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0422] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0423] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0424] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0425] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0426] The present invention provides a feedback collection system utilizing image authentication, which includes a process for providing feedback through image authentication when a user accesses a specific information system. This system is realized through the cooperation of three parties: a server, a terminal, and a user.

[0427] The server first generates a randomly selected set of image verification images for the website the user has accessed. This set of images includes images that the user must select according to specific criteria. For example, the user might be instructed to "select all images of cats." The server then sends these instructions and image information to the user's device.

[0428] The device displays the received image authentication to the user. The user selects the correct image from the displayed images according to the instructions. For example, if instructed to select images of cats, the user clicks on all the images of cats. Once the user has completed their selection, the device sends the selection data back to the server.

[0429] The server analyzes the selection data collected from users and compiles data on which images were selected and how many times. Based on this compiled data, the server updates and improves the training dataset for the automated learning system. For example, if other animals that are easily mistaken for cats are repeatedly selected, this information is recorded as training data so that the AI ​​can correct the misidentification.

[0430] Furthermore, the server uses security algorithms to check for any unauthorized access during this process. If suspicious patterns are detected, it can issue a warning and, if necessary, request additional authentication.

[0431] One concrete example of this system is its use as user authentication on a shopping platform. When a user logs in, they are presented with a CAPTCHA-style image verification, through which they provide feedback to the system. This feedback is used to improve the accuracy of visual product recognition. This approach makes it possible to efficiently collect data using user behavior and correct AI bias.

[0432] The following describes the processing flow.

[0433] Step 1:

[0434] The server generates a set of images for CAPTCHA for the digital platform accessed by the user. The image set is selected based on specific criteria and is prepared along with instructions to prompt the user to make a selection.

[0435] Step 2:

[0436] The server sends the generated image set and instructions to the terminal. The images are sent in an appropriate format to ensure clear display and rapid authentication.

[0437] Step 3:

[0438] The terminal presents the received image set and instructions to the user. These are visually clear and easy to understand on the web page the user accesses.

[0439] Step 4:

[0440] The user views the images displayed on the device screen and selects the correct image according to the instructions. For example, if the instruction is "Select all the cat images," the user will click on the cat images.

[0441] Step 5:

[0442] The device sends information about the image selections made by the user to the server. This information includes the IDs and selection order of the selected images.

[0443] Step 6:

[0444] The server analyzes the received selection information and compiles data on which images were selected correctly and which were selected incorrectly. Based on these results, feedback data is generated.

[0445] Step 7:

[0446] The server uses the collected feedback data to update the training dataset of the automated learning system and improve the model's bias. This improves the AI's recognition accuracy.

[0447] Step 8:

[0448] The server analyzes user behavior through the CAPTCHA process and executes security algorithms. This verifies that there are no unnatural selection patterns or unauthorized access attempts, thus ensuring security.

[0449] (Example 1)

[0450] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0451] Conventional image recognition systems have faced challenges such as the potential for misidentification and reduced system learning efficiency due to insufficient utilization of user feedback information. Furthermore, detecting unauthorized access was difficult, potentially leading to security vulnerabilities.

[0452] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0453] In this invention, the server includes means for generating a randomly selected set of image authentications to present to the user; means for transmitting the generated image authentications and instructions to a terminal; means for collecting the user's image selection actions via the terminal and transmitting them to the server; means for analyzing and aggregating the collected feedback information; means for updating and optimizing the data set of the automatic learning device based on the analysis and aggregation results; and means for performing security management to detect abnormalities in behavioral patterns and identify unauthorized access. This makes it possible to effectively utilize user feedback to reduce misidentification, improve the learning efficiency of the system, and enable the detection of unauthorized access and strengthen security.

[0454] A "user" refers to a person who operates a device in order to solve a problem provided through image authentication.

[0455] "Image verification" refers to a collection of randomly selected images in which a user is instructed to choose a specific object.

[0456] "Terminal" refers to a device or equipment that provides image authentication to a user and transmits the user's input to a server.

[0457] A "server" refers to a central processing unit that generates image authentication data, sends instructions, analyzes feedback information, and updates the automated learning device.

[0458] "Feedback information" refers to data that includes the results of the user's selections through image authentication.

[0459] An "automatic learning device" refers to an artificial intelligence system that uses feedback information to learn and improve its recognition capabilities.

[0460] "Unauthorized access" refers to any act that disrupts the normal functioning of a system or intrudes into the system without permission.

[0461] "Security management" refers to the process of detecting unauthorized access and taking appropriate countermeasures to ensure the security of a system.

[0462] This invention relates to a system for providing feedback through image authentication when a user accesses a specific information system. The following describes embodiments for carrying out this invention.

[0463] The server generates a randomly selected set of image authentication images for the information system accessed by the user. This process utilizes an image processing library to prevent misrecognition by combining random images. For example, OpenCV could be chosen as the library.

[0464] The generated image authentication data and accompanying instructions (such as "Select all the cat images") are sent from the server to the terminal. The HTTPS protocol is used to ensure the security of this data transmission. The terminal's role is to present the image authentication data received from the server to the user. On the terminal, a web browser uses HTML and JavaScript to construct a user interface, allowing the user to select images.

[0465] The user selects the correct image from those displayed on the device, following the instructions. This operation is often performed using common input devices such as a mouse or touchscreen. After the user completes their selection, the device sends the selection result back to the server.

[0466] The server analyzes feedback information collected from users and compiles data on which images were selected and how many times. This analysis involves an automated learning system, and the dataset is updated and optimized based on the analysis results. Images with many misidentifications are subjected to more detailed training to improve the accuracy of the automated learning system. Machine learning libraries such as TensorFlow and PyTorch are used in this process.

[0467] Furthermore, the server implements security measures to identify unauthorized access. For this purpose, security algorithms are in place, and additional authentication procedures may be required if unusual patterns are detected.

[0468] A concrete example of its use is user authentication on a shopping platform. When users log in, they are provided with CAPTCHA-style image verification, and this feedback is used to improve the AI ​​model. An example of a prompt might be, "Please provide a step-by-step detailed explanation of the process of an image verification system that identifies cats."

[0469] This system leverages user feedback to improve AI accuracy and enhance security, thereby achieving stable user authentication.

[0470] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0471] Step 1:

[0472] The server generates a randomly selected set of image authentications based on the user's request for the information system they accessed. This step uses an image processing library to randomly extract images from a database. The input is a specific authentication request from the user, and the output is a set of images and instructions (e.g., "Select all images of cats").

[0473] Step 2:

[0474] The server sends the generated image authentication and instructions to the terminal. Here, the data is transmitted securely using the HTTPS protocol. The input is the image set and instructions generated in step 1, and the output is the data converted into a format viewable by the terminal.

[0475] Step 3:

[0476] The terminal displays images and instructions received from the server to the user. Using HTML and JavaScript on a web browser, the images are laid out in a user-friendly manner. Input consists of image data and instructions sent from the server. Output is a visual presentation to the user.

[0477] Step 4:

[0478] The user selects the correct image from those displayed on the terminal, following the instructions. The user makes the selection using an input device (e.g., mouse, touchscreen). The input is the instructions and image set displayed by the terminal, and the output is the image selection result by the user.

[0479] Step 5:

[0480] The terminal collects the user's selection results and sends them to the server. During this process, data such as the ID of the selected image and the selection time are formatted and securely transmitted to the server. The input is information about the image selected by the user, and the output is the data sent to the server.

[0481] Step 6:

[0482] The server analyzes the received data and aggregates the selection frequency of each image. An AI model is used to identify misrecognition trends and optimize the training dataset. The input is the user's selection data, and the output is the improved training dataset.

[0483] Step 7:

[0484] The server verifies that no unauthorized access has occurred during the selection process. If the security algorithm detects any unusual patterns or behavior, it issues a warning and requests additional authentication if necessary. Input is the operation history of all users, and output is logs and alerts for unauthorized detection.

[0485] (Application Example 1)

[0486] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0487] Current electronic transactions face challenges such as increasing unauthorized access, limitations in AI learning accuracy, and a lack of enhanced security. For example, there is a lack of efficient means to collect user feedback, and concerns remain about improving AI recognition accuracy. Furthermore, there is a need for methods to maintain high usability while achieving a high level of security.

[0488] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0489] In this invention, the server includes means for presenting visual authentication to collect visual selection feedback from the user, means for aggregating and processing the collected feedback information, and means for improving the learning information set of the automated learning system using the processed information. This enables efficient collection of user feedback and improvement of AI learning accuracy while maintaining a high level of security in electronic transaction settings.

[0490] "Visual selection feedback" refers to information resulting from a user's selection of visual information according to specific criteria.

[0491] "Visual authentication" is an authentication method that uses the visual recognition of an object to verify user input.

[0492] "Feedback information" is data generated based on user choices and actions, and is used to improve the system.

[0493] An "automatic learning system" is a system that has a learning algorithm that automatically improves the accuracy of recognition and judgment based on collected data.

[0494] "Unauthorized access" refers to unauthorized access to a system by someone without the proper authorization, and is considered a security threat.

[0495] "Electronic trading" refers to a form of transaction in which goods and services are bought and sold via a network.

[0496] "Security" refers to the protective measures and technologies used to safeguard information and systems from unauthorized access and data breaches.

[0497] "Usability" is a concept that refers to the ease of use and satisfaction a user experiences when using a system.

[0498] In an embodiment of this invention, a server plays a primary role. The server generates a randomly selected set of visual authentications to efficiently collect visual selection feedback from the user. This authentication set includes instructions and images that comply with specific conditions of the transaction, and is transmitted to the user's terminal.

[0499] The user's device has an interface that presents the received visual authentication to the user. Through this interface, the user selects the correct visual information according to specified conditions. Once the selection is complete, the device sends the selection result back to the server.

[0500] The server aggregates the submitted selection results and improves the automated learning system's training information set while performing processing based on the generated AI model. Specifically, it updates the AI's visual recognition algorithm based on the collected feedback information to improve recognition accuracy. In this process, in addition to the usual recognition tasks, unauthorized access detection is also performed simultaneously from a security perspective. If unauthorized access is detected, the terminal will be required to grant additional authorization.

[0501] Furthermore, this system provides a high level of security in the electronic transaction process, creating an environment where users can trade with peace of mind. For example, when a user purchases goods through online shopping, a visual authentication system is implemented before payment, and the user is instructed to "select all the fruit images."

[0502] An example of a prompt might be, "Please tell me how to use this image recognition system to prevent unauthorized access while improving the AI's visual recognition capabilities." This makes it easier to develop solutions for specific problems.

[0503] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0504] Step 1:

[0505] When the server receives a request necessary for user authentication, it initiates the visual authentication process. During this process, the server generates a randomly selected set of visual authentication credentials. It receives user ID and access status as input, creates the visual authentication credentials based on this information, and prepares to present them to the user.

[0506] Step 2:

[0507] The server sends the generated visual authentication set to the user's device. This includes image data and recognition conditions (e.g., "Select all images of apples"). The output is then ready for the device to present authentication information to the user.

[0508] Step 3:

[0509] The terminal displays the received visual authentication set to the user. The user makes the appropriate selection from the presented images according to the given conditions. The terminal receives image data and instructions sent from the server as input and prepares the image information selected by the user as output.

[0510] Step 4:

[0511] The user selects images as instructed on the device. For example, if the instruction is "Select all images of apples," the user clicks on the apple images to complete the selection. The selected images are then saved on the device.

[0512] Step 5:

[0513] The terminal sends the user's selection results back to the server. It receives the user's selection data as input and prepares to send it to the server as output. No data processing is performed during this process, but reliable data transmission is required.

[0514] Step 6:

[0515] The server aggregates and analyzes the selection data sent from the terminals. Here, a generative AI model is used to process the data and improve the training information set of the automated learning system to enhance visual recognition accuracy. The server takes the selection data as input and maintains the improved training information set as output.

[0516] Step 7:

[0517] The server checks for unauthorized access based on the aggregated data. It receives selected data as input, checks for any abnormal selection patterns, and requests additional authentication if necessary. As output, it decides whether to send an instruction to the user for additional authentication.

[0518] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0519] The system of the present invention is implemented by combining image recognition, feedback collection, improvement of training data for an automated learning system, and an emotion engine. When a user accesses the system, the server first generates a set of image recognition images. These image recognition images include instructions to prompt selections that meet specific conditions.

[0520] The device displays a set of image authentication data received from the server, allowing the user to select an image according to the instructions. As the user selects an image, the device uses its built-in emotion engine to analyze the user's facial expressions. The emotion engine estimates the user's emotions based on their facial features and collects this data.

[0521] After the selection is complete, the device sends the image selection data and sentiment data to the server. The server aggregates this data and incorporates it into the training dataset of the automated learning system. Based on the feedback data and the user's sentiment data, the automated learning system's decisions become more accurate and flexible.

[0522] As a concrete example, consider a case where the system of the present invention is used in an online learning platform. When a learner answers a specific quiz, CAPTCHA-style image authentication is performed, and user sentiment data is collected during this process. This sentiment data is used to appropriately adjust the difficulty level of the quiz. If the learner finds it difficult, the overall flexibility of the system is increased so that a simpler set of images is provided for the next authentication.

[0523] The introduction of this system will not only increase the amount of information obtained from user feedback, but will also make the image recognition process itself more interactive and personalized. By utilizing emotion recognition, the quality of the user experience will improve, contributing to improved AI performance.

[0524] The following describes the processing flow.

[0525] Step 1:

[0526] When a user accesses the server, it generates a corresponding set of image recognition data. This set prompts the user to make a selection based on specific recognition criteria and includes any necessary instructions.

[0527] Step 2:

[0528] The server sends the generated image set and authentication instructions to the terminal. The images are formatted to be easily visually verifiable by the user.

[0529] Step 3:

[0530] The terminal displays the image set and instructions received from the server on the user screen. This prepares the user to make a selection.

[0531] Step 4:

[0532] The user looks at the images displayed on the device and selects the correct image according to the instructions. For example, if the instructions say, "Select all the images of dogs," the user clicks on the images of dogs.

[0533] Step 5:

[0534] While the user is selecting an image, the device activates its built-in emotion engine to analyze the user's facial expressions and estimate their emotional state.

[0535] Step 6:

[0536] Once the user has made their selection, the device sends the selected image data and analyzed sentiment data to the server.

[0537] Step 7:

[0538] The server receives the transmitted image selection information and sentiment data. Based on this, it aggregates feedback data and incorporates it into the training dataset of the automated learning system.

[0539] Step 8:

[0540] The server uses this feedback data to update the automated learning system, correcting misrecognitions and improving model biases.

[0541] Step 9:

[0542] The server uses the collected sentiment data to assess the user's stress and anxiety levels and adjust the difficulty of the next image recognition test as needed. This ensures a comfortable user experience.

[0543] (Example 2)

[0544] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0545] While conventional image recognition systems function as security measures, they suffer from a uniform user experience and a lack of flexibility in responding to the individual needs and emotions of users. Furthermore, the ineffective use of feedback data limits the improvement of the learning system's accuracy. This presents a challenge, particularly in learning support settings, where it is difficult to provide users with the most suitable content.

[0546] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0547] In this invention, the server includes means for generating a visual dataset to be presented for collecting image selection feedback from the user, means for analyzing emotional data from the user's facial expressions using a terminal device, and means for aggregating the image selection data and emotional data and analyzing them to improve the learning dataset of the automated learning device. This makes it possible to flexibly adjust the service based on the user's emotions and improve the accuracy of the learning system.

[0548] A "visual dataset" is a set of images generated according to specific conditions in order to collect image selection feedback from users.

[0549] "Emotional data" refers to data that indicates the type and intensity of emotions estimated from the user's facial expression information analyzed by the terminal device.

[0550] An "automatic learning device" is a system that can autonomously learn and improve its accuracy based on collected data.

[0551] "Feedback data" refers to data provided based on user intentions and choices, and is used to improve the system.

[0552] "Flexible adjustment" means adaptively changing the settings and outputs of a service or system based on user feedback and sentiment data.

[0553] This invention improves the adaptability and accuracy of a system based on image selection feedback and emotion data acquired through user interaction. The following hardware and software are used to implement the invention.

[0554] The server generates a visual dataset. This involves the process of creating a series of images that are presented as image authentication when a user accesses the system. There are no specific restrictions on the server platform used; any platform capable of efficiently generating and distributing image data is acceptable. In this process, the server uses image generation software to generate a set of images based on a specific algorithm.

[0555] The terminal displays images received from the server to the user, supporting the user's selection process. The terminal has an emotion engine installed that analyzes the user's facial expressions in real time, and this engine uses software that utilizes image recognition technology. Specifically, it uses a camera and CPU resources.

[0556] When a user engages in image recognition and makes a selection, the facial expression data is analyzed by an emotion engine. Emotions such as joy or confusion are estimated from the user's facial expressions and collected as emotion data. This enables data collection based on the user's reactions.

[0557] One concrete example is its use in online learning platforms. When learners take a quiz, their emotions at that moment can be analyzed through image recognition, and the difficulty of the quiz can be adjusted based on the estimated emotions. As a result, it becomes possible to provide a flexible learning experience tailored to the user.

[0558] An example of a prompt message is, "How can I use image recognition and emotion recognition together to adjust the difficulty of quizzes based on a user's learning progress on an online learning platform?" This demonstrates one method of improving the user experience through the use of generative AI models.

[0559] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0560] Step 1:

[0561] The server generates a visual dataset when a user accesses the system. Inputs include user profile information and existing data within the system. Based on this information, image generation software is used to generate a set of images containing specific selection instructions, which are then output as an image authentication set for use in authentication. Specifically, an algorithm is used to randomly or conditionally select images.

[0562] Step 2:

[0563] The terminal displays the image authentication set received from the server to the user. The input here is the image data sent from the server. This image is displayed on the screen, and screen interaction elements are prepared so that the user can select an image. The output is the configuration of the interaction interface related to the displayed image. Specifically, it generates a UI that displays the images side by side to show the appropriate choice.

[0564] Step 3:

[0565] The user selects an image from the images displayed on the device according to the instructions. The input consists of the displayed image and its instructions. The user selects one or more images based on this and confirms their selection. At this time, the user's selection information is output.

[0566] Step 4:

[0567] The device analyzes the user's facial expressions in real time using its built-in emotion engine. The input is a facial image of the user captured by the device's camera. This is processed by analysis software to estimate emotions from the facial expressions. As a result of the analysis, estimated emotion data is output. Specifically, it generates numerical emotion parameters based on eye movements and changes in mouth shape using an AI model.

[0568] Step 5:

[0569] The terminal sends the user's image selection data and emotion data to the server. The input consists of the image selection information obtained in step 3 and the emotion data obtained in step 4. This information is configured as a data packet and transferred to the server. The transmitted information becomes the output.

[0570] Step 6:

[0571] The server aggregates the received image selection data and sentiment data, and updates the dataset of the automated learning device. It uses data sent from the terminal as input. Using this information, it updates the training data based on the AI ​​algorithm and outputs the results. A specific example is analyzing the reaction trends of each user and reflecting this in the generation of subsequent datasets.

[0572] (Application Example 2)

[0573] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0574] Conventional factory automation systems have limited interaction with workers, making it difficult to respond flexibly to individual work situations and workers' emotions. This has resulted in increased worker burden without optimized work efficiency. To solve this, it is necessary to incorporate dynamic task adjustments based on the work environment and feedback based on the worker's state.

[0575] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0576] In this invention, the server includes means for providing a function to present image authentication and collect feedback from the user based on candidate selection; means for aggregating and analyzing the collected feedback information; means for using the analyzed information to enhance the learning dataset of the automatic learning function; and means for collecting image information and sentiment information from information terminals in the work environment and adjusting work conditions. This enables dynamic work adjustment according to the worker's state.

[0577] "Image verification" is a method of verifying a user's identity by having them select a specific image.

[0578] "Candidate selection" refers to the process of choosing the appropriate option from the choices presented to the user.

[0579] "Feedback information" refers to information obtained based on user choices and responses.

[0580] "Aggregation" is the process of combining and organizing multiple data sets.

[0581] "Analysis" is the process of examining collected data in detail to derive results and trends.

[0582] "Automatic learning function" refers to a function that allows a machine to learn and improve independently using the data it has collected.

[0583] A "training dataset" is a collection of data used to train a model in machine learning.

[0584] An "information terminal" is an electronic device used to collect and process information.

[0585] "Image information" refers to information that includes visual data.

[0586] "Emotional information" refers to emotional data inferred from a user's facial expressions and reactions.

[0587] "Adjusting working conditions" means appropriately changing tasks and methods according to the work environment and the condition of the workers.

[0588] This invention is a system for flexibly adjusting the operation of robots in a factory work environment. The system uses image recognition and emotion recognition to collect and analyze worker feedback information. This makes it possible to adjust working conditions according to the worker's state. A specific embodiment is shown below.

[0589] The server prompts the user to select a specific image via an image recognition system. Once the user selects an image, the terminal generates feedback information based on that selection. Simultaneously, a camera on the information terminal captures the user's facial expression, and the collected image information is analyzed using emotion recognition technology. This emotion recognition is performed using machine learning libraries and emotion engines.

[0590] The collected feedback and sentiment information is aggregated on a server. The server utilizes standard database systems and data analysis tools for data aggregation and analysis. The analyzed data is saved as a training dataset enhanced by automated learning functions, which is then used to refine future work. Specifically, data processing is performed using Python machine learning libraries and streaming frameworks.

[0591] For example, if a worker is deemed to have difficulty concentrating during a particular assembly task, the next work instructions will be simplified and the work speed adjusted. This reduces the burden on the worker and improves efficiency.

[0592] Specific examples of prompts for a generated AI model include: "Design an assistant function for a robot performing assembly work in a factory that can determine if the worker is tired and adjust the work if they are. Specifically, consider a system that combines image recognition and emotion recognition."

[0593] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0594] Step 1:

[0595] The server presents the user with image authentication. At this time, it generates candidate images and sends them to the terminal along with instructions prompting the user to make a specific selection. The input consists of the user's authentication information and the image selection options, while the output consists of instructions and a set of images.

[0596] Step 2:

[0597] The terminal displays a set of images received from the server to the user. When the user selects an image, the user's selected image information is returned to the terminal as input, and this selection data is output.

[0598] Step 3:

[0599] The device captures the user's facial expression during selection using its camera and analyzes it with its built-in emotion engine. The input is the captured image information, and the device performs facial expression analysis using emotion recognition technology, generating user emotion data as output.

[0600] Step 4:

[0601] Once the user completes their image selection, the device sends the image selection data and sentiment data to the server. The input is the selection data and sentiment data, and the output is this dataset. Data transmission takes place over the network.

[0602] Step 5:

[0603] The server aggregates the received data into a database and performs aggregation and analysis. It receives image selection data and sentiment data as input data, aggregates it using the database system, and outputs the analysis results using data analysis tools. The analysis results are then used as a training dataset for future development.

[0604] Step 6:

[0605] The server improves its automatic learning function based on the analysis results and incorporates these improvements into the next work adjustment. The input is the analyzed data, and the output is the improved training dataset. Specifically, the model is readjusted using machine learning libraries.

[0606] This entire process enables dynamic adjustment of working conditions according to the user's state, providing a more efficient work environment.

[0607] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0608] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0609] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0610] [Fourth Embodiment]

[0611] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0612] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0613] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0614] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0615] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0616] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0617] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0618] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0619] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0620] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0621] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0622] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0623] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0624] The present invention provides a feedback collection system utilizing image authentication, which includes a process for providing feedback through image authentication when a user accesses a specific information system. This system is realized through the cooperation of three parties: a server, a terminal, and a user.

[0625] The server first generates a randomly selected set of image verification images for the website the user has accessed. This set of images includes images that the user must select according to specific criteria. For example, the user might be instructed to "select all images of cats." The server then sends these instructions and image information to the user's device.

[0626] The device displays the received image authentication to the user. The user selects the correct image from the displayed images according to the instructions. For example, if instructed to select images of cats, the user clicks on all the images of cats. Once the user has completed their selection, the device sends the selection data back to the server.

[0627] The server analyzes the selection data collected from users and compiles data on which images were selected and how many times. Based on this compiled data, the server updates and improves the training dataset for the automated learning system. For example, if other animals that are easily mistaken for cats are repeatedly selected, this information is recorded as training data so that the AI ​​can correct the misidentification.

[0628] Furthermore, the server uses security algorithms to check for any unauthorized access during this process. If suspicious patterns are detected, it can issue a warning and, if necessary, request additional authentication.

[0629] One concrete example of this system is its use as user authentication on a shopping platform. When a user logs in, they are presented with a CAPTCHA-style image verification, through which they provide feedback to the system. This feedback is used to improve the accuracy of visual product recognition. This approach makes it possible to efficiently collect data using user behavior and correct AI bias.

[0630] The following describes the processing flow.

[0631] Step 1:

[0632] The server generates a set of images for CAPTCHA for the digital platform accessed by the user. The image set is selected based on specific criteria and is prepared along with instructions to prompt the user to make a selection.

[0633] Step 2:

[0634] The server sends the generated image set and instructions to the terminal. The images are sent in an appropriate format to ensure clear display and rapid authentication.

[0635] Step 3:

[0636] The terminal presents the received image set and instructions to the user. These are visually clear and easy to understand on the web page the user accesses.

[0637] Step 4:

[0638] The user views the images displayed on the device screen and selects the correct image according to the instructions. For example, if the instruction is "Select all the cat images," the user will click on the cat images.

[0639] Step 5:

[0640] The device sends information about the image selections made by the user to the server. This information includes the IDs and selection order of the selected images.

[0641] Step 6:

[0642] The server analyzes the received selection information and compiles data on which images were selected correctly and which were selected incorrectly. Based on these results, feedback data is generated.

[0643] Step 7:

[0644] The server uses the collected feedback data to update the training dataset of the automated learning system and improve the model's bias. This improves the AI's recognition accuracy.

[0645] Step 8:

[0646] The server analyzes user behavior through the CAPTCHA process and executes security algorithms. This verifies that there are no unnatural selection patterns or unauthorized access attempts, thus ensuring security.

[0647] (Example 1)

[0648] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0649] Conventional image recognition systems have faced challenges such as the potential for misidentification and reduced system learning efficiency due to insufficient utilization of user feedback information. Furthermore, detecting unauthorized access was difficult, potentially leading to security vulnerabilities.

[0650] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0651] In this invention, the server includes means for generating a randomly selected set of image authentications to present to the user; means for transmitting the generated image authentications and instructions to a terminal; means for collecting the user's image selection actions via the terminal and transmitting them to the server; means for analyzing and aggregating the collected feedback information; means for updating and optimizing the data set of the automatic learning device based on the analysis and aggregation results; and means for performing security management to detect abnormalities in behavioral patterns and identify unauthorized access. This makes it possible to effectively utilize user feedback to reduce misidentification, improve the learning efficiency of the system, and enable the detection of unauthorized access and strengthen security.

[0652] A "user" refers to a person who operates a device in order to solve a problem provided through image authentication.

[0653] "Image verification" refers to a collection of randomly selected images in which a user is instructed to choose a specific object.

[0654] "Terminal" refers to a device or equipment that provides image authentication to a user and transmits the user's input to a server.

[0655] A "server" refers to a central processing unit that generates image authentication data, sends instructions, analyzes feedback information, and updates the automated learning device.

[0656] "Feedback information" refers to data that includes the results of the user's selections through image authentication.

[0657] An "automatic learning device" refers to an artificial intelligence system that uses feedback information to learn and improve its recognition capabilities.

[0658] "Unauthorized access" refers to any act that disrupts the normal functioning of a system or intrudes into the system without permission.

[0659] "Security management" refers to the process of detecting unauthorized access and taking appropriate countermeasures to ensure the security of a system.

[0660] This invention relates to a system for providing feedback through image authentication when a user accesses a specific information system. The following describes embodiments for carrying out this invention.

[0661] The server generates a randomly selected set of image authentication images for the information system accessed by the user. This process utilizes an image processing library to prevent misrecognition by combining random images. For example, OpenCV could be chosen as the library.

[0662] The generated image authentication data and accompanying instructions (such as "Select all the cat images") are sent from the server to the terminal. The HTTPS protocol is used to ensure the security of this data transmission. The terminal's role is to present the image authentication data received from the server to the user. On the terminal, a web browser uses HTML and JavaScript to construct a user interface, allowing the user to select images.

[0663] The user selects the correct image from those displayed on the device, following the instructions. This operation is often performed using common input devices such as a mouse or touchscreen. After the user completes their selection, the device sends the selection result back to the server.

[0664] The server analyzes feedback information collected from users and compiles data on which images were selected and how many times. This analysis involves an automated learning system, and the dataset is updated and optimized based on the analysis results. Images with many misidentifications are subjected to more detailed training to improve the accuracy of the automated learning system. Machine learning libraries such as TensorFlow and PyTorch are used in this process.

[0665] Furthermore, the server implements security measures to identify unauthorized access. For this purpose, security algorithms are in place, and additional authentication procedures may be required if unusual patterns are detected.

[0666] A concrete example of its use is user authentication on a shopping platform. When users log in, they are provided with CAPTCHA-style image verification, and this feedback is used to improve the AI ​​model. An example of a prompt might be, "Please provide a step-by-step detailed explanation of the process of an image verification system that identifies cats."

[0667] This system leverages user feedback to improve AI accuracy and enhance security, thereby achieving stable user authentication.

[0668] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0669] Step 1:

[0670] The server generates a randomly selected set of image authentications based on the user's request for the information system they accessed. This step uses an image processing library to randomly extract images from a database. The input is a specific authentication request from the user, and the output is a set of images and instructions (e.g., "Select all images of cats").

[0671] Step 2:

[0672] The server sends the generated image authentication and instructions to the terminal. Here, the data is transmitted securely using the HTTPS protocol. The input is the image set and instructions generated in step 1, and the output is the data converted into a format viewable by the terminal.

[0673] Step 3:

[0674] The terminal displays images and instructions received from the server to the user. Using HTML and JavaScript on a web browser, the images are laid out in a user-friendly manner. Input consists of image data and instructions sent from the server. Output is a visual presentation to the user.

[0675] Step 4:

[0676] The user selects the correct image from those displayed on the terminal, following the instructions. The user makes the selection using an input device (e.g., mouse, touchscreen). The input is the instructions and image set displayed by the terminal, and the output is the image selection result by the user.

[0677] Step 5:

[0678] The terminal collects the user's selection results and sends them to the server. During this process, data such as the ID of the selected image and the selection time are formatted and securely transmitted to the server. The input is information about the image selected by the user, and the output is the data sent to the server.

[0679] Step 6:

[0680] The server analyzes the received data and aggregates the selection frequency of each image. An AI model is used to identify misrecognition trends and optimize the training dataset. The input is the user's selection data, and the output is the improved training dataset.

[0681] Step 7:

[0682] The server verifies that no unauthorized access has occurred during the selection process. If the security algorithm detects any unusual patterns or behavior, it issues a warning and requests additional authentication if necessary. Input is the operation history of all users, and output is logs and alerts for unauthorized detection.

[0683] (Application Example 1)

[0684] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0685] Current electronic transactions face challenges such as increasing unauthorized access, limitations in AI learning accuracy, and a lack of enhanced security. For example, there is a lack of efficient means to collect user feedback, and concerns remain about improving AI recognition accuracy. Furthermore, there is a need for methods to maintain high usability while achieving a high level of security.

[0686] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0687] In this invention, the server includes means for presenting visual authentication to collect visual selection feedback from the user, means for aggregating and processing the collected feedback information, and means for improving the learning information set of the automated learning system using the processed information. This enables efficient collection of user feedback and improvement of AI learning accuracy while maintaining a high level of security in electronic transaction settings.

[0688] "Visual selection feedback" refers to information resulting from a user's selection of visual information according to specific criteria.

[0689] "Visual authentication" is an authentication method that uses the visual recognition of an object to verify user input.

[0690] "Feedback information" is data generated based on user choices and actions, and is used to improve the system.

[0691] An "automatic learning system" is a system that has a learning algorithm that automatically improves the accuracy of recognition and judgment based on collected data.

[0692] "Unauthorized access" refers to unauthorized access to a system by someone without the proper authorization, and is considered a security threat.

[0693] "Electronic trading" refers to a form of transaction in which goods and services are bought and sold via a network.

[0694] "Security" refers to the protective measures and technologies used to safeguard information and systems from unauthorized access and data breaches.

[0695] "Usability" is a concept that refers to the ease of use and satisfaction a user experiences when using a system.

[0696] In an embodiment of this invention, a server plays a primary role. The server generates a randomly selected set of visual authentications to efficiently collect visual selection feedback from the user. This authentication set includes instructions and images that comply with specific conditions of the transaction, and is transmitted to the user's terminal.

[0697] The user's device has an interface that presents the received visual authentication to the user. Through this interface, the user selects the correct visual information according to specified conditions. Once the selection is complete, the device sends the selection result back to the server.

[0698] The server aggregates the submitted selection results and improves the automated learning system's training information set while performing processing based on the generated AI model. Specifically, it updates the AI's visual recognition algorithm based on the collected feedback information to improve recognition accuracy. In this process, in addition to the usual recognition tasks, unauthorized access detection is also performed simultaneously from a security perspective. If unauthorized access is detected, the terminal will be required to grant additional authorization.

[0699] Furthermore, this system provides a high level of security in the electronic transaction process, creating an environment where users can trade with peace of mind. For example, when a user purchases goods through online shopping, a visual authentication system is implemented before payment, and the user is instructed to "select all the fruit images."

[0700] An example of a prompt might be, "Please tell me how to use this image recognition system to prevent unauthorized access while improving the AI's visual recognition capabilities." This makes it easier to develop solutions for specific problems.

[0701] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0702] Step 1:

[0703] When the server receives a request necessary for user authentication, it initiates the visual authentication process. During this process, the server generates a randomly selected set of visual authentication credentials. It receives user ID and access status as input, creates the visual authentication credentials based on this information, and prepares to present them to the user.

[0704] Step 2:

[0705] The server sends the generated visual authentication set to the user's device. This includes image data and recognition conditions (e.g., "Select all images of apples"). The output is then ready for the device to present authentication information to the user.

[0706] Step 3:

[0707] The terminal displays the received visual authentication set to the user. The user makes the appropriate selection from the presented images according to the given conditions. The terminal receives image data and instructions sent from the server as input and prepares the image information selected by the user as output.

[0708] Step 4:

[0709] The user selects images as instructed on the device. For example, if the instruction is "Select all images of apples," the user clicks on the apple images to complete the selection. The selected images are then saved on the device.

[0710] Step 5:

[0711] The terminal sends the user's selection results back to the server. It receives the user's selection data as input and prepares to send it to the server as output. No data processing is performed during this process, but reliable data transmission is required.

[0712] Step 6:

[0713] The server aggregates and analyzes the selection data sent from the terminals. Here, a generative AI model is used to process the data and improve the training information set of the automated learning system to enhance visual recognition accuracy. The server takes the selection data as input and maintains the improved training information set as output.

[0714] Step 7:

[0715] The server checks for unauthorized access based on the aggregated data. It receives selected data as input, checks for any abnormal selection patterns, and requests additional authentication if necessary. As output, it decides whether to send an instruction to the user for additional authentication.

[0716] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0717] The system of the present invention is implemented by combining image recognition, feedback collection, improvement of training data for an automated learning system, and an emotion engine. When a user accesses the system, the server first generates a set of image recognition images. These image recognition images include instructions to prompt selections that meet specific conditions.

[0718] The device displays a set of image authentication data received from the server, allowing the user to select an image according to the instructions. As the user selects an image, the device uses its built-in emotion engine to analyze the user's facial expressions. The emotion engine estimates the user's emotions based on their facial features and collects this data.

[0719] After the selection is complete, the device sends the image selection data and sentiment data to the server. The server aggregates this data and incorporates it into the training dataset of the automated learning system. Based on the feedback data and the user's sentiment data, the automated learning system's decisions become more accurate and flexible.

[0720] As a concrete example, consider a case where the system of the present invention is used in an online learning platform. When a learner answers a specific quiz, CAPTCHA-style image authentication is performed, and user sentiment data is collected during this process. This sentiment data is used to appropriately adjust the difficulty level of the quiz. If the learner finds it difficult, the overall flexibility of the system is increased so that a simpler set of images is provided for the next authentication.

[0721] The introduction of this system will not only increase the amount of information obtained from user feedback, but will also make the image recognition process itself more interactive and personalized. By utilizing emotion recognition, the quality of the user experience will improve, contributing to improved AI performance.

[0722] The following describes the processing flow.

[0723] Step 1:

[0724] When a user accesses the server, it generates a corresponding set of image recognition data. This set prompts the user to make a selection based on specific recognition criteria and includes any necessary instructions.

[0725] Step 2:

[0726] The server sends the generated image set and authentication instructions to the terminal. The images are formatted to be easily visually verifiable by the user.

[0727] Step 3:

[0728] The terminal displays the image set and instructions received from the server on the user screen. This prepares the user to make a selection.

[0729] Step 4:

[0730] The user looks at the images displayed on the device and selects the correct image according to the instructions. For example, if the instructions say, "Select all the images of dogs," the user clicks on the images of dogs.

[0731] Step 5:

[0732] While the user is selecting an image, the device activates its built-in emotion engine to analyze the user's facial expressions and estimate their emotional state.

[0733] Step 6:

[0734] Once the user has made their selection, the device sends the selected image data and analyzed sentiment data to the server.

[0735] Step 7:

[0736] The server receives the transmitted image selection information and sentiment data. Based on this, it aggregates feedback data and incorporates it into the training dataset of the automated learning system.

[0737] Step 8:

[0738] The server uses this feedback data to update the automated learning system, correcting misrecognitions and improving model biases.

[0739] Step 9:

[0740] The server uses the collected sentiment data to assess the user's stress and anxiety levels and adjust the difficulty of the next image recognition test as needed. This ensures a comfortable user experience.

[0741] (Example 2)

[0742] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0743] While conventional image recognition systems function as security measures, they suffer from a uniform user experience and a lack of flexibility in responding to the individual needs and emotions of users. Furthermore, the ineffective use of feedback data limits the improvement of the learning system's accuracy. This presents a challenge, particularly in learning support settings, where it is difficult to provide users with the most suitable content.

[0744] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0745] In this invention, the server includes means for generating a visual dataset to be presented for collecting image selection feedback from the user, means for analyzing emotional data from the user's facial expressions using a terminal device, and means for aggregating the image selection data and emotional data and analyzing them to improve the learning dataset of the automated learning device. This makes it possible to flexibly adjust the service based on the user's emotions and improve the accuracy of the learning system.

[0746] A "visual dataset" is a set of images generated according to specific conditions in order to collect image selection feedback from users.

[0747] "Emotional data" refers to data that indicates the type and intensity of emotions estimated from the user's facial expression information analyzed by the terminal device.

[0748] An "automatic learning device" is a system that can autonomously learn and improve its accuracy based on collected data.

[0749] "Feedback data" refers to data provided based on user intentions and choices, and is used to improve the system.

[0750] "Flexible adjustment" means adaptively changing the settings and outputs of a service or system based on user feedback and sentiment data.

[0751] This invention improves the adaptability and accuracy of a system based on image selection feedback and emotion data acquired through user interaction. The following hardware and software are used to implement the invention.

[0752] The server generates a visual dataset. This involves the process of creating a series of images that are presented as image authentication when a user accesses the system. There are no specific restrictions on the server platform used; any platform capable of efficiently generating and distributing image data is acceptable. In this process, the server uses image generation software to generate a set of images based on a specific algorithm.

[0753] The terminal displays images received from the server to the user, supporting the user's selection process. The terminal has an emotion engine installed that analyzes the user's facial expressions in real time, and this engine uses software that utilizes image recognition technology. Specifically, it uses a camera and CPU resources.

[0754] When a user engages in image recognition and makes a selection, the facial expression data is analyzed by an emotion engine. Emotions such as joy or confusion are estimated from the user's facial expressions and collected as emotion data. This enables data collection based on the user's reactions.

[0755] One concrete example is its use in online learning platforms. When learners take a quiz, their emotions at that moment can be analyzed through image recognition, and the difficulty of the quiz can be adjusted based on the estimated emotions. As a result, it becomes possible to provide a flexible learning experience tailored to the user.

[0756] An example of a prompt message is, "How can I use image recognition and emotion recognition together to adjust the difficulty of quizzes based on a user's learning progress on an online learning platform?" This demonstrates one method of improving the user experience through the use of generative AI models.

[0757] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0758] Step 1:

[0759] The server generates a visual dataset when a user accesses the system. Inputs include user profile information and existing data within the system. Based on this information, image generation software is used to generate a set of images containing specific selection instructions, which are then output as an image authentication set for use in authentication. Specifically, an algorithm is used to randomly or conditionally select images.

[0760] Step 2:

[0761] The terminal displays the image authentication set received from the server to the user. The input here is the image data sent from the server. This image is displayed on the screen, and screen interaction elements are prepared so that the user can select an image. The output is the configuration of the interaction interface related to the displayed image. Specifically, it generates a UI that displays the images side by side to show the appropriate choice.

[0762] Step 3:

[0763] The user selects an image from the images displayed on the device according to the instructions. The input consists of the displayed image and its instructions. The user selects one or more images based on this and confirms their selection. At this time, the user's selection information is output.

[0764] Step 4:

[0765] The device analyzes the user's facial expressions in real time using its built-in emotion engine. The input is a facial image of the user captured by the device's camera. This is processed by analysis software to estimate emotions from the facial expressions. As a result of the analysis, estimated emotion data is output. Specifically, it generates numerical emotion parameters based on eye movements and changes in mouth shape using an AI model.

[0766] Step 5:

[0767] The terminal sends the user's image selection data and emotion data to the server. The input consists of the image selection information obtained in step 3 and the emotion data obtained in step 4. This information is configured as a data packet and transferred to the server. The transmitted information becomes the output.

[0768] Step 6:

[0769] The server aggregates the received image selection data and sentiment data, and updates the dataset of the automated learning device. It uses data sent from the terminal as input. Using this information, it updates the training data based on the AI ​​algorithm and outputs the results. A specific example is analyzing the reaction trends of each user and reflecting this in the generation of subsequent datasets.

[0770] (Application Example 2)

[0771] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0772] Conventional factory automation systems have limited interaction with workers, making it difficult to respond flexibly to individual work situations and workers' emotions. This has resulted in increased worker burden without optimized work efficiency. To solve this, it is necessary to incorporate dynamic task adjustments based on the work environment and feedback based on the worker's state.

[0773] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0774] In this invention, the server includes means for providing a function to present image authentication and collect feedback from the user based on candidate selection; means for aggregating and analyzing the collected feedback information; means for using the analyzed information to enhance the learning dataset of the automatic learning function; and means for collecting image information and sentiment information from information terminals in the work environment and adjusting work conditions. This enables dynamic work adjustment according to the worker's state.

[0775] "Image verification" is a method of verifying a user's identity by having them select a specific image.

[0776] "Candidate selection" refers to the process of choosing the appropriate option from the choices presented to the user.

[0777] "Feedback information" refers to information obtained based on user choices and responses.

[0778] "Aggregation" is the process of combining and organizing multiple data sets.

[0779] "Analysis" is the process of examining collected data in detail to derive results and trends.

[0780] "Automatic learning function" refers to a function that allows a machine to learn and improve independently using the data it has collected.

[0781] A "training dataset" is a collection of data used to train a model in machine learning.

[0782] An "information terminal" is an electronic device used to collect and process information.

[0783] "Image information" refers to information that includes visual data.

[0784] "Emotional information" refers to emotional data inferred from a user's facial expressions and reactions.

[0785] "Adjusting working conditions" means appropriately changing tasks and methods according to the work environment and the condition of the workers.

[0786] This invention is a system for flexibly adjusting the operation of robots in a factory work environment. The system uses image recognition and emotion recognition to collect and analyze worker feedback information. This makes it possible to adjust working conditions according to the worker's state. A specific embodiment is shown below.

[0787] The server prompts the user to select a specific image via an image recognition system. Once the user selects an image, the terminal generates feedback information based on that selection. Simultaneously, a camera on the information terminal captures the user's facial expression, and the collected image information is analyzed using emotion recognition technology. This emotion recognition is performed using machine learning libraries and emotion engines.

[0788] The collected feedback and sentiment information is aggregated on a server. The server utilizes standard database systems and data analysis tools for data aggregation and analysis. The analyzed data is saved as a training dataset enhanced by automated learning functions, which is then used to refine future work. Specifically, data processing is performed using Python machine learning libraries and streaming frameworks.

[0789] For example, if a worker is deemed to have difficulty concentrating during a particular assembly task, the next work instructions will be simplified and the work speed adjusted. This reduces the burden on the worker and improves efficiency.

[0790] Specific examples of prompts for a generated AI model include: "Design an assistant function for a robot performing assembly work in a factory that can determine if the worker is tired and adjust the work if they are. Specifically, consider a system that combines image recognition and emotion recognition."

[0791] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0792] Step 1:

[0793] The server presents the user with image authentication. At this time, it generates candidate images and sends them to the terminal along with instructions prompting the user to make a specific selection. The input consists of the user's authentication information and the image selection options, while the output consists of instructions and a set of images.

[0794] Step 2:

[0795] The terminal displays a set of images received from the server to the user. When the user selects an image, the user's selected image information is returned to the terminal as input, and this selection data is output.

[0796] Step 3:

[0797] The device captures the user's facial expression during selection using its camera and analyzes it with its built-in emotion engine. The input is the captured image information, and the device performs facial expression analysis using emotion recognition technology, generating user emotion data as output.

[0798] Step 4:

[0799] Once the user completes their image selection, the device sends the image selection data and sentiment data to the server. The input is the selection data and sentiment data, and the output is this dataset. Data transmission takes place over the network.

[0800] Step 5:

[0801] The server aggregates the received data into a database and performs aggregation and analysis. It receives image selection data and sentiment data as input data, aggregates it using the database system, and outputs the analysis results using data analysis tools. The analysis results are then used as a training dataset for future development.

[0802] Step 6:

[0803] The server improves its automatic learning function based on the analysis results and incorporates these improvements into the next work adjustment. The input is the analyzed data, and the output is the improved training dataset. Specifically, the model is readjusted using machine learning libraries.

[0804] This entire process enables dynamic adjustment of working conditions according to the user's state, providing a more efficient work environment.

[0805] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0806] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0807] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0808] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0809] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0810] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0811] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0812] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0813] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0814] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0815] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0816] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0817] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0818] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0819] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0820] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0821] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0822] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0823] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0824] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0825] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0826] The following is further disclosed regarding the embodiments described above.

[0827] (Claim 1)

[0828] A means of presenting image authentication in order to collect image selection feedback from users,

[0829] A means of aggregating and analyzing the collected feedback data,

[0830] A means of improving the training dataset of an automated learning system using the analyzed data,

[0831] Means for implementing security measures to distinguish between humans and automated programs,

[0832] A system that includes this.

[0833] (Claim 2)

[0834] The system according to claim 1, further comprising means for causing the automatic learning system to retrain based on the analysis of the aforementioned feedback data.

[0835] (Claim 3)

[0836] The system according to claim 1, further comprising means for detecting unauthorized access based on the user's image authentication behavior.

[0837] "Example 1"

[0838] (Claim 1)

[0839] A means for generating a randomly selected set of image authentications to be presented to a user,

[0840] A means for transmitting the generated image authentication and instructions to the terminal,

[0841] A means for collecting the user's image selection actions via the terminal and sending them to a server,

[0842] A means of analyzing and aggregating the collected feedback information,

[0843] A means for updating and optimizing the data set of an automated learning device based on the analysis and aggregation results,

[0844] A means of implementing security measures to detect abnormal behavioral patterns and identify unauthorized access,

[0845] A system that includes this.

[0846] (Claim 2)

[0847] The system according to claim 1, further comprising means for retraining an automated learning device based on the results of the analysis of the aforementioned feedback information.

[0848] (Claim 3)

[0849] The system according to claim 1, further comprising means for identifying unauthorized access based on the user's image authentication operation history.

[0850] "Application Example 1"

[0851] (Claim 1)

[0852] A means of presenting visual authentication to collect visual selection feedback from users,

[0853] A means for aggregating and processing the collected feedback information,

[0854] A means of improving the learning information set of an automated learning system using processed information,

[0855] A means for implementing a process to detect unauthorized access based on the user's authentication behavior,

[0856] Means to enhance security when conducting electronic transactions,

[0857] A device that includes this.

[0858] (Claim 2)

[0859] The apparatus according to claim 1, further comprising means for causing the automatic learning system to retrain based on the processing of the aforementioned feedback information.

[0860] (Claim 3)

[0861] The apparatus according to claim 1, comprising means for detecting unauthorized access and requesting additional authorization through a visual selection authentication process.

[0862] "Example 2 of combining an emotion engine"

[0863] (Claim 1)

[0864] A means for generating a visual dataset to be presented in order to collect image selection feedback from users,

[0865] A means of analyzing emotional data from a user's facial expressions using a terminal device,

[0866] A means for aggregating image selection data and sentiment data and analyzing them to improve the training dataset of an automated learning device,

[0867] A means for flexibly adjusting the service according to user requests using the aforementioned training dataset,

[0868] A system that includes this.

[0869] (Claim 2)

[0870] The system according to claim 1, further comprising means for causing the automatic learning device to retrain based on the analysis of the aforementioned feedback data and emotional data.

[0871] (Claim 3)

[0872] The system according to claim 1, further comprising means for detecting unauthorized access based on the user's image recognition behavior and analyzed sentiment data.

[0873] "Application example 2 when combining with an emotional engine"

[0874] (Claim 1)

[0875] A means of providing a function that presents image authentication and collects feedback from users based on their selection of candidates,

[0876] A means of aggregating and analyzing the collected feedback information,

[0877] A means of using the analyzed information to enhance the training dataset of an automated learning function,

[0878] A means of collecting image information and emotional information from information terminals in the work environment and adjusting work conditions,

[0879] A system that includes this.

[0880] (Claim 2)

[0881] The system according to claim 1, further comprising means for retraining the automatic learning function based on the results of the analysis of the aforementioned feedback information.

[0882] (Claim 3)

[0883] The system according to claim 1, further comprising means for adjusting operations based on the worker's intentions in the work environment. [Explanation of symbols]

[0884] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of presenting image authentication in order to collect image selection feedback from users, A means of aggregating and analyzing the collected feedback data, A means of improving the training dataset of an automated learning system using the analyzed data, Means for implementing security measures to distinguish between humans and automated programs, A system that includes this.

2. The system according to claim 1, further comprising means for causing the automatic learning system to retrain based on the analysis of the aforementioned feedback data.

3. The system according to claim 1, further comprising means for detecting unauthorized access based on the user's image authentication behavior.