System

A system using a generative AI model to analyze product images and descriptions addresses inaccuracies in auction and flea market services, enhancing transaction reliability through accurate descriptions and pricing suggestions.

JP2026034200APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137321
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Current auction and flea market services face challenges in verifying the accuracy of product descriptions, leading to misleading information for buyers and inefficiencies in creating effective descriptions and setting prices for sellers.

Method used

A system utilizing a generative AI model to analyze product images and descriptions, providing evaluation results to ensure consistency, detect false statements, and suggest appropriate pricing.

Benefits of technology

Enhances transaction reliability by ensuring accurate product descriptions and pricing suggestions, improving the overall trading experience for both buyers and sellers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026034200000001_ABST
    Figure 2026034200000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for a user to upload an item image and an item description; means for a server to receive and store the item image and the item description; means for the server to analyze the item image and the item description using a generative AI model to generate an assessment result; and means for the server to provide the assessment result to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Current auction and flea market services have the problem that it is difficult for buyers to verify whether the product description matches the actual condition of the product. There is also the risk that they may be misled by false or exaggerated descriptions and make a wrong purchasing decision. Meanwhile, sellers face the challenge of having to spend a lot of time and effort writing effective product descriptions and setting appropriate prices. The purpose of this invention is to solve these problems and realize efficient and reliable transactions. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems by providing a system including the following means: a means for a user to upload product images and product descriptions; a means for a server to receive and store the product images and product descriptions; a means for the server to analyze the product images and product descriptions using a generative AI model and generate evaluation results; and a means for the server to provide the evaluation results to the user. Furthermore, for buyers, the system includes a means for checking the product images and product descriptions for inconsistencies, false statements, and exaggerations. For sellers, the system includes a means for generating appropriate product descriptions based on the product images and product descriptions, and a means for evaluating the value of the product and presenting pricing suggestions. In this way, buyers can make purchasing decisions based on the evaluation results, and sellers can efficiently provide accurate product descriptions and set appropriate prices.

[0006] A "user" is a person who uses the system to upload product images and product descriptions, or who receives evaluation results.

[0007] "Product image" is an image file that shows visual information of the product that the user is selling.

[0008] "Product description" is text information that describes the features, condition, specifications, and other related information of the product that the user is selling.

[0009] A "server" is a computer system that receives product images and product descriptions, stores them, analyzes them, and generates and provides evaluation results.

[0010] A "generative AI model" is an algorithm that uses machine learning and data analysis techniques to analyze images and text.

[0011] "Analysis" refers to the process of analyzing product images and product descriptions and evaluating their consistency and appropriateness.

[0012] The "evaluation results" are feedback information generated through the analysis process that indicates whether the product description matches the product image and whether there are any false representations.

[0013] An "appropriate product description" is text information automatically generated by an AI model that accurately describes the product's features and condition.

[0014] "Pricing suggestions" are suggestions for appropriate starting prices at the time of listing, calculated by an AI model based on the market value of the item.

[0015] "Inconsistencies" refer to discrepancies in information between product images and product descriptions.

[0016] "False representation" refers to information contained in a product description that is factually incorrect or inaccurate.

[0017] "Exaggerated representation" refers to a description of a product that excessively exaggerates the actual condition or value of the product. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] As an embodiment of the present invention, a system that provides convenience to both buyers and sellers in auction and flea market services is shown. The following describes the processing contents of a specific program.

[0040] Overall system configuration

[0041] The system includes a user's device, a server, and a generative AI model. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback.

[0042] Program processing explanation

[0043] Uploading data

[0044] On the device: The seller launches the app on their device, takes a picture of the product, and enters a brief description of the product. After confirming the information entered, they press the "Upload" button to send the data to the server.

[0045] Receiving and storing data

[0046] Server: The server receives the product images and descriptions sent by the user and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[0047] Image and description analysis

[0048] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0049] Generate analysis results

[0050] Server: The results of image and text analysis are integrated to generate feedback for buyers and sellers. For buyers, feedback is provided on the degree of agreement between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, appropriate product descriptions and pricing suggestions based on the product's value are presented.

[0051] Providing analysis results

[0052] Server: The generated analysis results are sent to the user's device and displayed in a specified format. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer, providing reference information to assist in making a purchasing decision.

[0053] Specific examples

[0054] Examples for sellers

[0055] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not tested."

[0056] Device: Press the upload button to send the data to the server.

[0057] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description: "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." A starting price of 3,000 yen is suggested.

[0058] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[0059] Specific examples for buyers

[0060] User: A buyer clicks on a product they're interested in and sees the product image and description: "High-performance smartphone, virtually no scratches."

[0061] Device: Press the analyze button to send the data to the server.

[0062] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[0063] Server: Generates a rating that reads, "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate."

[0064] Server: The evaluation results are provided to the buyer and displayed on the device.

[0065] This system allows buyers and sellers to enjoy a highly reliable trading environment, improving convenience for both parties.

[0066] The processing flow will be explained below.

[0067] Step 1:

[0068] User: The seller launches the app on their device, takes a picture of the product, and enters a simple description of the product. For example, they can enter a description such as "Old camera, operation not confirmed."

[0069] Step 2:

[0070] Terminal: Checks the entered product image and description, converts them into the required format, and displays the "Upload" button. When the user presses the "Upload" button, the data is sent to the server.

[0071] Step 3:

[0072] Server: Receives product images and descriptions and stores them in a database, which also stores related information about the received images and descriptions.

[0073] Step 4:

[0074] Server: The saved product images and descriptions are passed to the generative AI model and analysis begins. Specifically, two processes, image analysis and text analysis, are performed.

[0075] Step 5:

[0076] Server: The image analysis process involves extracting features from product images, such as identifying the camera model, year of manufacture, and external condition.

[0077] Step 6:

[0078] Server: The text analysis process analyzes the product description, extracts keywords, and checks for inconsistencies, false statements, and exaggerations.

[0079] Step 7:

[0080] Server: Integrates the analysis results and generates feedback for buyers and sellers. For buyers, it evaluates the consistency between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates an appropriate product description and suggests a starting price.

[0081] Step 8:

[0082] Server: The generated analysis results are sent to the user's device. Sellers are sent new product descriptions and pricing suggestions. Buyers are sent product image and description evaluation results.

[0083] Step 9:

[0084] Terminal: The received analysis results are displayed in a specified format and notified to the user. The seller is shown the generated product description and pricing suggestions. The buyer is shown any discrepancies and evaluation results, which are provided as reference information to assist in purchasing decisions.

[0085] This process allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information.

[0086] Example 1

[0087] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0088] Conventional auction and flea market services lack the means to verify the authenticity of product images and descriptions uploaded by users, resulting in problems such as fraud and inappropriate language. Furthermore, sellers often lack support for creating appropriate product descriptions and setting prices, resulting in transactions that do not proceed smoothly. The objective of this invention is to provide a system that analyzes product images and descriptions and provides evaluation results to solve these problems.

[0089] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0090] In this invention, the server includes a means for analyzing product images and product descriptions using a generative AI model, performing feature extraction and text analysis, a means for generating evaluation results based on the results of the feature extraction and text analysis, and a means for providing the evaluation results to the user. This ensures the reliability of the product images and descriptions, allowing users to conduct transactions with peace of mind. Furthermore, the server can also provide support for the appropriateness of product descriptions and pricing, which is expected to facilitate smooth transactions.

[0091] "User" means a person or organization that uses the System to upload product images and product descriptions.

[0092] "Product image" refers to an image file that shows visual information about a product and is uploaded to the system by a seller.

[0093] "Product description" is text information that describes the characteristics and condition of a product and is uploaded by the seller to the system.

[0094] "Server" refers to a set of hardware and software systems that receives, stores, and analyzes product images and product descriptions, generates evaluation results, and provides them to users.

[0095] A "generative AI model" is an artificial intelligence model used to analyze product images and product descriptions and perform feature extraction and text analysis.

[0096] "Feature extraction" is the process of using a generative AI model to extract feature information such as product category and condition from product images.

[0097] "Text analysis" is the process of using generative AI models to extract keywords from product descriptions and check for contextual inconsistencies, misrepresentations, and exaggerations.

[0098] The "evaluation results" are feedback information regarding the reliability of a product, appropriate descriptions, and pricing, created based on the results of feature extraction and text analysis.

[0099] The "means for providing to the user" refers to a series of processes and system configuration for transmitting the evaluation results generated by the server to the user's terminal and displaying them.

[0100] This invention is a system that provides convenience to buyers and sellers in auction and flea market services. The system includes a user terminal, a server, and a generative AI model, and the specific program processing content will be explained below.

[0101] Overall system configuration

[0102] The system includes a user's device, a server, and a generative AI model. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback.

[0103] Uploading data

[0104] As a seller, the user launches the terminal app, takes a picture of the product, and enters a brief description of the product. After the device checks the entered product image and description, the user presses the "Upload" button, which sends this data to the server.

[0105] As a specific example of operation, a user takes an image with an old camera, enters a simple description such as "old camera, operation not confirmed," and presses the "upload" button.

[0106] Receiving and storing data

[0107] The server receives product images and descriptions submitted by users, and stores the data in a database, ready for analysis by the generative AI model.

[0108] Image and description analysis

[0109] The server uses a generative AI model to analyze product images and descriptions. For product images, object recognition technology is used to extract features and evaluate the product's category and condition. For product descriptions, text analysis is performed to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0110] As a specific example of how it works, the generative AI model analyzes an image of an old camera, generates a description such as "A vintage camera from the 1940s. The shutter function has not been confirmed, but the appearance is good," and suggests a starting price of 3,000 yen.

[0111] Generate analysis results

[0112] The server combines the results of image and text analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and false statements. For sellers, it presents appropriate product descriptions and pricing suggestions based on the product's value.

[0113] As a specific example of how it works, when a buyer checks a product described as "a high-performance smartphone with almost no scratches" and presses the analysis button, the server generates an evaluation that reads, "There is no inconsistency between the description and the image. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate."

[0114] Providing analysis results

[0115] The server sends the generated analysis results to the user's device and displays them in a specified format. The seller is shown the generated product description and pricing suggestions. The buyer is shown the analysis results of the product image and description, which are provided as reference information to assist in purchasing decisions.

[0116] This system allows buyers and sellers to enjoy a highly reliable trading environment, improving convenience for both parties.

[0117] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0118] Step 1:

[0119] Uploading data

[0120] User: The seller launches the app on their device, takes a picture of the product, and enters a description of the product. The user then presses the "Upload" button to send the data.

[0121] Input: Product image and description

[0122] Output: A data packet of product images and descriptions sent to the server

[0123] Specific behavior:

[0124] The seller launches the terminal app.

[0125] The seller takes a picture of the product.

[0126] The seller enters the product description.

[0127] The seller presses the "Upload" button.

[0128] The device app sends product image and description data to the server.

[0129] Step 2:

[0130] Receiving and storing data

[0131] Server: The server receives the product images and product descriptions sent by the user and stores the received data in a database.

[0132] Input: Product image and description sent from the user's device

[0133] Output: Product images and descriptions stored in a database

[0134] Specific behavior:

[0135] The server receives the data from the terminal.

[0136] The server converts the received data into an appropriate format for analysis.

[0137] The server stores the converted data in a database.

[0138] Step 3:

[0139] Image and description analysis

[0140] Server: The server uses generative AI models to analyze product images and product descriptions. It performs object recognition on product images to assess product category and condition, and performs text analysis on product descriptions to extract keywords and check for inconsistencies, false statements, and exaggerations.

[0141] Input: Product images and descriptions stored in the database

[0142] Output: Analyzed feature information and text analysis results

[0143] Specific behavior:

[0144] The server inputs product images into a generative AI model.

[0145] The generative AI model performs object recognition on the image and extracts the product's features.

[0146] The server inputs the product description into the generative AI model.

[0147] The generative AI model analyzes the text, extracts keywords, and checks for contextual inconsistencies.

[0148] Step 4:

[0149] Generate analysis results

[0150] Server: The server combines the results of image analysis and text analysis to generate feedback for buyers and sellers.

[0151] For buyers, it provides evaluation results on the degree of agreement between the description and the image, as well as any inconsistencies and false statements, while for sellers, it presents suggestions for appropriate product descriptions and pricing.

[0152] Input: Image analysis and text analysis results

[0153] Output: Feedback for buyers and sellers

[0154] Specific behavior:

[0155] The server combines the results of image analysis and text analysis.

[0156] The server generates feedback for the buyer.

[0157] The server generates feedback for the seller.

[0158] Step 5:

[0159] Providing analysis results

[0160] Server: The server sends the generated analysis results to the user's device. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer.

[0161] Input: Feedback for buyers and sellers

[0162] Output: Feedback information displayed on the user's terminal

[0163] Specific behavior:

[0164] The server sends the feedback to the user's terminal.

[0165] The generated product description and pricing suggestions are displayed on the seller's device.

[0166] The analysis results of the product image and description are displayed on the buyer's device.

[0167] (Application example 1)

[0168] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0169] In modern auction and flea market services, many problems arise due to the lack of trust between buyers and sellers. This can make buyers unsure about the safety of their transactions, and sellers find it difficult to properly evaluate and price their products. Furthermore, there is a lack of means to detect inconsistencies and false representations between product images and descriptions, making it difficult to achieve transparent transactions.

[0170] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0171] In this invention, the server includes a means for users to upload product images and product descriptions, a means for the server to receive and store the product images and product descriptions, a means for analyzing the product images and product descriptions using a generative AI model to generate evaluation results and security check results, and a means for users to check the evaluation results and security check results. This makes it possible to check for inconsistencies and false statements between product images and descriptions, generate appropriate product descriptions, present pricing options, and provide security check results.

[0172] "User" refers to sellers who use the system to upload product images and product descriptions, and buyers who check the analysis results.

[0173] "Product images" are photo files of products that users sell.

[0174] "Product description" is text data that describes in sentence format information about the product that the user is selling.

[0175] A "server" is a computer system that receives product images and product descriptions sent by users, analyzes them using a generative AI model, and stores the results.

[0176] "Generative AI model" refers to an algorithm that uses machine learning technology to analyze product images and product descriptions and generate evaluation results.

[0177] "Analysis" is the process by which the generative AI model analyzes product images and product descriptions to evaluate the accuracy of the product's features and descriptions.

[0178] "Evaluation results" are information about the accuracy of the product's characteristics and description obtained through analysis by the generative AI model.

[0179] "Security check results" are information that summarizes the results of detecting fraudulent expressions or potentially misleading statements through analysis of product images and product descriptions.

[0180] "Verifiable means" refers to the interface and functionality that allows users to view evaluation results and security check results.

[0181] "Inconsistencies" refers to inconsistencies or inconsistencies between product images and product descriptions.

[0182] "Misrepresentation" means a product description that is untrue or misleading.

[0183] "Product description generation" is the process in which a generative AI model automatically creates a more specific and attractive description based on the original description.

[0184] "Pricing suggestions" are the estimated value and price suggestions for a product that the generative AI model presents to the user based on the analysis results.

[0185] As an embodiment of the present invention, a system that provides convenience and security to both buyers and sellers in auction and flea market services is shown. The system includes a user terminal, a server, and a generative AI model.

[0186] Overall system configuration

[0187] 1. On the user's device:

[0188] Sellers use a terminal to take product images, enter product descriptions, and upload them.

[0189] The buyer uses a terminal to view the listed items and check the analysis results and security check results.

[0190] 2. Server:

[0191] The server receives the product images and product descriptions sent by the user and stores them in a database.

[0192] A generative AI model is used to analyze product images and descriptions, and generate evaluation and security check results.

[0193] The generated results are sent to the user's terminal and displayed in a predetermined format.

[0194] 3. Generative AI Model:

[0195] For product images, features are extracted using object recognition technology to evaluate the product category and condition.

[0196] For product descriptions, text analysis technology is used to extract keywords and detect contradictions, false statements, and exaggerated expressions.

[0197] Program processing explanation

[0198] Uploading data

[0199] On the device: The seller launches the app on their device, takes a picture of the product, and enters a brief description of the product. After confirming the information entered, they press the "Upload" button to send the data to the server.

[0200] Receiving and storing data

[0201] Server: The server receives the product images and descriptions sent by the user and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[0202] Image and description analysis

[0203] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0204] Generate analysis results

[0205] Server: The results of image and text analysis are integrated to generate feedback for buyers and sellers. For buyers, feedback is provided on the degree of agreement between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, appropriate product descriptions and pricing suggestions based on the product's value are presented.

[0206] Providing analysis results

[0207] Server: The generated analysis results are sent to the user's device and displayed in a specified format. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer, providing reference information to assist in making a purchasing decision.

[0208] Specific examples

[0209] Examples for sellers

[0210] The seller takes a picture of the old camera and enters a simple description such as "Old camera, operation not confirmed." Presses the upload button on the device to send the data to the server. The server analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "Vintage camera from the 1940s, shutter operation not confirmed, but appearance is good." A suggested starting price of 3,000 yen is presented. The server provides the generated product description and suggested starting price to the seller, which are displayed on the device.

[0211] Specific examples for buyers

[0212] A buyer clicks on a product they are interested in and checks the product image and description: "High-performance smartphone, almost no scratches." They then press the analysis button on their device to send the data to the server. The server analyzes the product image and description and checks whether the actual appearance matches the description. The server generates a rating: "There is no discrepancy between the description and the image. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate." The server then provides the rating result to the buyer, which is displayed on the device.

[0213] Prompt Sentence Examples

[0214] Seller prompt:

[0215] Product description: Old camera, operation not confirmed

[0216] Please change this to a more specific and compelling description.

[0217] Buyer prompt:

[0218] Description: High-performance smartphone, almost no scratches

[0219] Please check for discrepancies between images and descriptions and provide feedback on the actual condition.

[0220] These features allow users to conduct transactions with high reliability and transparency.

[0221] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0222] Step 1:

[0223] The user uses a device to take a picture of the product and enter a description. The user starts the application on the device, takes a picture of the product, and enters a description. After checking the entered information, the user presses the "Upload" button.

[0224] Input: Product image file, product description

[0225] Output: Upload request

[0226] Step 2:

[0227] The device sends the product image taken by the user and the product description entered by the user to the server. During this process, the image file and text data are transferred to the server in an appropriate format.

[0228] Input: Upload request (product image file, product description)

[0229] Output: Data sent to the server

[0230] Step 3:

[0231] The server stores the product images and description received from the device in a database, and converts the received data into the required format for analysis.

[0232] Input: Data to be sent (product image, product description)

[0233] Output: Data saved to database

[0234] Step 4:

[0235] The server uses a generative AI model to analyze product images and product descriptions. For product images, object recognition technology (e.g., Tensorflow® or PyTorch) is used to extract features and evaluate the product's category and condition. For product descriptions, text analysis technology (e.g., Hugging Face's T5 model) is used to extract keywords and detect inconsistencies, false statements, and exaggerations.

[0236] Input: Product images and product descriptions stored in the database

[0237] Output: Analysis results (verification results of product category, condition, description)

[0238] Step 5:

[0239] The server combines the results of image and text analysis to generate feedback for buyers and sellers. For buyers, it generates an evaluation result on the degree of match between the image and description, as well as any inconsistencies and false statements. For sellers, it presents appropriate product descriptions and pricing suggestions based on the value of the product.

[0240] Input: Analysis results (product category, condition, and description verification results)

[0241] Output: Feedback (for buyers and sellers)

[0242] Step 6:

[0243] The server sends the generated feedback to the user's terminal. The seller is provided with the generated product description and pricing suggestions, and the buyer is provided with the analysis results of the product image and description.

[0244] Input: Feedback (for buyers, for sellers)

[0245] Output: Feedback data to the user terminal

[0246] Step 7:

[0247] The terminal displays the feedback received from the server. The seller checks the generated product description and pricing suggestions, and the buyer views the analysis results of the product image and description.

[0248] Input: Feedback data sent from the server

[0249] Output: Displayed analysis results and suggestions

[0250] In this way, sellers and buyers can conduct transactions with high reliability and transparency.

[0251] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0252] The present invention is a system that provides convenience to both buyers and sellers in auction and flea market services, and in particular, by combining it with an emotion engine, it further improves the user experience. Below, we will explain the specific program processing content.

[0253] Overall system configuration

[0254] The system includes a user's device, a server, a generative AI model, and an emotion engine. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[0255] Program processing explanation

[0256] Uploading data

[0257] On-device: The seller launches the app on their device, takes a photo of the product, and enters a simple description. For example, they can enter a description such as "old camera, operation not confirmed." In addition, the emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[0258] Receiving and storing data

[0259] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[0260] Image and description analysis

[0261] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0262] Sentiment-based analysis

[0263] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[0264] Generate feedback

[0265] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[0266] Providing analysis results

[0267] Server: The generated analysis results are sent to the user's device and displayed in a specified format. When the generated product description and pricing suggestions are displayed to the seller, they are displayed in a tone based on the analysis results of the emotion engine. The analysis results of the product image and description are displayed to the buyer, providing reference information to support their purchasing decision.

[0268] Specific examples

[0269] Examples for sellers

[0270] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not working." The emotion engine recognizes that the user is relaxed.

[0271] Terminal: Sends data to the server.

[0272] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." Feedback is provided in a relaxed tone based on the emotional information.

[0273] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[0274] Specific examples for buyers

[0275] User: A buyer clicks on a product they're interested in, sees the product image and description, "High-performance smartphone, virtually flawless." The emotion engine recognizes the user's doubts.

[0276] Device: Press the analyze button to send the data to the server.

[0277] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[0278] Server: Generates a rating like "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate." Based on sentiment, additional information is provided to resolve any doubts.

[0279] Server: The evaluation results are provided to the buyer and displayed on the device.

[0280] This system allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information and receiving feedback that takes into account their emotions.

[0281] The processing flow will be explained below.

[0282] Step 1:

[0283] User: The seller launches the app on their device, takes a picture of the product, and writes a brief description of the product. The emotion engine then analyzes the seller's emotional state (e.g., nervous, relaxed, anxious, etc.).

[0284] Step 2:

[0285] Terminal: Checks the entered product image, description, and emotion information, converts them into the required format, and displays the "Upload" button. When the user presses the "Upload" button, the data is sent to the server.

[0286] Step 3:

[0287] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[0288] Step 4:

[0289] Server: The saved product images and descriptions are passed to the generative AI model and analysis begins. Specifically, two processes, image analysis and text analysis, are performed.

[0290] Step 5:

[0291] Server: The image analysis process involves extracting features from product images, such as identifying the camera model, year of manufacture, and external condition.

[0292] Step 6:

[0293] Server: The text analysis process analyzes the product description, extracts keywords, and checks for inconsistencies, misrepresentations, and exaggerations.

[0294] Step 7:

[0295] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[0296] Step 8:

[0297] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[0298] Step 9:

[0299] Server: The generated analysis results are sent to the user's device. Sellers are sent new product descriptions and pricing suggestions. Buyers are sent product image and description evaluation results.

[0300] Step 10:

[0301] Terminal: The received analysis results are displayed in a specified format and notified to the user. For sellers, the generated product description and pricing suggestions are displayed in a tone based on the emotion engine's analysis results. For buyers, discrepancies and evaluation results are displayed, providing reference information to assist in purchasing decisions.

[0302] This process allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information and emotionally sensitive feedback.

[0303] Example 2

[0304] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0305] Traditional auction and flea market services have faced challenges in terms of the reliability of product descriptions and improving the experience for buyers and sellers. In particular, there were insufficient means to check the consistency between product descriptions and images, and whether or not there were false or exaggerated statements, making it difficult for buyers to make decisions based on these. Furthermore, the impersonal feedback provided meant that services were not provided that took user feelings into consideration.

[0306] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0307] In this invention, the server includes a means for a user to upload product images and product descriptions, a means for the server to receive and store the product images and product descriptions, a means for analyzing the product images and product descriptions using a generative AI model to generate evaluation results, and a means for an emotion engine to analyze the user's emotional information and generate appropriate feedback by adjusting the tone and approach based on the emotional information. This enables the provision of highly reliable product evaluations as well as feedback that takes the user's emotions into consideration.

[0308] "User" refers to an individual or corporation that intends to sell or purchase items using the auction and flea market services.

[0309] "Terminal" refers to the electronic device (smartphone, PC, tablet, etc.) used by the user to input product images and product descriptions and send them to the system.

[0310] "Server" refers to a central management system that receives, stores, and analyzes data sent from user devices, generates evaluation results, and provides them to users.

[0311] "Generative AI model" refers to the artificial intelligence model used by the server to analyze product images and product descriptions and generate evaluation results.

[0312] "Product image" refers to a photograph or image data of the product that a user is listing for sale.

[0313] "Product description" refers to text data that explains the condition, features, price, etc. of the product that a user is selling.

[0314] An "emotion engine" is a system that analyzes emotional information from a user's facial expressions and voice, and provides appropriate feedback based on the results.

[0315] "Feedback" refers to comments, advice, notifications, etc. provided to users based on the evaluation results generated by the server.

[0316] The present invention is a system that provides convenience to both buyers and sellers in auction and flea market services, and in particular improves the user experience by combining an emotion engine. Specific embodiments are described below.

[0317] The system includes a user device, a server, a generative AI model, and an emotion engine. The user device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[0318] First, the seller launches the device app, takes a picture of the product, and enters a simple description. For example, the seller might enter a description such as "old camera, operation not confirmed." The device temporarily saves the captured image and the entered description in local storage. Next, the emotion engine uses the device's camera and microphone to analyze the seller's emotions.

[0319] Next, the device sends the product image, product description, and emotion information to the server. The server receives this data and stores it in a database. The product image is stored as image data, the description as text data, and the emotion information as structured data.

[0320] The server uses a generative AI model to analyze product images. Specifically, it uses object recognition technology to identify the product category and evaluate the product's condition. The generative AI model also performs text analysis of the product description to extract important keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0321] The emotion engine then generates feedback for sellers and buyers based on the emotional information analyzed. For sellers, it presents pricing suggestions based on appropriate product descriptions and the value of the product. For example, the server generates a description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good," and provides it to the seller. For buyers, it provides feedback on the degree of match between the product image and description, as well as any inconsistencies and false statements.

[0322] Furthermore, the tone and content of the feedback is adjusted based on the results of the emotion engine. For example, if the buyer is skeptical, detailed feedback such as "The description and images are consistent. The phone looks good, but there are some small scratches in the photos. The description is generally accurate."

[0323] Finally, the analysis results and feedback generated by the server are sent to the user's device and displayed in an appropriate format. The seller can review the generated product description and pricing suggestions and make any necessary corrections. The buyer can make a purchase decision based on the analysis results provided and with reliable information.

[0324] For example, if a seller takes a picture of an old camera and enters the description "Old camera, not working," the emotion engine recognizes that the seller is relaxed. The server analyzes the image and text, and the generative AI model generates an appropriate description, such as "Vintage camera from the 1940s, shutter not working, but looks good," and provides feedback in a relaxed tone.

[0325] If a buyer clicks on a smartphone product and sees the product image and the description "High-performance smartphone, almost no scratches," the emotion engine recognizes the buyer's doubts. The server analyzes the image and text and generates a rating result: "The description and the image are consistent. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate," providing detailed and reliable feedback.

[0326] This system allows both users and sellers to obtain reliable information, sellers to efficiently generate appropriate product descriptions, and buyers to make purchasing decisions by receiving feedback that takes into account their emotions.

[0327] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0328] Step 1:

[0329] User: The seller launches the app on their device and takes a picture of the product. For example, they take a picture of an "old camera from the 1940s."

[0330] Input: The captured image.

[0331] Specific operation: The device app uses the camera function to temporarily save the captured image to local storage.

[0332] Output: Product images saved to local storage.

[0333] Step 2:

[0334] User: The seller enters the product description into the terminal. For example, they might enter "old camera, not working."

[0335] Input: The entered product description.

[0336] Specific operation: The terminal app temporarily saves the entered text data in local storage.

[0337] Output: Product description saved to local storage.

[0338] Step 3:

[0339] On the device: The emotion engine uses the device's camera and microphone to analyze the seller's facial expressions and voice.

[0340] Input: Seller's facial expression and voice data.

[0341] How it works: The sentiment engine uses machine learning algorithms to analyze the emotional state of sellers in real time.

[0342] Output: Seller emotional state data.

[0343] Step 4:

[0344] Device: Sends product images, product descriptions, and emotion information to the server.

[0345] Input: Product image, product description, sentiment information.

[0346] Specific operation: The device uses the network and sends the data as a data packet to the server.

[0347] Output: The server receives the product image, description, and emotion information as a data packet.

[0348] Step 5:

[0349] Server: Stores the received product images, product descriptions, and emotion information in a database.

[0350] Input: Received data (product image, product description, emotional information).

[0351] Specific operation: The server converts the data into a data format and stores it in a database. Product images are saved as image data, descriptions as text data, and emotional information as structured data.

[0352] Output: Product images, product descriptions, and sentiment information stored in a database.

[0353] Step 6:

[0354] Server: Analyzes product images using a generative AI model and identifies the product category using object recognition technology.

[0355] Input: Product images stored in the database.

[0356] How it works: The generative AI model uses object recognition algorithms to extract features in an image and identify the product category.

[0357] Output: Identified product category and product condition.

[0358] Step 7:

[0359] Server: Uses a generative AI model to analyze the text of product descriptions, extracting important keywords and checking for contextual inconsistencies, misrepresentations, and exaggerations.

[0360] Input: Product description stored in the database.

[0361] How it works: The generative AI model uses natural language processing techniques to analyze text, extract keywords, analyze context, and check for misrepresentations.

[0362] Output: Keyword extraction results, contextual inconsistencies, misrepresentations, and exaggerations.

[0363] Step 8:

[0364] Server: Generates feedback for sellers and buyers based on the emotional information analyzed by the emotion engine.

[0365] Input: Emotion information, product image analysis results, product description analysis results.

[0366] Specific operation: The server takes into account emotional information and generates appropriate product description and pricing suggestions for the seller, and generates feedback for the buyer regarding the degree of match between the description and the image, any inconsistencies, and whether there are any misrepresentations.

[0367] Output: Feedback for seller (product description, suggested pricing), feedback for buyer (match between image and description, inconsistencies, misrepresentations).

[0368] Step 9:

[0369] Server: Sends the generated feedback to the user's device.

[0370] Input: The generated feedback.

[0371] Specific operation: The server sends feedback data to the terminal via the network.

[0372] Output: Feedback data received by the user terminal.

[0373] Step 10:

[0374] Terminal: Displays the received feedback data in an appropriate format.

[0375] Input: The received feedback data.

[0376] Specific operation: The device analyzes the feedback data and displays it on the user interface. The seller is shown the generated product description and pricing suggestions, and the buyer is shown the analysis results.

[0377] Output: Feedback displayed on the device.

[0378] (Application example 2)

[0379] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0380] In conventional auction and flea market services, if the product descriptions and images provided by sellers contain inconsistencies or falsehoods, buyers' trust is often damaged. It is also difficult for sellers to set appropriate prices and product descriptions, and there is a lack of advice and feedback to stimulate purchasing motivation. Another problem is the lack of a system that provides feedback based on the user's emotional state. Given this background, there is a need for a system that can provide accurate and reliable information while taking user emotions into consideration.

[0381] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0382] In this invention, the server includes: means for a user to upload product images and product descriptions; means for the server to receive and store the product images and product descriptions; means for the server to analyze the product images and product descriptions using a generative AI model and generate evaluation results; means for the server to provide the evaluation results to the user; means for an emotion engine to analyze the user's emotions; and means for generating and providing feedback to the user based on the analysis results and emotion information. This makes it possible to provide highly reliable feedback that takes the user's emotions into consideration.

[0383] "User emotion" refers to the user's psychological state and emotions as read and analyzed by the emotion engine.

[0384] "Emotion engine" is a general term for software and algorithms used to analyze emotions based on data such as a user's facial expressions and voice.

[0385] A "generative AI model" is an artificial intelligence model that analyzes product images and descriptions and generates the necessary feedback.

[0386] "Product images" refer to photographs and image data of the products being offered for sale.

[0387] A "product description" is a text description of the characteristics and condition of the product being offered for sale.

[0388] "Analysis results" refer to the evaluations and feedback generated by generative AI models and emotion engines.

[0389] "Feedback" refers to advice and evaluation information provided to the user based on the analysis results and emotional information.

[0390] A "server" is a centralized processing unit for receiving, storing, analyzing, and providing feedback to the user of data.

[0391] "Upload" refers to the act of a user sending data from their own terminal to a server.

[0392] "Receiving" means that the server takes in data such as product images and product descriptions sent from the user's terminal.

[0393] "Storing" means storing the received data in a storage device such as a database.

[0394] "Analysis" refers to using generative AI models and emotion engines to process product images and descriptions to generate ratings and feedback.

[0395] "Inconsistencies" refer to discrepancies or inconsistencies between product images and descriptions.

[0396] "False representation" refers to information contained in a product description that is incorrect and different from the facts.

[0397] "Exaggeration" refers to descriptions that exaggerate the actual characteristics of a product.

[0398] The present invention is a system that includes a means for a user to upload product images and product descriptions, a means for a server to receive and store the product images and product descriptions, a means for the server to analyze the product images and product descriptions using a generative AI model and generate an evaluation result, a means for the server to provide the evaluation result to the user, a means for an emotion engine to analyze the user's emotions, and a means for generating feedback based on the analysis result and emotion information and providing it to the user.

[0399] Overall system configuration

[0400] The system includes a user's device, a server, a generative AI model, and an emotion engine. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[0401] Program processing explanation

[0402] Uploading data

[0403] On-device: The seller launches the app on their device, takes a photo of the product, and enters a simple description. For example, they can enter a description such as "old camera, operation not confirmed." In addition, the emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[0404] Receiving and storing data

[0405] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[0406] Image and description analysis

[0407] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0408] Sentiment-based analysis

[0409] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[0410] Generate feedback

[0411] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[0412] Hardware and software used

[0413] Hardware: The smartphone used by the user

[0414] Software: Python, OpenCV, Transformers library, EmotionRecognition library

[0415] Specific examples

[0416] Examples for sellers

[0417] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not working." The emotion engine recognizes that the user is relaxed.

[0418] Terminal: Sends data to the server.

[0419] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." Feedback is provided in a relaxed tone based on the emotional information.

[0420] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[0421] Specific examples for buyers

[0422] User: A buyer clicks on a product they're interested in, sees the product image and description, "High-performance smartphone, virtually flawless." The emotion engine recognizes the user's doubts.

[0423] Device: Press the analyze button to send the data to the server.

[0424] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[0425] Server: Generates a rating like "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate." Based on sentiment, additional information is provided to resolve any doubts.

[0426] Server: The evaluation results are provided to the buyer and displayed on the device.

[0427] Prompt Sentence Examples

[0428] Description of the invention:

[0429] I would like to develop a system to improve the user experience of auction and flea market services. This system combines an emotion engine and a generative AI model to analyze product images and descriptions and provide appropriate feedback to both buyers and sellers. It also includes a function to check for inconsistencies and false statements in product descriptions and suggest optimal pricing to sellers.

[0430] Prerequisites:

[0431] Sellers upload product images and descriptions.

[0432] Analyzing user emotions with an emotion engine

[0433] Analyze and generate product descriptions using a generative AI model

[0434] Give feedback in an emotional tone

[0435] Expected output:

[0436] Relaxed tone of voice and optimal product description and pricing suggestions for sellers

[0437] Feedback to buyers regarding the consistency of product descriptions and images, as well as any inconsistencies

[0438] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0439] Step 1:

[0440] Uploading data

[0441] Device: The seller uses a smartphone to take a picture of the product and enter a brief description. For example, they might enter "old camera, not yet operational." The emotion engine then activates and analyzes the seller's facial expressions and voice to obtain emotional information.

[0442] Input: Product images, product descriptions, seller's facial expressions and voice data

[0443] Output: Product images, product descriptions, emotional information

[0444] Step 2:

[0445] Receiving and storing data

[0446] Server: Receives product images, product descriptions, and emotion information sent from the device. The received data is stored in a database using transaction processing to prepare for analysis.

[0447] Input: Product image, product description, emotional information

[0448] Output: Product images, product descriptions, and emotional information stored in a database

[0449] Step 3:

[0450] Image and description analysis

[0451] Server: Using a generative AI model, it analyzes product images and descriptions. For product images, it uses object recognition technology to evaluate the product category and condition, and for descriptions, it performs text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0452] Input: Product image, product description

[0453] Output: Analysis results of product images, analysis results of product descriptions

[0454] Step 4:

[0455] Sentiment-based analysis

[0456] Server: Based on the emotional information analyzed by the emotion engine, the system processes the user's emotional state. For example, if the user is nervous, the system generates feedback using a gentle tone.

[0457] Input: Emotion information

[0458] Output: Tone information for emotion-based feedback

[0459] Step 5:

[0460] Generate feedback

[0461] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and false statements. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[0462] Input: Analysis results of product images, analysis results of product descriptions, tone information of feedback based on emotions

[0463] Output: Buyer and seller feedback data

[0464] Step 6:

[0465] Providing feedback

[0466] Server: The generated feedback and rating results are sent to the user's device and provided to the seller and buyer. The seller is shown the generated product description and suggested starting price, and the buyer is shown the analysis results of the product image and description.

[0467] Input: Buyer and seller feedback data

[0468] Output: Feedback and evaluation results displayed on the user's device

[0469] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0470] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0471] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0472] [Second embodiment]

[0473] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0474] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0475] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0476] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0477] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0478] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0479] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0480] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0481] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0482] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0483] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0484] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0485] As an embodiment of the present invention, a system that provides convenience to both buyers and sellers in auction and flea market services is shown. The following describes the processing contents of a specific program.

[0486] Overall system configuration

[0487] The system includes a user's device, a server, and a generative AI model. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback.

[0488] Program processing explanation

[0489] Uploading data

[0490] On the device: The seller launches the app on their device, takes a picture of the product, and enters a brief description of the product. After confirming the information entered, they press the "Upload" button to send the data to the server.

[0491] Receiving and storing data

[0492] Server: The server receives the product images and descriptions sent by the user and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[0493] Image and description analysis

[0494] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0495] Generate analysis results

[0496] Server: The results of image and text analysis are integrated to generate feedback for buyers and sellers. For buyers, feedback is provided on the degree of agreement between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, appropriate product descriptions and pricing suggestions based on the product's value are presented.

[0497] Providing analysis results

[0498] Server: The generated analysis results are sent to the user's device and displayed in a specified format. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer, providing reference information to assist in making a purchasing decision.

[0499] Specific examples

[0500] Examples for sellers

[0501] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not tested."

[0502] Device: Press the upload button to send the data to the server.

[0503] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description: "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." A starting price of 3,000 yen is suggested.

[0504] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[0505] Specific examples for buyers

[0506] User: A buyer clicks on a product they're interested in and sees the product image and description: "High-performance smartphone, virtually no scratches."

[0507] Device: Press the analyze button to send the data to the server.

[0508] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[0509] Server: Generates a rating that reads, "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate."

[0510] Server: The evaluation results are provided to the buyer and displayed on the device.

[0511] This system allows buyers and sellers to enjoy a highly reliable trading environment, improving convenience for both parties.

[0512] The processing flow will be explained below.

[0513] Step 1:

[0514] User: The seller launches the app on their device, takes a picture of the product, and enters a simple description of the product. For example, they can enter a description such as "Old camera, operation not confirmed."

[0515] Step 2:

[0516] Terminal: Checks the entered product image and description, converts them into the required format, and displays the "Upload" button. When the user presses the "Upload" button, the data is sent to the server.

[0517] Step 3:

[0518] Server: Receives product images and descriptions and stores them in a database, which also stores related information about the received images and descriptions.

[0519] Step 4:

[0520] Server: The saved product images and descriptions are passed to the generative AI model and analysis begins. Specifically, two processes, image analysis and text analysis, are performed.

[0521] Step 5:

[0522] Server: The image analysis process involves extracting features from product images, such as identifying the camera model, year of manufacture, and external condition.

[0523] Step 6:

[0524] Server: The text analysis process analyzes the product description, extracts keywords, and checks for inconsistencies, false statements, and exaggerations.

[0525] Step 7:

[0526] Server: Integrates the analysis results and generates feedback for buyers and sellers. For buyers, it evaluates the consistency between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates an appropriate product description and suggests a starting price.

[0527] Step 8:

[0528] Server: The generated analysis results are sent to the user's device. Sellers are sent new product descriptions and pricing suggestions. Buyers are sent product image and description evaluation results.

[0529] Step 9:

[0530] Terminal: The received analysis results are displayed in a specified format and notified to the user. The seller is shown the generated product description and pricing suggestions. The buyer is shown any discrepancies and evaluation results, which are provided as reference information to assist in purchasing decisions.

[0531] This process allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information.

[0532] Example 1

[0533] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0534] Conventional auction and flea market services lack the means to verify the authenticity of product images and descriptions uploaded by users, resulting in problems such as fraud and inappropriate language. Furthermore, sellers often lack support for creating appropriate product descriptions and setting prices, resulting in transactions that do not proceed smoothly. The objective of this invention is to provide a system that analyzes product images and descriptions and provides evaluation results to solve these problems.

[0535] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0536] In this invention, the server includes a means for analyzing product images and product descriptions using a generative AI model, performing feature extraction and text analysis, a means for generating evaluation results based on the results of the feature extraction and text analysis, and a means for providing the evaluation results to the user. This ensures the reliability of the product images and descriptions, allowing users to conduct transactions with peace of mind. Furthermore, the server can also provide support for the appropriateness of product descriptions and pricing, which is expected to facilitate smooth transactions.

[0537] "User" means a person or organization that uses the System to upload product images and product descriptions.

[0538] "Product image" refers to an image file that shows visual information about a product and is uploaded to the system by a seller.

[0539] "Product description" is text information that describes the characteristics and condition of a product and is uploaded by the seller to the system.

[0540] "Server" refers to a set of hardware and software systems that receives, stores, and analyzes product images and product descriptions, generates evaluation results, and provides them to users.

[0541] A "generative AI model" is an artificial intelligence model used to analyze product images and product descriptions and perform feature extraction and text analysis.

[0542] "Feature extraction" is the process of using a generative AI model to extract feature information such as product category and condition from product images.

[0543] "Text analysis" is the process of using generative AI models to extract keywords from product descriptions and check for contextual inconsistencies, misrepresentations, and exaggerations.

[0544] The "evaluation results" are feedback information regarding the reliability of a product, appropriate descriptions, and pricing, created based on the results of feature extraction and text analysis.

[0545] The "means for providing to the user" refers to a series of processes and system configuration for transmitting the evaluation results generated by the server to the user's terminal and displaying them.

[0546] This invention is a system that provides convenience to buyers and sellers in auction and flea market services. The system includes a user terminal, a server, and a generative AI model, and the specific program processing content will be explained below.

[0547] Overall system configuration

[0548] The system includes a user's device, a server, and a generative AI model. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback.

[0549] Uploading data

[0550] As a seller, the user launches the terminal app, takes a picture of the product, and enters a brief description of the product. After the device checks the entered product image and description, the user presses the "Upload" button, which sends this data to the server.

[0551] As a specific example of operation, a user takes an image with an old camera, enters a simple description such as "old camera, operation not confirmed," and presses the "upload" button.

[0552] Receiving and storing data

[0553] The server receives product images and descriptions submitted by users, and stores the data in a database, ready for analysis by the generative AI model.

[0554] Image and description analysis

[0555] The server uses a generative AI model to analyze product images and descriptions. For product images, object recognition technology is used to extract features and evaluate the product's category and condition. For product descriptions, text analysis is performed to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0556] As a specific example of how it works, the generative AI model analyzes an image of an old camera, generates a description such as "A vintage camera from the 1940s. The shutter function has not been confirmed, but the appearance is good," and suggests a starting price of 3,000 yen.

[0557] Generate analysis results

[0558] The server combines the results of image and text analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and false statements. For sellers, it presents appropriate product descriptions and pricing suggestions based on the product's value.

[0559] As a specific example of how it works, when a buyer checks a product described as "a high-performance smartphone with almost no scratches" and presses the analysis button, the server generates an evaluation that reads, "There is no inconsistency between the description and the image. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate."

[0560] Providing analysis results

[0561] The server sends the generated analysis results to the user's device and displays them in a specified format. The seller is shown the generated product description and pricing suggestions. The buyer is shown the analysis results of the product image and description, which are provided as reference information to assist in purchasing decisions.

[0562] This system allows buyers and sellers to enjoy a highly reliable trading environment, improving convenience for both parties.

[0563] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0564] Step 1:

[0565] Uploading data

[0566] User: The seller launches the app on their device, takes a picture of the product, and enters a description of the product. The user then presses the "Upload" button to send the data.

[0567] Input: Product image and description

[0568] Output: A data packet of product images and descriptions sent to the server

[0569] Specific behavior:

[0570] The seller launches the terminal app.

[0571] The seller takes a picture of the product.

[0572] The seller enters the product description.

[0573] The seller presses the "Upload" button.

[0574] The device app sends product image and description data to the server.

[0575] Step 2:

[0576] Receiving and storing data

[0577] Server: The server receives the product images and product descriptions sent by the user and stores the received data in a database.

[0578] Input: Product image and description sent from the user's device

[0579] Output: Product images and descriptions stored in a database

[0580] Specific behavior:

[0581] The server receives the data from the terminal.

[0582] The server converts the received data into an appropriate format for analysis.

[0583] The server stores the converted data in a database.

[0584] Step 3:

[0585] Image and description analysis

[0586] Server: The server uses generative AI models to analyze product images and product descriptions. It performs object recognition on product images to assess product category and condition, and performs text analysis on product descriptions to extract keywords and check for inconsistencies, false statements, and exaggerations.

[0587] Input: Product images and descriptions stored in the database

[0588] Output: Analyzed feature information and text analysis results

[0589] Specific behavior:

[0590] The server inputs product images into a generative AI model.

[0591] The generative AI model performs object recognition on the image and extracts the product's features.

[0592] The server inputs the product description into the generative AI model.

[0593] The generative AI model analyzes the text, extracts keywords, and checks for contextual inconsistencies.

[0594] Step 4:

[0595] Generate analysis results

[0596] Server: The server combines the results of image analysis and text analysis to generate feedback for buyers and sellers.

[0597] For buyers, it provides evaluation results on the degree of agreement between the description and the image, as well as any inconsistencies and false statements, while for sellers, it presents suggestions for appropriate product descriptions and pricing.

[0598] Input: Image analysis and text analysis results

[0599] Output: Feedback for buyers and sellers

[0600] Specific behavior:

[0601] The server combines the results of image analysis and text analysis.

[0602] The server generates feedback for the buyer.

[0603] The server generates feedback for the seller.

[0604] Step 5:

[0605] Providing analysis results

[0606] Server: The server sends the generated analysis results to the user's device. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer.

[0607] Input: Feedback for buyers and sellers

[0608] Output: Feedback information displayed on the user's terminal

[0609] Specific behavior:

[0610] The server sends the feedback to the user's terminal.

[0611] The generated product description and pricing suggestions are displayed on the seller's device.

[0612] The analysis results of the product image and description are displayed on the buyer's device.

[0613] (Application example 1)

[0614] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0615] In modern auction and flea market services, many problems arise due to the lack of trust between buyers and sellers. This can make buyers unsure about the safety of their transactions, and sellers find it difficult to properly evaluate and price their products. Furthermore, there is a lack of means to detect inconsistencies and false representations between product images and descriptions, making it difficult to achieve transparent transactions.

[0616] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0617] In this invention, the server includes a means for users to upload product images and product descriptions, a means for the server to receive and store the product images and product descriptions, a means for analyzing the product images and product descriptions using a generative AI model to generate evaluation results and security check results, and a means for users to check the evaluation results and security check results. This makes it possible to check for inconsistencies and false statements between product images and descriptions, generate appropriate product descriptions, present pricing options, and provide security check results.

[0618] "User" refers to sellers who use the system to upload product images and product descriptions, and buyers who check the analysis results.

[0619] "Product images" are photo files of products that users sell.

[0620] "Product description" is text data that describes in sentence format information about the product that the user is selling.

[0621] A "server" is a computer system that receives product images and product descriptions sent by users, analyzes them using a generative AI model, and stores the results.

[0622] "Generative AI model" refers to an algorithm that uses machine learning technology to analyze product images and product descriptions and generate evaluation results.

[0623] "Analysis" is the process by which the generative AI model analyzes product images and product descriptions to evaluate the accuracy of the product's features and descriptions.

[0624] "Evaluation results" are information about the accuracy of the product's characteristics and description obtained through analysis by the generative AI model.

[0625] "Security check results" are information that summarizes the results of detecting fraudulent expressions or potentially misleading statements through analysis of product images and product descriptions.

[0626] "Verifiable means" refers to the interface and functionality that allows users to view evaluation results and security check results.

[0627] "Inconsistencies" refers to inconsistencies or inconsistencies between product images and product descriptions.

[0628] "Misrepresentation" means a product description that is untrue or misleading.

[0629] "Product description generation" is the process in which a generative AI model automatically creates a more specific and attractive description based on the original description.

[0630] "Pricing suggestions" are the estimated value and price suggestions for a product that the generative AI model presents to the user based on the analysis results.

[0631] As an embodiment of the present invention, a system that provides convenience and security to both buyers and sellers in auction and flea market services is shown. The system includes a user terminal, a server, and a generative AI model.

[0632] Overall system configuration

[0633] 1. On the user's device:

[0634] Sellers use a terminal to take product images, enter product descriptions, and upload them.

[0635] The buyer uses a terminal to view the listed items and check the analysis results and security check results.

[0636] 2. Server:

[0637] The server receives the product images and product descriptions sent by the user and stores them in a database.

[0638] A generative AI model is used to analyze product images and descriptions, and generate evaluation and security check results.

[0639] The generated results are sent to the user's terminal and displayed in a predetermined format.

[0640] 3. Generative AI Model:

[0641] For product images, features are extracted using object recognition technology to evaluate the product category and condition.

[0642] For product descriptions, text analysis technology is used to extract keywords and detect contradictions, false statements, and exaggerated expressions.

[0643] Program processing explanation

[0644] Uploading data

[0645] On the device: The seller launches the app on their device, takes a picture of the product, and enters a brief description of the product. After confirming the information entered, they press the "Upload" button to send the data to the server.

[0646] Receiving and storing data

[0647] Server: The server receives the product images and descriptions sent by the user and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[0648] Image and description analysis

[0649] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0650] Generate analysis results

[0651] Server: The results of image and text analysis are integrated to generate feedback for buyers and sellers. For buyers, feedback is provided on the degree of agreement between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, appropriate product descriptions and pricing suggestions based on the product's value are presented.

[0652] Providing analysis results

[0653] Server: The generated analysis results are sent to the user's device and displayed in a specified format. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer, providing reference information to assist in making a purchasing decision.

[0654] Specific examples

[0655] Examples for sellers

[0656] The seller takes a picture of the old camera and enters a simple description such as "Old camera, operation not confirmed." Presses the upload button on the device to send the data to the server. The server analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "Vintage camera from the 1940s, shutter operation not confirmed, but appearance is good." A suggested starting price of 3,000 yen is presented. The server provides the generated product description and suggested starting price to the seller, which are displayed on the device.

[0657] Specific examples for buyers

[0658] A buyer clicks on a product they are interested in and checks the product image and description: "High-performance smartphone, almost no scratches." They then press the analysis button on their device to send the data to the server. The server analyzes the product image and description and checks whether the actual appearance matches the description. The server generates a rating: "There is no discrepancy between the description and the image. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate." The server then provides the rating result to the buyer, which is displayed on the device.

[0659] Prompt Sentence Examples

[0660] Seller prompt:

[0661] Product description: Old camera, operation not confirmed

[0662] Please change this to a more specific and compelling description.

[0663] Buyer prompt:

[0664] Description: High-performance smartphone, almost no scratches

[0665] Please check for discrepancies between images and descriptions and provide feedback on the actual condition.

[0666] These features allow users to conduct transactions with high reliability and transparency.

[0667] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0668] Step 1:

[0669] The user uses a device to take a picture of the product and enter a description. The user starts the application on the device, takes a picture of the product, and enters a description. After checking the entered information, the user presses the "Upload" button.

[0670] Input: Product image file, product description

[0671] Output: Upload request

[0672] Step 2:

[0673] The device sends the product image taken by the user and the product description entered by the user to the server. During this process, the image file and text data are transferred to the server in an appropriate format.

[0674] Input: Upload request (product image file, product description)

[0675] Output: Data sent to the server

[0676] Step 3:

[0677] The server stores the product images and description received from the device in a database, and converts the received data into the required format for analysis.

[0678] Input: Data to be sent (product image, product description)

[0679] Output: Data saved to database

[0680] Step 4:

[0681] The server uses a generative AI model to analyze product images and product descriptions. For product images, object recognition technology (e.g., TensorFlow or PyTorch) is used to extract features and evaluate the product's category and condition. For product descriptions, text analysis technology (e.g., Hugging Face's T5 model) is used to extract keywords and detect contradictions, false statements, and exaggerations.

[0682] Input: Product images and product descriptions stored in the database

[0683] Output: Analysis results (verification results of product category, condition, description)

[0684] Step 5:

[0685] The server combines the results of image and text analysis to generate feedback for buyers and sellers. For buyers, it generates an evaluation result on the degree of match between the image and description, as well as any inconsistencies and false statements. For sellers, it presents appropriate product descriptions and pricing suggestions based on the value of the product.

[0686] Input: Analysis results (product category, condition, and description verification results)

[0687] Output: Feedback (for buyers and sellers)

[0688] Step 6:

[0689] The server sends the generated feedback to the user's terminal. The seller is provided with the generated product description and pricing suggestions, and the buyer is provided with the analysis results of the product image and description.

[0690] Input: Feedback (for buyers, for sellers)

[0691] Output: Feedback data to the user terminal

[0692] Step 7:

[0693] The terminal displays the feedback received from the server. The seller checks the generated product description and pricing suggestions, and the buyer views the analysis results of the product image and description.

[0694] Input: Feedback data sent from the server

[0695] Output: Displayed analysis results and suggestions

[0696] In this way, sellers and buyers can conduct transactions with high reliability and transparency.

[0697] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0698] The present invention is a system that provides convenience to both buyers and sellers in auction and flea market services, and in particular, by combining it with an emotion engine, it further improves the user experience. Below, we will explain the specific program processing content.

[0699] Overall system configuration

[0700] The system includes a user's device, a server, a generative AI model, and an emotion engine. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[0701] Program processing explanation

[0702] Uploading data

[0703] On-device: The seller launches the app on their device, takes a photo of the product, and enters a simple description. For example, they can enter a description such as "old camera, operation not confirmed." In addition, the emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[0704] Receiving and storing data

[0705] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[0706] Image and description analysis

[0707] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0708] Sentiment-based analysis

[0709] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[0710] Generate feedback

[0711] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[0712] Providing analysis results

[0713] Server: The generated analysis results are sent to the user's device and displayed in a specified format. When the generated product description and pricing suggestions are displayed to the seller, they are displayed in a tone based on the analysis results of the emotion engine. The analysis results of the product image and description are displayed to the buyer, providing reference information to support their purchasing decision.

[0714] Specific examples

[0715] Examples for sellers

[0716] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not working." The emotion engine recognizes that the user is relaxed.

[0717] Terminal: Sends data to the server.

[0718] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." Feedback is provided in a relaxed tone based on the emotional information.

[0719] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[0720] Specific examples for buyers

[0721] User: A buyer clicks on a product they're interested in, sees the product image and description, "High-performance smartphone, virtually flawless." The emotion engine recognizes the user's doubts.

[0722] Device: Press the analyze button to send the data to the server.

[0723] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[0724] Server: Generates a rating like "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate." Based on sentiment, additional information is provided to resolve any doubts.

[0725] Server: The evaluation results are provided to the buyer and displayed on the device.

[0726] This system allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information and receiving feedback that takes into account their emotions.

[0727] The processing flow will be explained below.

[0728] Step 1:

[0729] User: The seller launches the app on their device, takes a picture of the product, and writes a brief description of the product. The emotion engine then analyzes the seller's emotional state (e.g., nervous, relaxed, anxious, etc.).

[0730] Step 2:

[0731] Terminal: Checks the entered product image, description, and emotion information, converts them into the required format, and displays the "Upload" button. When the user presses the "Upload" button, the data is sent to the server.

[0732] Step 3:

[0733] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[0734] Step 4:

[0735] Server: The saved product images and descriptions are passed to the generative AI model and analysis begins. Specifically, two processes, image analysis and text analysis, are performed.

[0736] Step 5:

[0737] Server: The image analysis process involves extracting features from product images, such as identifying the camera model, year of manufacture, and external condition.

[0738] Step 6:

[0739] Server: The text analysis process analyzes the product description, extracts keywords, and checks for inconsistencies, misrepresentations, and exaggerations.

[0740] Step 7:

[0741] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[0742] Step 8:

[0743] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[0744] Step 9:

[0745] Server: The generated analysis results are sent to the user's device. Sellers are sent new product descriptions and pricing suggestions. Buyers are sent product image and description evaluation results.

[0746] Step 10:

[0747] Terminal: The received analysis results are displayed in a specified format and notified to the user. For sellers, the generated product description and pricing suggestions are displayed in a tone based on the emotion engine's analysis results. For buyers, discrepancies and evaluation results are displayed, providing reference information to assist in purchasing decisions.

[0748] This process allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information and emotionally sensitive feedback.

[0749] Example 2

[0750] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0751] Traditional auction and flea market services have faced challenges in terms of the reliability of product descriptions and improving the experience for buyers and sellers. In particular, there were insufficient means to check the consistency between product descriptions and images, and whether or not there were false or exaggerated statements, making it difficult for buyers to make decisions based on these. Furthermore, the impersonal feedback provided meant that services were not provided that took user feelings into consideration.

[0752] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0753] In this invention, the server includes a means for a user to upload product images and product descriptions, a means for the server to receive and store the product images and product descriptions, a means for analyzing the product images and product descriptions using a generative AI model to generate evaluation results, and a means for an emotion engine to analyze the user's emotional information and generate appropriate feedback by adjusting the tone and approach based on the emotional information. This enables the provision of highly reliable product evaluations as well as feedback that takes the user's emotions into consideration.

[0754] "User" refers to an individual or corporation that intends to sell or purchase items using the auction and flea market services.

[0755] "Terminal" refers to the electronic device (smartphone, PC, tablet, etc.) used by the user to input product images and product descriptions and send them to the system.

[0756] "Server" refers to a central management system that receives, stores, and analyzes data sent from user devices, generates evaluation results, and provides them to users.

[0757] "Generative AI model" refers to the artificial intelligence model used by the server to analyze product images and product descriptions and generate evaluation results.

[0758] "Product image" refers to a photograph or image data of the product that a user is listing for sale.

[0759] "Product description" refers to text data that explains the condition, features, price, etc. of the product that a user is selling.

[0760] An "emotion engine" is a system that analyzes emotional information from a user's facial expressions and voice, and provides appropriate feedback based on the results.

[0761] "Feedback" refers to comments, advice, notifications, etc. provided to users based on the evaluation results generated by the server.

[0762] The present invention is a system that provides convenience to both buyers and sellers in auction and flea market services, and in particular improves the user experience by combining an emotion engine. Specific embodiments are described below.

[0763] The system includes a user device, a server, a generative AI model, and an emotion engine. The user device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[0764] First, the seller launches the device app, takes a picture of the product, and enters a simple description. For example, the seller might enter a description such as "old camera, operation not confirmed." The device temporarily saves the captured image and the entered description in local storage. Next, the emotion engine uses the device's camera and microphone to analyze the seller's emotions.

[0765] Next, the device sends the product image, product description, and emotion information to the server. The server receives this data and stores it in a database. The product image is stored as image data, the description as text data, and the emotion information as structured data.

[0766] The server uses a generative AI model to analyze product images. Specifically, it uses object recognition technology to identify the product category and evaluate the product's condition. The generative AI model also performs text analysis of the product description to extract important keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0767] The emotion engine then generates feedback for sellers and buyers based on the emotional information analyzed. For sellers, it presents pricing suggestions based on appropriate product descriptions and the value of the product. For example, the server generates a description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good," and provides it to the seller. For buyers, it provides feedback on the degree of match between the product image and description, as well as any inconsistencies and false statements.

[0768] Furthermore, the tone and content of the feedback is adjusted based on the results of the emotion engine. For example, if the buyer is skeptical, detailed feedback such as "The description and images are consistent. The phone looks good, but there are some small scratches in the photos. The description is generally accurate."

[0769] Finally, the analysis results and feedback generated by the server are sent to the user's device and displayed in an appropriate format. The seller can review the generated product description and pricing suggestions and make any necessary corrections. The buyer can make a purchase decision based on the analysis results provided and with reliable information.

[0770] For example, if a seller takes a picture of an old camera and enters the description "Old camera, not working," the emotion engine recognizes that the seller is relaxed. The server analyzes the image and text, and the generative AI model generates an appropriate description, such as "Vintage camera from the 1940s, shutter not working, but looks good," and provides feedback in a relaxed tone.

[0771] If a buyer clicks on a smartphone product and sees the product image and the description "High-performance smartphone, almost no scratches," the emotion engine recognizes the buyer's doubts. The server analyzes the image and text and generates a rating result: "The description and the image are consistent. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate," providing detailed and reliable feedback.

[0772] This system allows both users and sellers to obtain reliable information, sellers to efficiently generate appropriate product descriptions, and buyers to make purchasing decisions by receiving feedback that takes into account their emotions.

[0773] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0774] Step 1:

[0775] User: The seller launches the app on their device and takes a picture of the product. For example, they take a picture of an "old camera from the 1940s."

[0776] Input: The captured image.

[0777] Specific operation: The device app uses the camera function to temporarily save the captured image to local storage.

[0778] Output: Product images saved to local storage.

[0779] Step 2:

[0780] User: The seller enters the product description into the terminal. For example, they might enter "old camera, not working."

[0781] Input: The entered product description.

[0782] Specific operation: The terminal app temporarily saves the entered text data in local storage.

[0783] Output: Product description saved to local storage.

[0784] Step 3:

[0785] On the device: The emotion engine uses the device's camera and microphone to analyze the seller's facial expressions and voice.

[0786] Input: Seller's facial expression and voice data.

[0787] How it works: The sentiment engine uses machine learning algorithms to analyze the emotional state of sellers in real time.

[0788] Output: Seller emotional state data.

[0789] Step 4:

[0790] Device: Sends product images, product descriptions, and emotion information to the server.

[0791] Input: Product image, product description, sentiment information.

[0792] Specific operation: The device uses the network and sends the data as a data packet to the server.

[0793] Output: The server receives the product image, description, and emotion information as a data packet.

[0794] Step 5:

[0795] Server: Stores the received product images, product descriptions, and emotion information in a database.

[0796] Input: Received data (product image, product description, emotional information).

[0797] Specific operation: The server converts the data into a data format and stores it in a database. Product images are saved as image data, descriptions as text data, and emotional information as structured data.

[0798] Output: Product images, product descriptions, and sentiment information stored in a database.

[0799] Step 6:

[0800] Server: Analyzes product images using a generative AI model and identifies the product category using object recognition technology.

[0801] Input: Product images stored in the database.

[0802] How it works: The generative AI model uses object recognition algorithms to extract features in an image and identify the product category.

[0803] Output: Identified product category and product condition.

[0804] Step 7:

[0805] Server: Uses a generative AI model to analyze the text of product descriptions, extracting important keywords and checking for contextual inconsistencies, misrepresentations, and exaggerations.

[0806] Input: Product description stored in the database.

[0807] How it works: The generative AI model uses natural language processing techniques to analyze text, extract keywords, analyze context, and check for misrepresentations.

[0808] Output: Keyword extraction results, contextual inconsistencies, misrepresentations, and exaggerations.

[0809] Step 8:

[0810] Server: Generates feedback for sellers and buyers based on the emotional information analyzed by the emotion engine.

[0811] Input: Emotion information, product image analysis results, product description analysis results.

[0812] Specific operation: The server takes into account emotional information and generates appropriate product description and pricing suggestions for the seller, and generates feedback for the buyer regarding the degree of match between the description and the image, any inconsistencies, and whether there are any misrepresentations.

[0813] Output: Feedback for seller (product description, suggested pricing), feedback for buyer (match between image and description, inconsistencies, misrepresentations).

[0814] Step 9:

[0815] Server: Sends the generated feedback to the user's device.

[0816] Input: The generated feedback.

[0817] Specific operation: The server sends feedback data to the terminal via the network.

[0818] Output: Feedback data received by the user terminal.

[0819] Step 10:

[0820] Terminal: Displays the received feedback data in an appropriate format.

[0821] Input: The received feedback data.

[0822] Specific operation: The device analyzes the feedback data and displays it on the user interface. The seller is shown the generated product description and pricing suggestions, and the buyer is shown the analysis results.

[0823] Output: Feedback displayed on the device.

[0824] (Application example 2)

[0825] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0826] In conventional auction and flea market services, if the product descriptions and images provided by sellers contain inconsistencies or falsehoods, buyers' trust is often damaged. It is also difficult for sellers to set appropriate prices and product descriptions, and there is a lack of advice and feedback to stimulate purchasing motivation. Another problem is the lack of a system that provides feedback based on the user's emotional state. Given this background, there is a need for a system that can provide accurate and reliable information while taking user emotions into consideration.

[0827] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0828] In this invention, the server includes: means for a user to upload product images and product descriptions; means for the server to receive and store the product images and product descriptions; means for the server to analyze the product images and product descriptions using a generative AI model and generate evaluation results; means for the server to provide the evaluation results to the user; means for an emotion engine to analyze the user's emotions; and means for generating and providing feedback to the user based on the analysis results and emotion information. This makes it possible to provide highly reliable feedback that takes the user's emotions into consideration.

[0829] "User emotion" refers to the user's psychological state and emotions as read and analyzed by the emotion engine.

[0830] "Emotion engine" is a general term for software and algorithms used to analyze emotions based on data such as a user's facial expressions and voice.

[0831] A "generative AI model" is an artificial intelligence model that analyzes product images and descriptions and generates the necessary feedback.

[0832] "Product images" refer to photographs and image data of the products being offered for sale.

[0833] A "product description" is a text description of the characteristics and condition of the product being offered for sale.

[0834] "Analysis results" refer to the evaluations and feedback generated by generative AI models and emotion engines.

[0835] "Feedback" refers to advice and evaluation information provided to the user based on the analysis results and emotional information.

[0836] A "server" is a centralized processing unit for receiving, storing, analyzing, and providing feedback to the user of data.

[0837] "Upload" refers to the act of a user sending data from their own terminal to a server.

[0838] "Receiving" means that the server takes in data such as product images and product descriptions sent from the user's terminal.

[0839] "Storing" means storing the received data in a storage device such as a database.

[0840] "Analysis" refers to using generative AI models and emotion engines to process product images and descriptions to generate ratings and feedback.

[0841] "Inconsistencies" refer to discrepancies or inconsistencies between product images and descriptions.

[0842] "False representation" refers to information contained in a product description that is incorrect and different from the facts.

[0843] "Exaggeration" refers to descriptions that exaggerate the actual characteristics of a product.

[0844] The present invention is a system that includes a means for a user to upload product images and product descriptions, a means for a server to receive and store the product images and product descriptions, a means for the server to analyze the product images and product descriptions using a generative AI model and generate an evaluation result, a means for the server to provide the evaluation result to the user, a means for an emotion engine to analyze the user's emotions, and a means for generating feedback based on the analysis result and emotion information and providing it to the user.

[0845] Overall system configuration

[0846] The system includes a user's device, a server, a generative AI model, and an emotion engine. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[0847] Program processing explanation

[0848] Uploading data

[0849] On-device: The seller launches the app on their device, takes a photo of the product, and enters a simple description. For example, they can enter a description such as "old camera, operation not confirmed." In addition, the emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[0850] Receiving and storing data

[0851] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[0852] Image and description analysis

[0853] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0854] Sentiment-based analysis

[0855] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[0856] Generate feedback

[0857] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[0858] Hardware and software used

[0859] Hardware: The smartphone used by the user

[0860] Software: Python, OpenCV, Transformers library, EmotionRecognition library

[0861] Specific examples

[0862] Examples for sellers

[0863] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not working." The emotion engine recognizes that the user is relaxed.

[0864] Terminal: Sends data to the server.

[0865] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." Feedback is provided in a relaxed tone based on the emotional information.

[0866] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[0867] Specific examples for buyers

[0868] User: A buyer clicks on a product they're interested in, sees the product image and description, "High-performance smartphone, virtually flawless." The emotion engine recognizes the user's doubts.

[0869] Device: Press the analyze button to send the data to the server.

[0870] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[0871] Server: Generates a rating like "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate." Based on sentiment, additional information is provided to resolve any doubts.

[0872] Server: The evaluation results are provided to the buyer and displayed on the device.

[0873] Prompt Sentence Examples

[0874] Description of the invention:

[0875] I would like to develop a system to improve the user experience of auction and flea market services. This system combines an emotion engine and a generative AI model to analyze product images and descriptions and provide appropriate feedback to both buyers and sellers. It also includes a function to check for inconsistencies and false statements in product descriptions and suggest optimal pricing to sellers.

[0876] Prerequisites:

[0877] Sellers upload product images and descriptions.

[0878] Analyzing user emotions with an emotion engine

[0879] Analyze and generate product descriptions using a generative AI model

[0880] Give feedback in an emotional tone

[0881] Expected output:

[0882] Relaxed tone of voice and optimal product description and pricing suggestions for sellers

[0883] Feedback to buyers regarding the consistency of product descriptions and images, as well as any inconsistencies

[0884] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0885] Step 1:

[0886] Uploading data

[0887] Device: The seller uses a smartphone to take a picture of the product and enter a brief description. For example, they might enter "old camera, not yet operational." The emotion engine then activates and analyzes the seller's facial expressions and voice to obtain emotional information.

[0888] Input: Product images, product descriptions, seller's facial expressions and voice data

[0889] Output: Product images, product descriptions, emotional information

[0890] Step 2:

[0891] Receiving and storing data

[0892] Server: Receives product images, product descriptions, and emotion information sent from the device. The received data is stored in a database using transaction processing to prepare for analysis.

[0893] Input: Product image, product description, emotional information

[0894] Output: Product images, product descriptions, and emotional information stored in a database

[0895] Step 3:

[0896] Image and description analysis

[0897] Server: Using a generative AI model, it analyzes product images and descriptions. For product images, it uses object recognition technology to evaluate the product category and condition, and for descriptions, it performs text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0898] Input: Product image, product description

[0899] Output: Analysis results of product images, analysis results of product descriptions

[0900] Step 4:

[0901] Sentiment-based analysis

[0902] Server: Based on the emotional information analyzed by the emotion engine, the system processes the user's emotional state. For example, if the user is nervous, the system generates feedback using a gentle tone.

[0903] Input: Emotion information

[0904] Output: Tone information for emotion-based feedback

[0905] Step 5:

[0906] Generate feedback

[0907] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and false statements. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[0908] Input: Analysis results of product images, analysis results of product descriptions, tone information of feedback based on emotions

[0909] Output: Buyer and seller feedback data

[0910] Step 6:

[0911] Providing feedback

[0912] Server: The generated feedback and rating results are sent to the user's device and provided to the seller and buyer. The seller is shown the generated product description and suggested starting price, and the buyer is shown the analysis results of the product image and description.

[0913] Input: Buyer and seller feedback data

[0914] Output: Feedback and evaluation results displayed on the user's device

[0915] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0916] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0917] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0918] [Third embodiment]

[0919] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0920] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0921] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0922] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0923] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0924] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0925] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0926] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0927] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0928] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0929] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0930] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0931] As an embodiment of the present invention, a system that provides convenience to both buyers and sellers in auction and flea market services is shown. The following describes the processing contents of a specific program.

[0932] Overall system configuration

[0933] The system includes a user's device, a server, and a generative AI model. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback.

[0934] Program processing explanation

[0935] Uploading data

[0936] On the device: The seller launches the app on their device, takes a picture of the product, and enters a brief description of the product. After confirming the information entered, they press the "Upload" button to send the data to the server.

[0937] Receiving and storing data

[0938] Server: The server receives the product images and descriptions sent by the user and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[0939] Image and description analysis

[0940] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[0941] Generate analysis results

[0942] Server: The results of image and text analysis are integrated to generate feedback for buyers and sellers. For buyers, feedback is provided on the degree of agreement between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, appropriate product descriptions and pricing suggestions based on the product's value are presented.

[0943] Providing analysis results

[0944] Server: The generated analysis results are sent to the user's device and displayed in a specified format. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer, providing reference information to assist in making a purchasing decision.

[0945] Specific examples

[0946] Examples for sellers

[0947] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not tested."

[0948] Device: Press the upload button to send the data to the server.

[0949] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description: "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." A starting price of 3,000 yen is suggested.

[0950] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[0951] Specific examples for buyers

[0952] User: A buyer clicks on a product they're interested in and sees the product image and description: "High-performance smartphone, virtually no scratches."

[0953] Device: Press the analyze button to send the data to the server.

[0954] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[0955] Server: Generates a rating that reads, "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate."

[0956] Server: The evaluation results are provided to the buyer and displayed on the device.

[0957] This system allows buyers and sellers to enjoy a highly reliable trading environment, improving convenience for both parties.

[0958] The processing flow will be explained below.

[0959] Step 1:

[0960] User: The seller launches the app on their device, takes a picture of the product, and enters a simple description of the product. For example, they can enter a description such as "Old camera, operation not confirmed."

[0961] Step 2:

[0962] Terminal: Checks the entered product image and description, converts them into the required format, and displays the "Upload" button. When the user presses the "Upload" button, the data is sent to the server.

[0963] Step 3:

[0964] Server: Receives product images and descriptions and stores them in a database, which also stores related information about the received images and descriptions.

[0965] Step 4:

[0966] Server: The saved product images and descriptions are passed to the generative AI model and analysis begins. Specifically, two processes, image analysis and text analysis, are performed.

[0967] Step 5:

[0968] Server: The image analysis process involves extracting features from product images, such as identifying the camera model, year of manufacture, and external condition.

[0969] Step 6:

[0970] Server: The text analysis process analyzes the product description, extracts keywords, and checks for inconsistencies, false statements, and exaggerations.

[0971] Step 7:

[0972] Server: Integrates the analysis results and generates feedback for buyers and sellers. For buyers, it evaluates the consistency between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates an appropriate product description and suggests a starting price.

[0973] Step 8:

[0974] Server: The generated analysis results are sent to the user's device. Sellers are sent new product descriptions and pricing suggestions. Buyers are sent product image and description evaluation results.

[0975] Step 9:

[0976] Terminal: The received analysis results are displayed in a specified format and notified to the user. The seller is shown the generated product description and pricing suggestions. The buyer is shown any discrepancies and evaluation results, which are provided as reference information to assist in purchasing decisions.

[0977] This process allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information.

[0978] Example 1

[0979] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0980] Conventional auction and flea market services lack the means to verify the authenticity of product images and descriptions uploaded by users, resulting in problems such as fraud and inappropriate language. Furthermore, sellers often lack support for creating appropriate product descriptions and setting prices, resulting in transactions that do not proceed smoothly. The objective of this invention is to provide a system that analyzes product images and descriptions and provides evaluation results to solve these problems.

[0981] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0982] In this invention, the server includes a means for analyzing product images and product descriptions using a generative AI model, performing feature extraction and text analysis, a means for generating evaluation results based on the results of the feature extraction and text analysis, and a means for providing the evaluation results to the user. This ensures the reliability of the product images and descriptions, allowing users to conduct transactions with peace of mind. Furthermore, the server can also provide support for the appropriateness of product descriptions and pricing, which is expected to facilitate smooth transactions.

[0983] "User" means a person or organization that uses the System to upload product images and product descriptions.

[0984] "Product image" refers to an image file that shows visual information about a product and is uploaded to the system by a seller.

[0985] "Product description" is text information that describes the characteristics and condition of a product and is uploaded by the seller to the system.

[0986] "Server" refers to a set of hardware and software systems that receives, stores, and analyzes product images and product descriptions, generates evaluation results, and provides them to users.

[0987] A "generative AI model" is an artificial intelligence model used to analyze product images and product descriptions and perform feature extraction and text analysis.

[0988] "Feature extraction" is the process of using a generative AI model to extract feature information such as product category and condition from product images.

[0989] "Text analysis" is the process of using generative AI models to extract keywords from product descriptions and check for contextual inconsistencies, misrepresentations, and exaggerations.

[0990] The "evaluation results" are feedback information regarding the reliability of a product, appropriate descriptions, and pricing, created based on the results of feature extraction and text analysis.

[0991] The "means for providing to the user" refers to a series of processes and system configuration for transmitting the evaluation results generated by the server to the user's terminal and displaying them.

[0992] This invention is a system that provides convenience to buyers and sellers in auction and flea market services. The system includes a user terminal, a server, and a generative AI model, and the specific program processing content will be explained below.

[0993] Overall system configuration

[0994] The system includes a user's device, a server, and a generative AI model. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback.

[0995] Uploading data

[0996] As a seller, the user launches the terminal app, takes a picture of the product, and enters a brief description of the product. After the device checks the entered product image and description, the user presses the "Upload" button, which sends this data to the server.

[0997] As a specific example of operation, a user takes an image with an old camera, enters a simple description such as "old camera, operation not confirmed," and presses the "upload" button.

[0998] Receiving and storing data

[0999] The server receives product images and descriptions submitted by users, and stores the data in a database, ready for analysis by the generative AI model.

[1000] Image and description analysis

[1001] The server uses a generative AI model to analyze product images and descriptions. For product images, object recognition technology is used to extract features and evaluate the product's category and condition. For product descriptions, text analysis is performed to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1002] As a specific example of how it works, the generative AI model analyzes an image of an old camera, generates a description such as "A vintage camera from the 1940s. The shutter function has not been confirmed, but the appearance is good," and suggests a starting price of 3,000 yen.

[1003] Generate analysis results

[1004] The server combines the results of image and text analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and false statements. For sellers, it presents appropriate product descriptions and pricing suggestions based on the product's value.

[1005] As a specific example of how it works, when a buyer checks a product described as "a high-performance smartphone with almost no scratches" and presses the analysis button, the server generates an evaluation that reads, "There is no inconsistency between the description and the image. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate."

[1006] Providing analysis results

[1007] The server sends the generated analysis results to the user's device and displays them in a specified format. The seller is shown the generated product description and pricing suggestions. The buyer is shown the analysis results of the product image and description, which are provided as reference information to assist in purchasing decisions.

[1008] This system allows buyers and sellers to enjoy a highly reliable trading environment, improving convenience for both parties.

[1009] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1010] Step 1:

[1011] Uploading data

[1012] User: The seller launches the app on their device, takes a picture of the product, and enters a description of the product. The user then presses the "Upload" button to send the data.

[1013] Input: Product image and description

[1014] Output: A data packet of product images and descriptions sent to the server

[1015] Specific behavior:

[1016] The seller launches the terminal app.

[1017] The seller takes a picture of the product.

[1018] The seller enters the product description.

[1019] The seller presses the "Upload" button.

[1020] The device app sends product image and description data to the server.

[1021] Step 2:

[1022] Receiving and storing data

[1023] Server: The server receives the product images and product descriptions sent by the user and stores the received data in a database.

[1024] Input: Product image and description sent from the user's device

[1025] Output: Product images and descriptions stored in a database

[1026] Specific behavior:

[1027] The server receives the data from the terminal.

[1028] The server converts the received data into an appropriate format for analysis.

[1029] The server stores the converted data in a database.

[1030] Step 3:

[1031] Image and description analysis

[1032] Server: The server uses generative AI models to analyze product images and product descriptions. It performs object recognition on product images to assess product category and condition, and performs text analysis on product descriptions to extract keywords and check for inconsistencies, false statements, and exaggerations.

[1033] Input: Product images and descriptions stored in the database

[1034] Output: Analyzed feature information and text analysis results

[1035] Specific behavior:

[1036] The server inputs product images into a generative AI model.

[1037] The generative AI model performs object recognition on the image and extracts the product's features.

[1038] The server inputs the product description into the generative AI model.

[1039] The generative AI model analyzes the text, extracts keywords, and checks for contextual inconsistencies.

[1040] Step 4:

[1041] Generate analysis results

[1042] Server: The server combines the results of image analysis and text analysis to generate feedback for buyers and sellers.

[1043] For buyers, it provides evaluation results on the degree of agreement between the description and the image, as well as any inconsistencies and false statements, while for sellers, it presents suggestions for appropriate product descriptions and pricing.

[1044] Input: Image analysis and text analysis results

[1045] Output: Feedback for buyers and sellers

[1046] Specific behavior:

[1047] The server combines the results of image analysis and text analysis.

[1048] The server generates feedback for the buyer.

[1049] The server generates feedback for the seller.

[1050] Step 5:

[1051] Providing analysis results

[1052] Server: The server sends the generated analysis results to the user's device. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer.

[1053] Input: Feedback for buyers and sellers

[1054] Output: Feedback information displayed on the user's terminal

[1055] Specific behavior:

[1056] The server sends the feedback to the user's terminal.

[1057] The generated product description and pricing suggestions are displayed on the seller's device.

[1058] The analysis results of the product image and description are displayed on the buyer's device.

[1059] (Application example 1)

[1060] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1061] In modern auction and flea market services, many problems arise due to the lack of trust between buyers and sellers. This can make buyers unsure about the safety of their transactions, and sellers find it difficult to properly evaluate and price their products. Furthermore, there is a lack of means to detect inconsistencies and false representations between product images and descriptions, making it difficult to achieve transparent transactions.

[1062] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1063] In this invention, the server includes a means for users to upload product images and product descriptions, a means for the server to receive and store the product images and product descriptions, a means for analyzing the product images and product descriptions using a generative AI model to generate evaluation results and security check results, and a means for users to check the evaluation results and security check results. This makes it possible to check for inconsistencies and false statements between product images and descriptions, generate appropriate product descriptions, present pricing options, and provide security check results.

[1064] "User" refers to sellers who use the system to upload product images and product descriptions, and buyers who check the analysis results.

[1065] "Product images" are photo files of products that users sell.

[1066] "Product description" is text data that describes in sentence format information about the product that the user is selling.

[1067] A "server" is a computer system that receives product images and product descriptions sent by users, analyzes them using a generative AI model, and stores the results.

[1068] "Generative AI model" refers to an algorithm that uses machine learning technology to analyze product images and product descriptions and generate evaluation results.

[1069] "Analysis" is the process by which the generative AI model analyzes product images and product descriptions to evaluate the accuracy of the product's features and descriptions.

[1070] "Evaluation results" are information about the accuracy of the product's characteristics and description obtained through analysis by the generative AI model.

[1071] "Security check results" are information that summarizes the results of detecting fraudulent expressions or potentially misleading statements through analysis of product images and product descriptions.

[1072] "Verifiable means" refers to the interface and functionality that allows users to view evaluation results and security check results.

[1073] "Inconsistencies" refers to inconsistencies or inconsistencies between product images and product descriptions.

[1074] "Misrepresentation" means a product description that is untrue or misleading.

[1075] "Product description generation" is the process in which a generative AI model automatically creates a more specific and attractive description based on the original description.

[1076] "Pricing suggestions" are the estimated value and price suggestions for a product that the generative AI model presents to the user based on the analysis results.

[1077] As an embodiment of the present invention, a system that provides convenience and security to both buyers and sellers in auction and flea market services is shown. The system includes a user terminal, a server, and a generative AI model.

[1078] Overall system configuration

[1079] 1. On the user's device:

[1080] Sellers use a terminal to take product images, enter product descriptions, and upload them.

[1081] The buyer uses a terminal to view the listed items and check the analysis results and security check results.

[1082] 2. Server:

[1083] The server receives the product images and product descriptions sent by the user and stores them in a database.

[1084] A generative AI model is used to analyze product images and descriptions, and generate evaluation and security check results.

[1085] The generated results are sent to the user's terminal and displayed in a predetermined format.

[1086] 3. Generative AI Model:

[1087] For product images, features are extracted using object recognition technology to evaluate the product category and condition.

[1088] For product descriptions, text analysis technology is used to extract keywords and detect contradictions, false statements, and exaggerated expressions.

[1089] Program processing explanation

[1090] Uploading data

[1091] On the device: The seller launches the app on their device, takes a picture of the product, and enters a brief description of the product. After confirming the information entered, they press the "Upload" button to send the data to the server.

[1092] Receiving and storing data

[1093] Server: The server receives the product images and descriptions sent by the user and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[1094] Image and description analysis

[1095] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1096] Generate analysis results

[1097] Server: The results of image and text analysis are integrated to generate feedback for buyers and sellers. For buyers, feedback is provided on the degree of agreement between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, appropriate product descriptions and pricing suggestions based on the product's value are presented.

[1098] Providing analysis results

[1099] Server: The generated analysis results are sent to the user's device and displayed in a specified format. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer, providing reference information to assist in making a purchasing decision.

[1100] Specific examples

[1101] Examples for sellers

[1102] The seller takes a picture of the old camera and enters a simple description such as "Old camera, operation not confirmed." Presses the upload button on the device to send the data to the server. The server analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "Vintage camera from the 1940s, shutter operation not confirmed, but appearance is good." A suggested starting price of 3,000 yen is presented. The server provides the generated product description and suggested starting price to the seller, which are displayed on the device.

[1103] Specific examples for buyers

[1104] A buyer clicks on a product they are interested in and checks the product image and description: "High-performance smartphone, almost no scratches." They then press the analysis button on their device to send the data to the server. The server analyzes the product image and description and checks whether the actual appearance matches the description. The server generates a rating: "There is no discrepancy between the description and the image. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate." The server then provides the rating result to the buyer, which is displayed on the device.

[1105] Prompt Sentence Examples

[1106] Seller prompt:

[1107] Product description: Old camera, operation not confirmed

[1108] Please change this to a more specific and compelling description.

[1109] Buyer prompt:

[1110] Description: High-performance smartphone, almost no scratches

[1111] Please check for discrepancies between images and descriptions and provide feedback on the actual condition.

[1112] These features allow users to conduct transactions with high reliability and transparency.

[1113] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1114] Step 1:

[1115] The user uses a device to take a picture of the product and enter a description. The user starts the application on the device, takes a picture of the product, and enters a description. After checking the entered information, the user presses the "Upload" button.

[1116] Input: Product image file, product description

[1117] Output: Upload request

[1118] Step 2:

[1119] The device sends the product image taken by the user and the product description entered by the user to the server. During this process, the image file and text data are transferred to the server in an appropriate format.

[1120] Input: Upload request (product image file, product description)

[1121] Output: Data sent to the server

[1122] Step 3:

[1123] The server stores the product images and description received from the device in a database, and converts the received data into the required format for analysis.

[1124] Input: Data to be sent (product image, product description)

[1125] Output: Data saved to database

[1126] Step 4:

[1127] The server uses a generative AI model to analyze product images and product descriptions. For product images, object recognition technology (e.g., TensorFlow or PyTorch) is used to extract features and evaluate the product's category and condition. For product descriptions, text analysis technology (e.g., Hugging Face's T5 model) is used to extract keywords and detect contradictions, false statements, and exaggerations.

[1128] Input: Product images and product descriptions stored in the database

[1129] Output: Analysis results (verification results of product category, condition, description)

[1130] Step 5:

[1131] The server combines the results of image and text analysis to generate feedback for buyers and sellers. For buyers, it generates an evaluation result on the degree of match between the image and description, as well as any inconsistencies and false statements. For sellers, it presents appropriate product descriptions and pricing suggestions based on the value of the product.

[1132] Input: Analysis results (product category, condition, and description verification results)

[1133] Output: Feedback (for buyers and sellers)

[1134] Step 6:

[1135] The server sends the generated feedback to the user's terminal. The seller is provided with the generated product description and pricing suggestions, and the buyer is provided with the analysis results of the product image and description.

[1136] Input: Feedback (for buyers, for sellers)

[1137] Output: Feedback data to the user terminal

[1138] Step 7:

[1139] The terminal displays the feedback received from the server. The seller checks the generated product description and pricing suggestions, and the buyer views the analysis results of the product image and description.

[1140] Input: Feedback data sent from the server

[1141] Output: Displayed analysis results and suggestions

[1142] In this way, sellers and buyers can conduct transactions with high reliability and transparency.

[1143] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1144] The present invention is a system that provides convenience to both buyers and sellers in auction and flea market services, and in particular, by combining it with an emotion engine, it further improves the user experience. Below, we will explain the specific program processing content.

[1145] Overall system configuration

[1146] The system includes a user's device, a server, a generative AI model, and an emotion engine. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[1147] Program processing explanation

[1148] Uploading data

[1149] On-device: The seller launches the app on their device, takes a photo of the product, and enters a simple description. For example, they can enter a description such as "old camera, operation not confirmed." In addition, the emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[1150] Receiving and storing data

[1151] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[1152] Image and description analysis

[1153] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1154] Sentiment-based analysis

[1155] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[1156] Generate feedback

[1157] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[1158] Providing analysis results

[1159] Server: The generated analysis results are sent to the user's device and displayed in a specified format. When the generated product description and pricing suggestions are displayed to the seller, they are displayed in a tone based on the analysis results of the emotion engine. The analysis results of the product image and description are displayed to the buyer, providing reference information to support their purchasing decision.

[1160] Specific examples

[1161] Examples for sellers

[1162] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not working." The emotion engine recognizes that the user is relaxed.

[1163] Terminal: Sends data to the server.

[1164] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." Feedback is provided in a relaxed tone based on the emotional information.

[1165] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[1166] Specific examples for buyers

[1167] User: A buyer clicks on a product they're interested in, sees the product image and description, "High-performance smartphone, virtually flawless." The emotion engine recognizes the user's doubts.

[1168] Device: Press the analyze button to send the data to the server.

[1169] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[1170] Server: Generates a rating like "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate." Based on sentiment, additional information is provided to resolve any doubts.

[1171] Server: The evaluation results are provided to the buyer and displayed on the device.

[1172] This system allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information and receiving feedback that takes into account their emotions.

[1173] The processing flow will be explained below.

[1174] Step 1:

[1175] User: The seller launches the app on their device, takes a picture of the product, and writes a brief description of the product. The emotion engine then analyzes the seller's emotional state (e.g., nervous, relaxed, anxious, etc.).

[1176] Step 2:

[1177] Terminal: Checks the entered product image, description, and emotion information, converts them into the required format, and displays the "Upload" button. When the user presses the "Upload" button, the data is sent to the server.

[1178] Step 3:

[1179] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[1180] Step 4:

[1181] Server: The saved product images and descriptions are passed to the generative AI model and analysis begins. Specifically, two processes, image analysis and text analysis, are performed.

[1182] Step 5:

[1183] Server: The image analysis process involves extracting features from product images, such as identifying the camera model, year of manufacture, and external condition.

[1184] Step 6:

[1185] Server: The text analysis process analyzes the product description, extracts keywords, and checks for inconsistencies, misrepresentations, and exaggerations.

[1186] Step 7:

[1187] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[1188] Step 8:

[1189] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[1190] Step 9:

[1191] Server: The generated analysis results are sent to the user's device. Sellers are sent new product descriptions and pricing suggestions. Buyers are sent product image and description evaluation results.

[1192] Step 10:

[1193] Terminal: The received analysis results are displayed in a specified format and notified to the user. For sellers, the generated product description and pricing suggestions are displayed in a tone based on the emotion engine's analysis results. For buyers, discrepancies and evaluation results are displayed, providing reference information to assist in purchasing decisions.

[1194] This process allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information and emotionally sensitive feedback.

[1195] Example 2

[1196] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1197] Traditional auction and flea market services have faced challenges in terms of the reliability of product descriptions and improving the experience for buyers and sellers. In particular, there were insufficient means to check the consistency between product descriptions and images, and whether or not there were false or exaggerated statements, making it difficult for buyers to make decisions based on these. Furthermore, the impersonal feedback provided meant that services were not provided that took user feelings into consideration.

[1198] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1199] In this invention, the server includes a means for a user to upload product images and product descriptions, a means for the server to receive and store the product images and product descriptions, a means for analyzing the product images and product descriptions using a generative AI model to generate evaluation results, and a means for an emotion engine to analyze the user's emotional information and generate appropriate feedback by adjusting the tone and approach based on the emotional information. This enables the provision of highly reliable product evaluations as well as feedback that takes the user's emotions into consideration.

[1200] "User" refers to an individual or corporation that intends to sell or purchase items using the auction and flea market services.

[1201] "Terminal" refers to the electronic device (smartphone, PC, tablet, etc.) used by the user to input product images and product descriptions and send them to the system.

[1202] "Server" refers to a central management system that receives, stores, and analyzes data sent from user devices, generates evaluation results, and provides them to users.

[1203] "Generative AI model" refers to the artificial intelligence model used by the server to analyze product images and product descriptions and generate evaluation results.

[1204] "Product image" refers to a photograph or image data of the product that a user is listing for sale.

[1205] "Product description" refers to text data that explains the condition, features, price, etc. of the product that a user is selling.

[1206] An "emotion engine" is a system that analyzes emotional information from a user's facial expressions and voice, and provides appropriate feedback based on the results.

[1207] "Feedback" refers to comments, advice, notifications, etc. provided to users based on the evaluation results generated by the server.

[1208] The present invention is a system that provides convenience to both buyers and sellers in auction and flea market services, and in particular improves the user experience by combining an emotion engine. Specific embodiments are described below.

[1209] The system includes a user device, a server, a generative AI model, and an emotion engine. The user device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[1210] First, the seller launches the device app, takes a picture of the product, and enters a simple description. For example, the seller might enter a description such as "old camera, operation not confirmed." The device temporarily saves the captured image and the entered description in local storage. Next, the emotion engine uses the device's camera and microphone to analyze the seller's emotions.

[1211] Next, the device sends the product image, product description, and emotion information to the server. The server receives this data and stores it in a database. The product image is stored as image data, the description as text data, and the emotion information as structured data.

[1212] The server uses a generative AI model to analyze product images. Specifically, it uses object recognition technology to identify the product category and evaluate the product's condition. The generative AI model also performs text analysis of the product description to extract important keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1213] The emotion engine then generates feedback for sellers and buyers based on the emotional information analyzed. For sellers, it presents pricing suggestions based on appropriate product descriptions and the value of the product. For example, the server generates a description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good," and provides it to the seller. For buyers, it provides feedback on the degree of match between the product image and description, as well as any inconsistencies and false statements.

[1214] Furthermore, the tone and content of the feedback is adjusted based on the results of the emotion engine. For example, if the buyer is skeptical, detailed feedback such as "The description and images are consistent. The phone looks good, but there are some small scratches in the photos. The description is generally accurate."

[1215] Finally, the analysis results and feedback generated by the server are sent to the user's device and displayed in an appropriate format. The seller can review the generated product description and pricing suggestions and make any necessary corrections. The buyer can make a purchase decision based on the analysis results provided and with reliable information.

[1216] For example, if a seller takes a picture of an old camera and enters the description "Old camera, not working," the emotion engine recognizes that the seller is relaxed. The server analyzes the image and text, and the generative AI model generates an appropriate description, such as "Vintage camera from the 1940s, shutter not working, but looks good," and provides feedback in a relaxed tone.

[1217] If a buyer clicks on a smartphone product and sees the product image and the description "High-performance smartphone, almost no scratches," the emotion engine recognizes the buyer's doubts. The server analyzes the image and text and generates a rating result: "The description and the image are consistent. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate," providing detailed and reliable feedback.

[1218] This system allows both users and sellers to obtain reliable information, sellers to efficiently generate appropriate product descriptions, and buyers to make purchasing decisions by receiving feedback that takes into account their emotions.

[1219] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1220] Step 1:

[1221] User: The seller launches the app on their device and takes a picture of the product. For example, they take a picture of an "old camera from the 1940s."

[1222] Input: The captured image.

[1223] Specific operation: The device app uses the camera function to temporarily save the captured image to local storage.

[1224] Output: Product images saved to local storage.

[1225] Step 2:

[1226] User: The seller enters the product description into the terminal. For example, they might enter "old camera, not working."

[1227] Input: The entered product description.

[1228] Specific operation: The terminal app temporarily saves the entered text data in local storage.

[1229] Output: Product description saved to local storage.

[1230] Step 3:

[1231] On the device: The emotion engine uses the device's camera and microphone to analyze the seller's facial expressions and voice.

[1232] Input: Seller's facial expression and voice data.

[1233] How it works: The sentiment engine uses machine learning algorithms to analyze the emotional state of sellers in real time.

[1234] Output: Seller emotional state data.

[1235] Step 4:

[1236] Device: Sends product images, product descriptions, and emotion information to the server.

[1237] Input: Product image, product description, sentiment information.

[1238] Specific operation: The device uses the network and sends the data as a data packet to the server.

[1239] Output: The server receives the product image, description, and emotion information as a data packet.

[1240] Step 5:

[1241] Server: Stores the received product images, product descriptions, and emotion information in a database.

[1242] Input: Received data (product image, product description, emotional information).

[1243] Specific operation: The server converts the data into a data format and stores it in a database. Product images are saved as image data, descriptions as text data, and emotional information as structured data.

[1244] Output: Product images, product descriptions, and sentiment information stored in a database.

[1245] Step 6:

[1246] Server: Analyzes product images using a generative AI model and identifies the product category using object recognition technology.

[1247] Input: Product images stored in the database.

[1248] How it works: The generative AI model uses object recognition algorithms to extract features in an image and identify the product category.

[1249] Output: Identified product category and product condition.

[1250] Step 7:

[1251] Server: Uses a generative AI model to analyze the text of product descriptions, extracting important keywords and checking for contextual inconsistencies, misrepresentations, and exaggerations.

[1252] Input: Product description stored in the database.

[1253] How it works: The generative AI model uses natural language processing techniques to analyze text, extract keywords, analyze context, and check for misrepresentations.

[1254] Output: Keyword extraction results, contextual inconsistencies, misrepresentations, and exaggerations.

[1255] Step 8:

[1256] Server: Generates feedback for sellers and buyers based on the emotional information analyzed by the emotion engine.

[1257] Input: Emotion information, product image analysis results, product description analysis results.

[1258] Specific operation: The server takes into account emotional information and generates appropriate product description and pricing suggestions for the seller, and generates feedback for the buyer regarding the degree of match between the description and the image, any inconsistencies, and whether there are any misrepresentations.

[1259] Output: Feedback for seller (product description, suggested pricing), feedback for buyer (match between image and description, inconsistencies, misrepresentations).

[1260] Step 9:

[1261] Server: Sends the generated feedback to the user's device.

[1262] Input: The generated feedback.

[1263] Specific operation: The server sends feedback data to the terminal via the network.

[1264] Output: Feedback data received by the user terminal.

[1265] Step 10:

[1266] Terminal: Displays the received feedback data in an appropriate format.

[1267] Input: The received feedback data.

[1268] Specific operation: The device analyzes the feedback data and displays it on the user interface. The seller is shown the generated product description and pricing suggestions, and the buyer is shown the analysis results.

[1269] Output: Feedback displayed on the device.

[1270] (Application example 2)

[1271] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1272] In conventional auction and flea market services, if the product descriptions and images provided by sellers contain inconsistencies or falsehoods, buyers' trust is often damaged. It is also difficult for sellers to set appropriate prices and product descriptions, and there is a lack of advice and feedback to stimulate purchasing motivation. Another problem is the lack of a system that provides feedback based on the user's emotional state. Given this background, there is a need for a system that can provide accurate and reliable information while taking user emotions into consideration.

[1273] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1274] In this invention, the server includes: means for a user to upload product images and product descriptions; means for the server to receive and store the product images and product descriptions; means for the server to analyze the product images and product descriptions using a generative AI model and generate evaluation results; means for the server to provide the evaluation results to the user; means for an emotion engine to analyze the user's emotions; and means for generating and providing feedback to the user based on the analysis results and emotion information. This makes it possible to provide highly reliable feedback that takes the user's emotions into consideration.

[1275] "User emotion" refers to the user's psychological state and emotions as read and analyzed by the emotion engine.

[1276] "Emotion engine" is a general term for software and algorithms used to analyze emotions based on data such as a user's facial expressions and voice.

[1277] A "generative AI model" is an artificial intelligence model that analyzes product images and descriptions and generates the necessary feedback.

[1278] "Product images" refer to photographs and image data of the products being offered for sale.

[1279] A "product description" is a text description of the characteristics and condition of the product being offered for sale.

[1280] "Analysis results" refer to the evaluations and feedback generated by generative AI models and emotion engines.

[1281] "Feedback" refers to advice and evaluation information provided to the user based on the analysis results and emotional information.

[1282] A "server" is a centralized processing unit for receiving, storing, analyzing, and providing feedback to the user of data.

[1283] "Upload" refers to the act of a user sending data from their own terminal to a server.

[1284] "Receiving" means that the server takes in data such as product images and product descriptions sent from the user's terminal.

[1285] "Storing" means storing the received data in a storage device such as a database.

[1286] "Analysis" refers to using generative AI models and emotion engines to process product images and descriptions to generate ratings and feedback.

[1287] "Inconsistencies" refer to discrepancies or inconsistencies between product images and descriptions.

[1288] "False representation" refers to information contained in a product description that is incorrect and different from the facts.

[1289] "Exaggeration" refers to descriptions that exaggerate the actual characteristics of a product.

[1290] The present invention is a system that includes a means for a user to upload product images and product descriptions, a means for a server to receive and store the product images and product descriptions, a means for the server to analyze the product images and product descriptions using a generative AI model and generate an evaluation result, a means for the server to provide the evaluation result to the user, a means for an emotion engine to analyze the user's emotions, and a means for generating feedback based on the analysis result and emotion information and providing it to the user.

[1291] Overall system configuration

[1292] The system includes a user's device, a server, a generative AI model, and an emotion engine. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[1293] Program processing explanation

[1294] Uploading data

[1295] On-device: The seller launches the app on their device, takes a photo of the product, and enters a simple description. For example, they can enter a description such as "old camera, operation not confirmed." In addition, the emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[1296] Receiving and storing data

[1297] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[1298] Image and description analysis

[1299] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1300] Sentiment-based analysis

[1301] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[1302] Generate feedback

[1303] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[1304] Hardware and software used

[1305] Hardware: The smartphone used by the user

[1306] Software: Python, OpenCV, Transformers library, EmotionRecognition library

[1307] Specific examples

[1308] Examples for sellers

[1309] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not working." The emotion engine recognizes that the user is relaxed.

[1310] Terminal: Sends data to the server.

[1311] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." Feedback is provided in a relaxed tone based on the emotional information.

[1312] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[1313] Specific examples for buyers

[1314] User: A buyer clicks on a product they're interested in, sees the product image and description, "High-performance smartphone, virtually flawless." The emotion engine recognizes the user's doubts.

[1315] Device: Press the analyze button to send the data to the server.

[1316] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[1317] Server: Generates a rating like "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate." Based on sentiment, additional information is provided to resolve any doubts.

[1318] Server: The evaluation results are provided to the buyer and displayed on the device.

[1319] Prompt Sentence Examples

[1320] Description of the invention:

[1321] I would like to develop a system to improve the user experience of auction and flea market services. This system combines an emotion engine and a generative AI model to analyze product images and descriptions and provide appropriate feedback to both buyers and sellers. It also includes a function to check for inconsistencies and false statements in product descriptions and suggest optimal pricing to sellers.

[1322] Prerequisites:

[1323] Sellers upload product images and descriptions.

[1324] Analyzing user emotions with an emotion engine

[1325] Analyze and generate product descriptions using a generative AI model

[1326] Give feedback in an emotional tone

[1327] Expected output:

[1328] Relaxed tone of voice and optimal product description and pricing suggestions for sellers

[1329] Feedback to buyers regarding the consistency of product descriptions and images, as well as any inconsistencies

[1330] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1331] Step 1:

[1332] Uploading data

[1333] Device: The seller uses a smartphone to take a picture of the product and enter a brief description. For example, they might enter "old camera, not yet operational." The emotion engine then activates and analyzes the seller's facial expressions and voice to obtain emotional information.

[1334] Input: Product images, product descriptions, seller's facial expressions and voice data

[1335] Output: Product images, product descriptions, emotional information

[1336] Step 2:

[1337] Receiving and storing data

[1338] Server: Receives product images, product descriptions, and emotion information sent from the device. The received data is stored in a database using transaction processing to prepare for analysis.

[1339] Input: Product image, product description, emotional information

[1340] Output: Product images, product descriptions, and emotional information stored in a database

[1341] Step 3:

[1342] Image and description analysis

[1343] Server: Using a generative AI model, it analyzes product images and descriptions. For product images, it uses object recognition technology to evaluate the product category and condition, and for descriptions, it performs text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1344] Input: Product image, product description

[1345] Output: Analysis results of product images, analysis results of product descriptions

[1346] Step 4:

[1347] Sentiment-based analysis

[1348] Server: Based on the emotional information analyzed by the emotion engine, the system processes the user's emotional state. For example, if the user is nervous, the system generates feedback using a gentle tone.

[1349] Input: Emotion information

[1350] Output: Tone information for emotion-based feedback

[1351] Step 5:

[1352] Generate feedback

[1353] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and false statements. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[1354] Input: Analysis results of product images, analysis results of product descriptions, tone information of feedback based on emotions

[1355] Output: Buyer and seller feedback data

[1356] Step 6:

[1357] Providing feedback

[1358] Server: The generated feedback and rating results are sent to the user's device and provided to the seller and buyer. The seller is shown the generated product description and suggested starting price, and the buyer is shown the analysis results of the product image and description.

[1359] Input: Buyer and seller feedback data

[1360] Output: Feedback and evaluation results displayed on the user's device

[1361] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1362] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1363] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1364] [Fourth embodiment]

[1365] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1366] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1367] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1368] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1369] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1370] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1371] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1372] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1373] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1374] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1375] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1376] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1377] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1378] As an embodiment of the present invention, a system that provides convenience to both buyers and sellers in auction and flea market services is shown. The following describes the processing contents of a specific program.

[1379] Overall system configuration

[1380] The system includes a user's device, a server, and a generative AI model. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback.

[1381] Program processing explanation

[1382] Uploading data

[1383] On the device: The seller launches the app on their device, takes a picture of the product, and enters a brief description of the product. After confirming the information entered, they press the "Upload" button to send the data to the server.

[1384] Receiving and storing data

[1385] Server: The server receives the product images and descriptions sent by the user and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[1386] Image and description analysis

[1387] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1388] Generate analysis results

[1389] Server: The results of image and text analysis are integrated to generate feedback for buyers and sellers. For buyers, feedback is provided on the degree of agreement between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, appropriate product descriptions and pricing suggestions based on the product's value are presented.

[1390] Providing analysis results

[1391] Server: The generated analysis results are sent to the user's device and displayed in a specified format. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer, providing reference information to assist in making a purchasing decision.

[1392] Specific examples

[1393] Examples for sellers

[1394] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not tested."

[1395] Device: Press the upload button to send the data to the server.

[1396] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description: "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." A starting price of 3,000 yen is suggested.

[1397] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[1398] Specific examples for buyers

[1399] User: A buyer clicks on a product they're interested in and sees the product image and description: "High-performance smartphone, virtually no scratches."

[1400] Device: Press the analyze button to send the data to the server.

[1401] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[1402] Server: Generates a rating that reads, "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate."

[1403] Server: The evaluation results are provided to the buyer and displayed on the device.

[1404] This system allows buyers and sellers to enjoy a highly reliable trading environment, improving convenience for both parties.

[1405] The processing flow will be explained below.

[1406] Step 1:

[1407] User: The seller launches the app on their device, takes a picture of the product, and enters a simple description of the product. For example, they can enter a description such as "Old camera, operation not confirmed."

[1408] Step 2:

[1409] Terminal: Checks the entered product image and description, converts them into the required format, and displays the "Upload" button. When the user presses the "Upload" button, the data is sent to the server.

[1410] Step 3:

[1411] Server: Receives product images and descriptions and stores them in a database, which also stores related information about the received images and descriptions.

[1412] Step 4:

[1413] Server: The saved product images and descriptions are passed to the generative AI model and analysis begins. Specifically, two processes, image analysis and text analysis, are performed.

[1414] Step 5:

[1415] Server: The image analysis process involves extracting features from product images, such as identifying the camera model, year of manufacture, and external condition.

[1416] Step 6:

[1417] Server: The text analysis process analyzes the product description, extracts keywords, and checks for inconsistencies, false statements, and exaggerations.

[1418] Step 7:

[1419] Server: Integrates the analysis results and generates feedback for buyers and sellers. For buyers, it evaluates the consistency between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates an appropriate product description and suggests a starting price.

[1420] Step 8:

[1421] Server: The generated analysis results are sent to the user's device. Sellers are sent new product descriptions and pricing suggestions. Buyers are sent product image and description evaluation results.

[1422] Step 9:

[1423] Terminal: The received analysis results are displayed in a specified format and notified to the user. The seller is shown the generated product description and pricing suggestions. The buyer is shown any discrepancies and evaluation results, which are provided as reference information to assist in purchasing decisions.

[1424] This process allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information.

[1425] Example 1

[1426] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1427] Conventional auction and flea market services lack the means to verify the authenticity of product images and descriptions uploaded by users, resulting in problems such as fraud and inappropriate language. Furthermore, sellers often lack support for creating appropriate product descriptions and setting prices, resulting in transactions that do not proceed smoothly. The objective of this invention is to provide a system that analyzes product images and descriptions and provides evaluation results to solve these problems.

[1428] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1429] In this invention, the server includes a means for analyzing product images and product descriptions using a generative AI model, performing feature extraction and text analysis, a means for generating evaluation results based on the results of the feature extraction and text analysis, and a means for providing the evaluation results to the user. This ensures the reliability of the product images and descriptions, allowing users to conduct transactions with peace of mind. Furthermore, the server can also provide support for the appropriateness of product descriptions and pricing, which is expected to facilitate smooth transactions.

[1430] "User" means a person or organization that uses the System to upload product images and product descriptions.

[1431] "Product image" refers to an image file that shows visual information about a product and is uploaded to the system by a seller.

[1432] "Product description" is text information that describes the characteristics and condition of a product and is uploaded by the seller to the system.

[1433] "Server" refers to a set of hardware and software systems that receives, stores, and analyzes product images and product descriptions, generates evaluation results, and provides them to users.

[1434] A "generative AI model" is an artificial intelligence model used to analyze product images and product descriptions and perform feature extraction and text analysis.

[1435] "Feature extraction" is the process of using a generative AI model to extract feature information such as product category and condition from product images.

[1436] "Text analysis" is the process of using generative AI models to extract keywords from product descriptions and check for contextual inconsistencies, misrepresentations, and exaggerations.

[1437] The "evaluation results" are feedback information regarding the reliability of a product, appropriate descriptions, and pricing, created based on the results of feature extraction and text analysis.

[1438] The "means for providing to the user" refers to a series of processes and system configuration for transmitting the evaluation results generated by the server to the user's terminal and displaying them.

[1439] This invention is a system that provides convenience to buyers and sellers in auction and flea market services. The system includes a user terminal, a server, and a generative AI model, and the specific program processing content will be explained below.

[1440] Overall system configuration

[1441] The system includes a user's device, a server, and a generative AI model. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback.

[1442] Uploading data

[1443] As a seller, the user launches the terminal app, takes a picture of the product, and enters a brief description of the product. After the device checks the entered product image and description, the user presses the "Upload" button, which sends this data to the server.

[1444] As a specific example of operation, a user takes an image with an old camera, enters a simple description such as "old camera, operation not confirmed," and presses the "upload" button.

[1445] Receiving and storing data

[1446] The server receives product images and descriptions submitted by users, and stores the data in a database, ready for analysis by the generative AI model.

[1447] Image and description analysis

[1448] The server uses a generative AI model to analyze product images and descriptions. For product images, object recognition technology is used to extract features and evaluate the product's category and condition. For product descriptions, text analysis is performed to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1449] As a specific example of how it works, the generative AI model analyzes an image of an old camera, generates a description such as "A vintage camera from the 1940s. The shutter function has not been confirmed, but the appearance is good," and suggests a starting price of 3,000 yen.

[1450] Generate analysis results

[1451] The server combines the results of image and text analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and false statements. For sellers, it presents appropriate product descriptions and pricing suggestions based on the product's value.

[1452] As a specific example of how it works, when a buyer checks a product described as "a high-performance smartphone with almost no scratches" and presses the analysis button, the server generates an evaluation that reads, "There is no inconsistency between the description and the image. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate."

[1453] Providing analysis results

[1454] The server sends the generated analysis results to the user's device and displays them in a specified format. The seller is shown the generated product description and pricing suggestions. The buyer is shown the analysis results of the product image and description, which are provided as reference information to assist in purchasing decisions.

[1455] This system allows buyers and sellers to enjoy a highly reliable trading environment, improving convenience for both parties.

[1456] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1457] Step 1:

[1458] Uploading data

[1459] User: The seller launches the app on their device, takes a picture of the product, and enters a description of the product. The user then presses the "Upload" button to send the data.

[1460] Input: Product image and description

[1461] Output: A data packet of product images and descriptions sent to the server

[1462] Specific behavior:

[1463] The seller launches the terminal app.

[1464] The seller takes a picture of the product.

[1465] The seller enters the product description.

[1466] The seller presses the "Upload" button.

[1467] The device app sends product image and description data to the server.

[1468] Step 2:

[1469] Receiving and storing data

[1470] Server: The server receives the product images and product descriptions sent by the user and stores the received data in a database.

[1471] Input: Product image and description sent from the user's device

[1472] Output: Product images and descriptions stored in a database

[1473] Specific behavior:

[1474] The server receives the data from the terminal.

[1475] The server converts the received data into an appropriate format for analysis.

[1476] The server stores the converted data in a database.

[1477] Step 3:

[1478] Image and description analysis

[1479] Server: The server uses generative AI models to analyze product images and product descriptions. It performs object recognition on product images to assess product category and condition, and performs text analysis on product descriptions to extract keywords and check for inconsistencies, false statements, and exaggerations.

[1480] Input: Product images and descriptions stored in the database

[1481] Output: Analyzed feature information and text analysis results

[1482] Specific behavior:

[1483] The server inputs product images into a generative AI model.

[1484] The generative AI model performs object recognition on the image and extracts the product's features.

[1485] The server inputs the product description into the generative AI model.

[1486] The generative AI model analyzes the text, extracts keywords, and checks for contextual inconsistencies.

[1487] Step 4:

[1488] Generate analysis results

[1489] Server: The server combines the results of image analysis and text analysis to generate feedback for buyers and sellers.

[1490] For buyers, it provides evaluation results on the degree of agreement between the description and the image, as well as any inconsistencies and false statements, while for sellers, it presents suggestions for appropriate product descriptions and pricing.

[1491] Input: Image analysis and text analysis results

[1492] Output: Feedback for buyers and sellers

[1493] Specific behavior:

[1494] The server combines the results of image analysis and text analysis.

[1495] The server generates feedback for the buyer.

[1496] The server generates feedback for the seller.

[1497] Step 5:

[1498] Providing analysis results

[1499] Server: The server sends the generated analysis results to the user's device. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer.

[1500] Input: Feedback for buyers and sellers

[1501] Output: Feedback information displayed on the user's terminal

[1502] Specific behavior:

[1503] The server sends the feedback to the user's terminal.

[1504] The generated product description and pricing suggestions are displayed on the seller's device.

[1505] The analysis results of the product image and description are displayed on the buyer's device.

[1506] (Application example 1)

[1507] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1508] In modern auction and flea market services, many problems arise due to the lack of trust between buyers and sellers. This can make buyers unsure about the safety of their transactions, and sellers find it difficult to properly evaluate and price their products. Furthermore, there is a lack of means to detect inconsistencies and false representations between product images and descriptions, making it difficult to achieve transparent transactions.

[1509] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1510] In this invention, the server includes a means for users to upload product images and product descriptions, a means for the server to receive and store the product images and product descriptions, a means for analyzing the product images and product descriptions using a generative AI model to generate evaluation results and security check results, and a means for users to check the evaluation results and security check results. This makes it possible to check for inconsistencies and false statements between product images and descriptions, generate appropriate product descriptions, present pricing options, and provide security check results.

[1511] "User" refers to sellers who use the system to upload product images and product descriptions, and buyers who check the analysis results.

[1512] "Product images" are photo files of products that users sell.

[1513] "Product description" is text data that describes in sentence format information about the product that the user is selling.

[1514] A "server" is a computer system that receives product images and product descriptions sent by users, analyzes them using a generative AI model, and stores the results.

[1515] "Generative AI model" refers to an algorithm that uses machine learning technology to analyze product images and product descriptions and generate evaluation results.

[1516] "Analysis" is the process by which the generative AI model analyzes product images and product descriptions to evaluate the accuracy of the product's features and descriptions.

[1517] "Evaluation results" are information about the accuracy of the product's characteristics and description obtained through analysis by the generative AI model.

[1518] "Security check results" are information that summarizes the results of detecting fraudulent expressions or potentially misleading statements through analysis of product images and product descriptions.

[1519] "Verifiable means" refers to the interface and functionality that allows users to view evaluation results and security check results.

[1520] "Inconsistencies" refers to inconsistencies or inconsistencies between product images and product descriptions.

[1521] "Misrepresentation" means a product description that is untrue or misleading.

[1522] "Product description generation" is the process in which a generative AI model automatically creates a more specific and attractive description based on the original description.

[1523] "Pricing suggestions" are the estimated value and price suggestions for a product that the generative AI model presents to the user based on the analysis results.

[1524] As an embodiment of the present invention, a system that provides convenience and security to both buyers and sellers in auction and flea market services is shown. The system includes a user terminal, a server, and a generative AI model.

[1525] Overall system configuration

[1526] 1. On the user's device:

[1527] Sellers use a terminal to take product images, enter product descriptions, and upload them.

[1528] The buyer uses a terminal to view the listed items and check the analysis results and security check results.

[1529] 2. Server:

[1530] The server receives the product images and product descriptions sent by the user and stores them in a database.

[1531] A generative AI model is used to analyze product images and descriptions, and generate evaluation and security check results.

[1532] The generated results are sent to the user's terminal and displayed in a predetermined format.

[1533] 3. Generative AI Model:

[1534] For product images, features are extracted using object recognition technology to evaluate the product category and condition.

[1535] For product descriptions, text analysis technology is used to extract keywords and detect contradictions, false statements, and exaggerated expressions.

[1536] Program processing explanation

[1537] Uploading data

[1538] On the device: The seller launches the app on their device, takes a picture of the product, and enters a brief description of the product. After confirming the information entered, they press the "Upload" button to send the data to the server.

[1539] Receiving and storing data

[1540] Server: The server receives the product images and descriptions sent by the user and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[1541] Image and description analysis

[1542] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1543] Generate analysis results

[1544] Server: The results of image and text analysis are integrated to generate feedback for buyers and sellers. For buyers, feedback is provided on the degree of agreement between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, appropriate product descriptions and pricing suggestions based on the product's value are presented.

[1545] Providing analysis results

[1546] Server: The generated analysis results are sent to the user's device and displayed in a specified format. The generated product description and pricing suggestions are displayed to the seller. The analysis results of the product image and description are displayed to the buyer, providing reference information to assist in making a purchasing decision.

[1547] Specific examples

[1548] Examples for sellers

[1549] The seller takes a picture of the old camera and enters a simple description such as "Old camera, operation not confirmed." Presses the upload button on the device to send the data to the server. The server analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "Vintage camera from the 1940s, shutter operation not confirmed, but appearance is good." A suggested starting price of 3,000 yen is presented. The server provides the generated product description and suggested starting price to the seller, which are displayed on the device.

[1550] Specific examples for buyers

[1551] A buyer clicks on a product they are interested in and checks the product image and description: "High-performance smartphone, almost no scratches." They then press the analysis button on their device to send the data to the server. The server analyzes the product image and description and checks whether the actual appearance matches the description. The server generates a rating: "There is no discrepancy between the description and the image. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate." The server then provides the rating result to the buyer, which is displayed on the device.

[1552] Prompt Sentence Examples

[1553] Seller prompt:

[1554] Product description: Old camera, operation not confirmed

[1555] Please change this to a more specific and compelling description.

[1556] Buyer prompt:

[1557] Description: High-performance smartphone, almost no scratches

[1558] Please check for discrepancies between images and descriptions and provide feedback on the actual condition.

[1559] These features allow users to conduct transactions with high reliability and transparency.

[1560] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1561] Step 1:

[1562] The user uses a device to take a picture of the product and enter a description. The user starts the application on the device, takes a picture of the product, and enters a description. After checking the entered information, the user presses the "Upload" button.

[1563] Input: Product image file, product description

[1564] Output: Upload request

[1565] Step 2:

[1566] The device sends the product image taken by the user and the product description entered by the user to the server. During this process, the image file and text data are transferred to the server in an appropriate format.

[1567] Input: Upload request (product image file, product description)

[1568] Output: Data sent to the server

[1569] Step 3:

[1570] The server stores the product images and description received from the device in a database, and converts the received data into the required format for analysis.

[1571] Input: Data to be sent (product image, product description)

[1572] Output: Data saved to database

[1573] Step 4:

[1574] The server uses a generative AI model to analyze product images and product descriptions. For product images, object recognition technology (e.g., TensorFlow or PyTorch) is used to extract features and evaluate the product's category and condition. For product descriptions, text analysis technology (e.g., Hugging Face's T5 model) is used to extract keywords and detect contradictions, false statements, and exaggerations.

[1575] Input: Product images and product descriptions stored in the database

[1576] Output: Analysis results (verification results of product category, condition, description)

[1577] Step 5:

[1578] The server combines the results of image and text analysis to generate feedback for buyers and sellers. For buyers, it generates an evaluation result on the degree of match between the image and description, as well as any inconsistencies and false statements. For sellers, it presents appropriate product descriptions and pricing suggestions based on the value of the product.

[1579] Input: Analysis results (product category, condition, and description verification results)

[1580] Output: Feedback (for buyers and sellers)

[1581] Step 6:

[1582] The server sends the generated feedback to the user's terminal. The seller is provided with the generated product description and pricing suggestions, and the buyer is provided with the analysis results of the product image and description.

[1583] Input: Feedback (for buyers, for sellers)

[1584] Output: Feedback data to the user terminal

[1585] Step 7:

[1586] The terminal displays the feedback received from the server. The seller checks the generated product description and pricing suggestions, and the buyer views the analysis results of the product image and description.

[1587] Input: Feedback data sent from the server

[1588] Output: Displayed analysis results and suggestions

[1589] In this way, sellers and buyers can conduct transactions with high reliability and transparency.

[1590] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1591] The present invention is a system that provides convenience to both buyers and sellers in auction and flea market services, and in particular, by combining it with an emotion engine, it further improves the user experience. Below, we will explain the specific program processing content.

[1592] Overall system configuration

[1593] The system includes a user's device, a server, a generative AI model, and an emotion engine. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[1594] Program processing explanation

[1595] Uploading data

[1596] On-device: The seller launches the app on their device, takes a photo of the product, and enters a simple description. For example, they can enter a description such as "old camera, operation not confirmed." In addition, the emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[1597] Receiving and storing data

[1598] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[1599] Image and description analysis

[1600] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1601] Sentiment-based analysis

[1602] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[1603] Generate feedback

[1604] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[1605] Providing analysis results

[1606] Server: The generated analysis results are sent to the user's device and displayed in a specified format. When the generated product description and pricing suggestions are displayed to the seller, they are displayed in a tone based on the analysis results of the emotion engine. The analysis results of the product image and description are displayed to the buyer, providing reference information to support their purchasing decision.

[1607] Specific examples

[1608] Examples for sellers

[1609] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not working." The emotion engine recognizes that the user is relaxed.

[1610] Terminal: Sends data to the server.

[1611] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." Feedback is provided in a relaxed tone based on the emotional information.

[1612] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[1613] Specific examples for buyers

[1614] User: A buyer clicks on a product they're interested in, sees the product image and description, "High-performance smartphone, virtually flawless." The emotion engine recognizes the user's doubts.

[1615] Device: Press the analyze button to send the data to the server.

[1616] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[1617] Server: Generates a rating like "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate." Based on sentiment, additional information is provided to resolve any doubts.

[1618] Server: The evaluation results are provided to the buyer and displayed on the device.

[1619] This system allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information and receiving feedback that takes into account their emotions.

[1620] The processing flow will be explained below.

[1621] Step 1:

[1622] User: The seller launches the app on their device, takes a picture of the product, and writes a brief description of the product. The emotion engine then analyzes the seller's emotional state (e.g., nervous, relaxed, anxious, etc.).

[1623] Step 2:

[1624] Terminal: Checks the entered product image, description, and emotion information, converts them into the required format, and displays the "Upload" button. When the user presses the "Upload" button, the data is sent to the server.

[1625] Step 3:

[1626] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[1627] Step 4:

[1628] Server: The saved product images and descriptions are passed to the generative AI model and analysis begins. Specifically, two processes, image analysis and text analysis, are performed.

[1629] Step 5:

[1630] Server: The image analysis process involves extracting features from product images, such as identifying the camera model, year of manufacture, and external condition.

[1631] Step 6:

[1632] Server: The text analysis process analyzes the product description, extracts keywords, and checks for inconsistencies, misrepresentations, and exaggerations.

[1633] Step 7:

[1634] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[1635] Step 8:

[1636] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[1637] Step 9:

[1638] Server: The generated analysis results are sent to the user's device. Sellers are sent new product descriptions and pricing suggestions. Buyers are sent product image and description evaluation results.

[1639] Step 10:

[1640] Terminal: The received analysis results are displayed in a specified format and notified to the user. For sellers, the generated product description and pricing suggestions are displayed in a tone based on the emotion engine's analysis results. For buyers, discrepancies and evaluation results are displayed, providing reference information to assist in purchasing decisions.

[1641] This process allows sellers to efficiently provide accurate product descriptions and set appropriate prices, and allows buyers to make purchasing decisions based on reliable information and emotionally sensitive feedback.

[1642] Example 2

[1643] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1644] Traditional auction and flea market services have faced challenges in terms of the reliability of product descriptions and improving the experience for buyers and sellers. In particular, there were insufficient means to check the consistency between product descriptions and images, and whether or not there were false or exaggerated statements, making it difficult for buyers to make decisions based on these. Furthermore, the impersonal feedback provided meant that services were not provided that took user feelings into consideration.

[1645] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1646] In this invention, the server includes a means for a user to upload product images and product descriptions, a means for the server to receive and store the product images and product descriptions, a means for analyzing the product images and product descriptions using a generative AI model to generate evaluation results, and a means for an emotion engine to analyze the user's emotional information and generate appropriate feedback by adjusting the tone and approach based on the emotional information. This enables the provision of highly reliable product evaluations as well as feedback that takes the user's emotions into consideration.

[1647] "User" refers to an individual or corporation that intends to sell or purchase items using the auction and flea market services.

[1648] "Terminal" refers to the electronic device (smartphone, PC, tablet, etc.) used by the user to input product images and product descriptions and send them to the system.

[1649] "Server" refers to a central management system that receives, stores, and analyzes data sent from user devices, generates evaluation results, and provides them to users.

[1650] "Generative AI model" refers to the artificial intelligence model used by the server to analyze product images and product descriptions and generate evaluation results.

[1651] "Product image" refers to a photograph or image data of the product that a user is listing for sale.

[1652] "Product description" refers to text data that explains the condition, features, price, etc. of the product that a user is selling.

[1653] An "emotion engine" is a system that analyzes emotional information from a user's facial expressions and voice, and provides appropriate feedback based on the results.

[1654] "Feedback" refers to comments, advice, notifications, etc. provided to users based on the evaluation results generated by the server.

[1655] The present invention is a system that provides convenience to both buyers and sellers in auction and flea market services, and in particular improves the user experience by combining an emotion engine. Specific embodiments are described below.

[1656] The system includes a user device, a server, a generative AI model, and an emotion engine. The user device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[1657] First, the seller launches the device app, takes a picture of the product, and enters a simple description. For example, the seller might enter a description such as "old camera, operation not confirmed." The device temporarily saves the captured image and the entered description in local storage. Next, the emotion engine uses the device's camera and microphone to analyze the seller's emotions.

[1658] Next, the device sends the product image, product description, and emotion information to the server. The server receives this data and stores it in a database. The product image is stored as image data, the description as text data, and the emotion information as structured data.

[1659] The server uses a generative AI model to analyze product images. Specifically, it uses object recognition technology to identify the product category and evaluate the product's condition. The generative AI model also performs text analysis of the product description to extract important keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1660] The emotion engine then generates feedback for sellers and buyers based on the emotional information analyzed. For sellers, it presents pricing suggestions based on appropriate product descriptions and the value of the product. For example, the server generates a description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good," and provides it to the seller. For buyers, it provides feedback on the degree of match between the product image and description, as well as any inconsistencies and false statements.

[1661] Furthermore, the tone and content of the feedback is adjusted based on the results of the emotion engine. For example, if the buyer is skeptical, detailed feedback such as "The description and images are consistent. The phone looks good, but there are some small scratches in the photos. The description is generally accurate."

[1662] Finally, the analysis results and feedback generated by the server are sent to the user's device and displayed in an appropriate format. The seller can review the generated product description and pricing suggestions and make any necessary corrections. The buyer can make a purchase decision based on the analysis results provided and with reliable information.

[1663] For example, if a seller takes a picture of an old camera and enters the description "Old camera, not working," the emotion engine recognizes that the seller is relaxed. The server analyzes the image and text, and the generative AI model generates an appropriate description, such as "Vintage camera from the 1940s, shutter not working, but looks good," and provides feedback in a relaxed tone.

[1664] If a buyer clicks on a smartphone product and sees the product image and the description "High-performance smartphone, almost no scratches," the emotion engine recognizes the buyer's doubts. The server analyzes the image and text and generates a rating result: "The description and the image are consistent. The smartphone looks good, but there are small scratches in some of the photos. The description is overall accurate," providing detailed and reliable feedback.

[1665] This system allows both users and sellers to obtain reliable information, sellers to efficiently generate appropriate product descriptions, and buyers to make purchasing decisions by receiving feedback that takes into account their emotions.

[1666] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1667] Step 1:

[1668] User: The seller launches the app on their device and takes a picture of the product. For example, they take a picture of an "old camera from the 1940s."

[1669] Input: The captured image.

[1670] Specific operation: The device app uses the camera function to temporarily save the captured image to local storage.

[1671] Output: Product images saved to local storage.

[1672] Step 2:

[1673] User: The seller enters the product description into the terminal. For example, they might enter "old camera, not working."

[1674] Input: The entered product description.

[1675] Specific operation: The terminal app temporarily saves the entered text data in local storage.

[1676] Output: Product description saved to local storage.

[1677] Step 3:

[1678] On the device: The emotion engine uses the device's camera and microphone to analyze the seller's facial expressions and voice.

[1679] Input: Seller's facial expression and voice data.

[1680] How it works: The sentiment engine uses machine learning algorithms to analyze the emotional state of sellers in real time.

[1681] Output: Seller emotional state data.

[1682] Step 4:

[1683] Device: Sends product images, product descriptions, and emotion information to the server.

[1684] Input: Product image, product description, sentiment information.

[1685] Specific operation: The device uses the network and sends the data as a data packet to the server.

[1686] Output: The server receives the product image, description, and emotion information as a data packet.

[1687] Step 5:

[1688] Server: Stores the received product images, product descriptions, and emotion information in a database.

[1689] Input: Received data (product image, product description, emotional information).

[1690] Specific operation: The server converts the data into a data format and stores it in a database. Product images are saved as image data, descriptions as text data, and emotional information as structured data.

[1691] Output: Product images, product descriptions, and sentiment information stored in a database.

[1692] Step 6:

[1693] Server: Analyzes product images using a generative AI model and identifies the product category using object recognition technology.

[1694] Input: Product images stored in the database.

[1695] How it works: The generative AI model uses object recognition algorithms to extract features in an image and identify the product category.

[1696] Output: Identified product category and product condition.

[1697] Step 7:

[1698] Server: Uses a generative AI model to analyze the text of product descriptions, extracting important keywords and checking for contextual inconsistencies, misrepresentations, and exaggerations.

[1699] Input: Product description stored in the database.

[1700] How it works: The generative AI model uses natural language processing techniques to analyze text, extract keywords, analyze context, and check for misrepresentations.

[1701] Output: Keyword extraction results, contextual inconsistencies, misrepresentations, and exaggerations.

[1702] Step 8:

[1703] Server: Generates feedback for sellers and buyers based on the emotional information analyzed by the emotion engine.

[1704] Input: Emotion information, product image analysis results, product description analysis results.

[1705] Specific operation: The server takes into account emotional information and generates appropriate product description and pricing suggestions for the seller, and generates feedback for the buyer regarding the degree of match between the description and the image, any inconsistencies, and whether there are any misrepresentations.

[1706] Output: Feedback for seller (product description, suggested pricing), feedback for buyer (match between image and description, inconsistencies, misrepresentations).

[1707] Step 9:

[1708] Server: Sends the generated feedback to the user's device.

[1709] Input: The generated feedback.

[1710] Specific operation: The server sends feedback data to the terminal via the network.

[1711] Output: Feedback data received by the user terminal.

[1712] Step 10:

[1713] Terminal: Displays the received feedback data in an appropriate format.

[1714] Input: The received feedback data.

[1715] Specific operation: The device analyzes the feedback data and displays it on the user interface. The seller is shown the generated product description and pricing suggestions, and the buyer is shown the analysis results.

[1716] Output: Feedback displayed on the device.

[1717] (Application example 2)

[1718] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1719] In conventional auction and flea market services, if the product descriptions and images provided by sellers contain inconsistencies or falsehoods, buyers' trust is often damaged. It is also difficult for sellers to set appropriate prices and product descriptions, and there is a lack of advice and feedback to stimulate purchasing motivation. Another problem is the lack of a system that provides feedback based on the user's emotional state. Given this background, there is a need for a system that can provide accurate and reliable information while taking user emotions into consideration.

[1720] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1721] In this invention, the server includes: means for a user to upload product images and product descriptions; means for the server to receive and store the product images and product descriptions; means for the server to analyze the product images and product descriptions using a generative AI model and generate evaluation results; means for the server to provide the evaluation results to the user; means for an emotion engine to analyze the user's emotions; and means for generating and providing feedback to the user based on the analysis results and emotion information. This makes it possible to provide highly reliable feedback that takes the user's emotions into consideration.

[1722] "User emotion" refers to the user's psychological state and emotions as read and analyzed by the emotion engine.

[1723] "Emotion engine" is a general term for software and algorithms used to analyze emotions based on data such as a user's facial expressions and voice.

[1724] A "generative AI model" is an artificial intelligence model that analyzes product images and descriptions and generates the necessary feedback.

[1725] "Product images" refer to photographs and image data of the products being offered for sale.

[1726] A "product description" is a text description of the characteristics and condition of the product being offered for sale.

[1727] "Analysis results" refer to the evaluations and feedback generated by generative AI models and emotion engines.

[1728] "Feedback" refers to advice and evaluation information provided to the user based on the analysis results and emotional information.

[1729] A "server" is a centralized processing unit for receiving, storing, analyzing, and providing feedback to the user of data.

[1730] "Upload" refers to the act of a user sending data from their own terminal to a server.

[1731] "Receiving" means that the server takes in data such as product images and product descriptions sent from the user's terminal.

[1732] "Storing" means storing the received data in a storage device such as a database.

[1733] "Analysis" refers to using generative AI models and emotion engines to process product images and descriptions to generate ratings and feedback.

[1734] "Inconsistencies" refer to discrepancies or inconsistencies between product images and descriptions.

[1735] "False representation" refers to information contained in a product description that is incorrect and different from the facts.

[1736] "Exaggeration" refers to descriptions that exaggerate the actual characteristics of a product.

[1737] The present invention is a system that includes a means for a user to upload product images and product descriptions, a means for a server to receive and store the product images and product descriptions, a means for the server to analyze the product images and product descriptions using a generative AI model and generate an evaluation result, a means for the server to provide the evaluation result to the user, a means for an emotion engine to analyze the user's emotions, and a means for generating feedback based on the analysis result and emotion information and providing it to the user.

[1738] Overall system configuration

[1739] The system includes a user's device, a server, a generative AI model, and an emotion engine. The user's device is used by sellers to upload product images and descriptions. It is also used by buyers to review listed products and obtain evaluation results. The server receives, stores, and analyzes the data, generates evaluation results, and provides them to users. The generative AI model analyzes product images and descriptions to generate the necessary feedback. The emotion engine recognizes the user's emotions and provides appropriate feedback according to their state.

[1740] Program processing explanation

[1741] Uploading data

[1742] On-device: The seller launches the app on their device, takes a photo of the product, and enters a simple description. For example, they can enter a description such as "old camera, operation not confirmed." In addition, the emotion engine analyzes the user's facial expressions and voice to determine their emotions.

[1743] Receiving and storing data

[1744] Server: Receives product images, product descriptions, and user emotion information analyzed by the emotion engine, and stores them in a database. The received data is prepared for analysis by the generative AI model and stored in an appropriate format.

[1745] Image and description analysis

[1746] Server: The generative AI model analyzes product images and descriptions. The product images are subjected to feature extraction using object recognition technology to evaluate the product's category and condition. The product descriptions are subjected to text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1747] Sentiment-based analysis

[1748] Server: Based on the emotional information analyzed by the emotion engine, processing is performed taking into account the user's emotional state. For example, if the user is nervous, the system will provide feedback using a gentle tone.

[1749] Generate feedback

[1750] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and misrepresentations. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[1751] Hardware and software used

[1752] Hardware: The smartphone used by the user

[1753] Software: Python, OpenCV, Transformers library, EmotionRecognition library

[1754] Specific examples

[1755] Examples for sellers

[1756] User: The seller takes a picture of an old camera and writes a simple description: "Old camera, not working." The emotion engine recognizes that the user is relaxed.

[1757] Terminal: Sends data to the server.

[1758] Server: Analyzes the received product image and description, identifies the camera model, and generates an appropriate product description such as "A vintage camera from the 1940s. The shutter is not working, but the appearance is good." Feedback is provided in a relaxed tone based on the emotional information.

[1759] Server: The seller is provided with a generated product description and suggested starting price, which are displayed on the device.

[1760] Specific examples for buyers

[1761] User: A buyer clicks on a product they're interested in, sees the product image and description, "High-performance smartphone, virtually flawless." The emotion engine recognizes the user's doubts.

[1762] Device: Press the analyze button to send the data to the server.

[1763] Server: Analyzes product images and descriptions to check whether the actual appearance matches the description.

[1764] Server: Generates a rating like "The description and images are consistent. The phone looks good, but there are small scratches in some of the photos. The description is generally accurate." Based on sentiment, additional information is provided to resolve any doubts.

[1765] Server: The evaluation results are provided to the buyer and displayed on the device.

[1766] Prompt Sentence Examples

[1767] Description of the invention:

[1768] I would like to develop a system to improve the user experience of auction and flea market services. This system combines an emotion engine and a generative AI model to analyze product images and descriptions and provide appropriate feedback to both buyers and sellers. It also includes a function to check for inconsistencies and false statements in product descriptions and suggest optimal pricing to sellers.

[1769] Prerequisites:

[1770] Sellers upload product images and descriptions.

[1771] Analyzing user emotions with an emotion engine

[1772] Analyze and generate product descriptions using a generative AI model

[1773] Give feedback in an emotional tone

[1774] Expected output:

[1775] Relaxed tone of voice and optimal product description and pricing suggestions for sellers

[1776] Feedback to buyers regarding the consistency of product descriptions and images, as well as any inconsistencies

[1777] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1778] Step 1:

[1779] Uploading data

[1780] Device: The seller uses a smartphone to take a picture of the product and enter a brief description. For example, they might enter "old camera, not yet operational." The emotion engine then activates and analyzes the seller's facial expressions and voice to obtain emotional information.

[1781] Input: Product images, product descriptions, seller's facial expressions and voice data

[1782] Output: Product images, product descriptions, emotional information

[1783] Step 2:

[1784] Receiving and storing data

[1785] Server: Receives product images, product descriptions, and emotion information sent from the device. The received data is stored in a database using transaction processing to prepare for analysis.

[1786] Input: Product image, product description, emotional information

[1787] Output: Product images, product descriptions, and emotional information stored in a database

[1788] Step 3:

[1789] Image and description analysis

[1790] Server: Using a generative AI model, it analyzes product images and descriptions. For product images, it uses object recognition technology to evaluate the product category and condition, and for descriptions, it performs text analysis to extract keywords and check for contextual inconsistencies, false statements, and exaggerations.

[1791] Input: Product image, product description

[1792] Output: Analysis results of product images, analysis results of product descriptions

[1793] Step 4:

[1794] Sentiment-based analysis

[1795] Server: Based on the emotional information analyzed by the emotion engine, the system processes the user's emotional state. For example, if the user is nervous, the system generates feedback using a gentle tone.

[1796] Input: Emotion information

[1797] Output: Tone information for emotion-based feedback

[1798] Step 5:

[1799] Generate feedback

[1800] Server: Integrates the results of image analysis, text analysis, and sentiment analysis to generate feedback for buyers and sellers. For buyers, it provides evaluation results on the degree of match between the description and the image, as well as any inconsistencies and false statements. For sellers, it generates appropriate product descriptions and suggests pricing based on the product's value. It also provides feedback in an appropriate tone and approach based on the results of the sentiment engine.

[1801] Input: Analysis results of product images, analysis results of product descriptions, tone information of feedback based on emotions

[1802] Output: Buyer and seller feedback data

[1803] Step 6:

[1804] Providing feedback

[1805] Server: The generated feedback and rating results are sent to the user's device and provided to the seller and buyer. The seller is shown the generated product description and suggested starting price, and the buyer is shown the analysis results of the product image and description.

[1806] Input: Buyer and seller feedback data

[1807] Output: Feedback and evaluation results displayed on the user's device

[1808] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1809] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1810] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1811] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1812] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1813] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1814] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1815] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1816] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1817] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1818] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1819] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1820] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1821] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1822] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1823] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1824] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1825] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1826] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1827] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1828] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1829] The following is further disclosed regarding the above embodiment.

[1830] (Claim 1)

[1831] A means for users to upload product images and product descriptions;

[1832] A server receives and stores the prod...

Claims

1. A means for users to upload product images and product descriptions; A server receives and stores the product images and product descriptions; A server uses a generated AI model to analyze the product image and product description and generate an evaluation result; a means for the server to provide the evaluation results to a user; A system including:

2. The system according to claim 1, further comprising means for checking the product images and product descriptions for inconsistencies, false statements, and exaggerations for a purchaser user.

3. The system according to claim 1, further comprising: a means for generating an appropriate product description for seller users based on the product image and product description; and a means for evaluating the value of the product and presenting pricing suggestions.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A