system
The system simplifies the listing of items on online markets by using a camera to identify objects, calculate market prices, and provide voice recommendations, addressing the challenges of cumbersome listing processes and promoting sustainability.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Many individuals miss opportunities to sell their items on online markets due to a lack of understanding of the selling market and a cumbersome listing process, leading to items being left at home or discarded without understanding their value, which affects space utilization and sustainability.
A system that uses a camera to identify objects, references a database for past sales history to calculate market prices, and provides visual and voice recommendations for listing items on online markets, automating the listing process through a generation device.
Enables users to easily identify and list items for sale quickly and efficiently, maximizing space utilization and promoting sustainability by streamlining the listing process and reducing effort.
Smart Images

Figure 2026069027000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, many people have missed the opportunity to sell their items on online markets while still holding them at home. The main reasons are that they do not understand the selling market or the listing process is cumbersome. Due to such problems, many items are left at home, preventing the effective use of space. Also, many cases occur where items are discarded without understanding their value, and improvement is required from the perspective of sustainability. The present invention aims to enable a user to easily identify the items they own, quickly and accurately grasp their selling value, and simplify the process of listing them on online markets.
Means for Solving the Problems
[0005] To solve this problem, the present invention provides a means for analyzing video data acquired using a camera and automatically identifying objects. Furthermore, it provides a means for referencing a database of past sales history based on the information of the identified object and calculating the market price for selling it. This price information is visually presented on a display device, and in addition, the listing of the item is suggested by voice. This allows the user to intuitively understand the selling value and encourages them to make a decision to list the item. Moreover, after obtaining the user's intention to list the item, a means is provided to automatically create a product description and image caption using a generation device and send the listing information to an online market application, thereby streamlining the entire listing process and reducing the effort involved.
[0006] "Object" refers to a specific object that is recognized and identified through a photographic device.
[0007] "Image capture device" refers to a camera or similar device used to acquire image data of an object.
[0008] "Video data" refers to image information acquired by a camera or camera and used for object recognition and analysis.
[0009] A "server" refers to a computer system that receives video data and performs tasks such as object identification and calculating market value.
[0010] "Identification AI" refers to artificial intelligence algorithms that analyze video data to identify objects.
[0011] "Market selling price" refers to monetary information that indicates the current market value of an object, calculated based on past sales history.
[0012] A "display device" refers to a visual output device that provides users with information such as market prices for selling items and recommendations for listing items.
[0013] "Voice recommendation" refers to audio suggestions made to encourage users to list items for sale.
[0014] The "generation device" refers to a device or system that automatically generates product descriptions and image captions to create product information.
[0015] The "online market application" refers to an e-commerce trading platform for users to buy and sell objects.
Brief Description of Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] The system of this invention enables efficient listing on online marketplaces by allowing users to easily identify objects they own and quickly determine their market value. This system consists of a camera, a terminal, a server, and an online marketplace application.
[0038] When a user uses the system, they first use a camera associated with their device to photograph the object they intend to sell. This camera is typically a built-in camera in a smartphone or AR glasses. The device then acquires image data of the object and transmits it to a server via the internet.
[0039] The server analyzes the received video data. Using identification AI, it identifies objects from the video and extracts their characteristics. This includes the object's appearance, shape, and color. Based on the extracted information, the server queries a database containing past sales history to calculate the current market price for selling the object.
[0040] The device visually displays market price information transmitted from the server on its display device. This display can be the screen of AR glasses or a smartphone. The device also uses voice recommendation functionality to suggest that the user list an item on an online marketplace. This voice assistant might make suggestions such as, "You can sell this item for approximately 5,000 yen. Why not list it?"
[0041] If a user wishes to list an item for sale following the terminal's suggestions, they can indicate their intention to list the item through the terminal's interface. In this case, the server automatically generates listing information for the item using a generator. Specifically, a product description, title, and accompanying image captions are generated. This data is then compiled and sent to the online marketplace application.
[0042] Finally, the server completes the listing process on the online marketplace application. This entire process allows users to sell their items quickly and easily. This system promotes the effective use of users' goods and enhances convenience and sustainability.
[0043] For example, if a user takes a picture of an old smartphone at home, the system will identify the smartphone and indicate its current market value of approximately 15,000 yen. If the user then wishes to list it for sale, the system will quickly complete the listing using automatically generated information. This allows users to maximize their profits while saving time and effort.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] The user points the object they want to sell at the camera using a camera device associated with the terminal. The terminal acquires video data of this object in real time.
[0047] Step 2:
[0048] The terminal efficiently compresses the video data it acquires and sends it to the server via the internet. The server stores the received data.
[0049] Step 3:
[0050] The server analyzes the received video data and uses identification AI to identify objects. The server then extracts the object's characteristics (e.g., shape, color, brand, etc.).
[0051] Step 4:
[0052] The server uses the extracted object information to access a database that holds past trading history and calculates the market selling price of the identified object. In this process, the server utilizes statistical analysis and machine learning techniques.
[0053] Step 5:
[0054] The server sends the calculated market price information to the terminal. The terminal receives this information and presents it visually to the user through a display device.
[0055] Step 6:
[0056] The device uses its voice output function to offer the user a voice message suggesting they list the item for sale, such as, "This item can currently be sold for approximately 15,000 yen. Would you like to list it?"
[0057] Step 7:
[0058] If a user wishes to list an item for sale, they express their intention via voice command or touch input through their device. The device then notifies the server of this intention.
[0059] Step 8:
[0060] The server uses a generator to automatically produce product descriptions and image captions, thus consolidating the listing information. This information is then prepared for use in online marketplace applications.
[0061] Step 9:
[0062] The server sends the listing information it has configured to the online marketplace application via a communication device, completing the listing process.
[0063] Step 10:
[0064] The device notifies the user that the listing is complete and reports to the user that the listing was successfully completed.
[0065] (Example 1)
[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0067] In today's information society, listing personal belongings on the market quickly and at a fair price remains a complex and time-consuming process. Furthermore, identifying items and calculating market prices requires specialized knowledge and skills, making it difficult for the average consumer to use. Therefore, there is a challenge in how to simplify and efficiently carry out these processes.
[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0069] In this invention, the server includes means for analyzing video data acquired by a shooting means and automatically identifying a specific object using image processing technology; means for calculating the transaction value of the identified object by referring to a recording medium containing past transaction history based on the visual characteristics of the identified object; and means for presenting the calculated value information through a display means and recommending the transaction of the object by voice. As a result, users can easily and quickly list items on the market and trade them at a fair price without having specialized knowledge.
[0070] "Means of shooting" refers to a device or function for acquiring image data of an object, and includes devices such as cameras for taking images or videos.
[0071] "Image processing technology" refers to methods for analyzing acquired video data, and involves the use of algorithms and artificial intelligence technologies to identify objects within images.
[0072] An "object" is an object that a user intends to sell, which is identified by the system and treated as the subject of a transaction.
[0073] A "recording medium" is a data storage system used to hold information such as past transaction history, and includes databases and digital storage devices.
[0074] "Value" refers to the market transaction price calculated based on the identified object, and is estimated based on past transaction data.
[0075] "Display means" refers to a device used to visually communicate calculated value or information to the user, and includes devices such as displays and screens.
[0076] "Voice-based recommendation methods" refer to devices or systems that have the function of using voice to suggest to users how to encourage them to trade an item.
[0077] "Communication means" refers to devices or systems for transferring information over a network, including protocols for internet connectivity and data communication.
[0078] An "e-commerce platform" is a software infrastructure that enables the online trading of goods and is a system for conducting buying and selling transactions.
[0079] This invention is a system for efficiently listing items owned by users on the market, and it functions through the coordinated operation of multiple devices and software. The user photographs the items they wish to sell using a photographic means. This photographic means uses a camera built into a smartphone or augmented reality glasses. The captured video data is acquired by a terminal, which processes the image and transfers it to a server.
[0080] The server analyzes the acquired video data using image processing technology. Here, an identification AI model is used to identify objects and extract their features. The AI models used are created using TENSORFLOW® and PyTorch. The extracted feature information is then compared with past transaction history stored on the recording medium. Based on this, the server calculates the market value of the object.
[0081] The calculated value information is transmitted to the terminal. The terminal uses a smartphone screen or augmented reality glasses display to provide the user with visual information. It also encourages the user to trade the item through voice prompts. For example, it might prompt, "You can sell this item for about 5,000 yen. Why not list it?"
[0082] When a user requests to list an item, the server automatically generates product descriptions and image annotations using a generative AI model, and then constructs the transaction information. This generated transaction information is then transferred to the e-commerce platform via communication. Finally, the server completes the listing process, allowing the user to easily and quickly list their items on the marketplace.
[0083] For example, if a user holds an old smartphone from their home up to the system and takes a picture, the system analyzes the captured data and indicates a market value of approximately 15,000 yen. If the user wishes to list the item for sale, the system automatically generates listing information, enabling a quick transaction. Through this process, users can make optimal transactions while saving time and effort.
[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0085] Step 1:
[0086] The user photographs the items they are considering selling using a photographic device. Specifically, they use the camera built into their smartphone or augmented reality glasses to acquire video data of the objects. This video data is then used for processing within the device.
[0087] Step 2:
[0088] The terminal formats the acquired video data into a predetermined format and sends it to the server via the internet. At this stage, the output is video data formatted in a format that the server can receive.
[0089] Step 3:
[0090] The server inputs the received video data into an identification AI model. It performs image processing for video analysis, extracting features such as the appearance and shape of objects. Through this process, the server identifies what the objects are and obtains feature information of the identified objects as output.
[0091] Step 4:
[0092] The server uses the extracted characteristic information to refer to past transaction history stored on the recording medium. Using database queries, it extracts historical data on transactions of similar objects and calculates the market value based on this data. The output is market price information representing the current market value of the object.
[0093] Step 5:
[0094] The server sends the calculated market price information to the terminal. The terminal visually displays the output market price on the smartphone screen or the display of augmented reality glasses. This allows the user to check the current market price of an object.
[0095] Step 6:
[0096] The device uses a voice assistant to suggest selling an item to the user. Specifically, it communicates the suggestion by voice and prompts the user to take the next action. The output of this step is that a valid selling suggestion is made to the user.
[0097] Step 7:
[0098] If a user wishes to sell based on the terminal's suggestions, they input their intention into the terminal. The terminal transmits this information to the server. Based on this input, the server uses a generative AI model to automatically generate product descriptions and image annotations, creating the listing information. This output contains all the listing information necessary for the transaction.
[0099] Step 8:
[0100] The server transmits the generated listing information to the e-commerce platform via communication means, completing the listing process. This final output means the item is listed on the marketplace and ready for immediate trading.
[0101] (Application Example 1)
[0102] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0103] Current online trading platforms present challenges in quickly and easily listing individual items, requiring users to manually input vast amounts of information and necessitating time-consuming market price searches. As a result, users often find the trading process burdensome, making it difficult to sell items quickly at appropriate market prices. Furthermore, the lack of mechanisms to ensure users can list items at fair prices can delay the listing process, potentially leading to missed sales opportunities.
[0104] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0105] In this invention, the server includes means for analyzing video data acquired by a camera to recognize an object and automatically identifying a specific object; means for calculating the market selling price of the identified object by referring to accumulated information of past transaction history based on the information of the identified object; and means for transmitting the generated listing information to an online trading platform via a communication device and completing the procedure for selling the object. This enables users to quickly grasp market prices and immediately list objects on the online platform at a fair price.
[0106] An "object" is something that the user intends to sell and that is recognized by the camera.
[0107] A "camera" is hardware used to acquire video data, and is mainly installed in mobile devices and other devices with camera functions.
[0108] "Video data" refers to visual information acquired by a camera or camera, and is fundamental information for identifying objects.
[0109] "Automatic identification means" refers to a combination of software or hardware that uses acquired video data to perform a process for identifying a specific object.
[0110] "Market selling price" is an estimate of the current market price of an object, calculated based on past transaction history.
[0111] A "generation device" refers to a software or hardware configuration for automatically generating listing information.
[0112] A "visual display device" is a digital screen or display used to present information to a user visually.
[0113] A "communication device" is a device used to send and receive data via a network connection.
[0114] An "online trading platform" refers to a website or application where goods and services are bought and sold via the internet.
[0115] A "generative AI model" is an artificial intelligence technology used to automatically generate text and information.
[0116] The system of this invention provides an efficient means for users to quickly put their owned objects up for market. The system mainly consists of a camera, a terminal, a server, and an online trading platform.
[0117] First, the user acquires video data of the object they wish to sell using the camera function of their portable information terminal. The camera function of a smartphone or tablet is used as the camera. This video data is then transmitted to a server via the internet through the terminal.
[0118] The server analyzes the received video data using identification AI to automatically identify specific objects. Information about the identified objects is then referenced in a database containing past transaction history to calculate the market price for those objects. This price information is transmitted to the terminal and displayed on the user's device. A speech synthesis system is also utilized, allowing users to receive voice-based listing suggestions.
[0119] When a user indicates their willingness to follow voice suggestions during the listing process, the server generates listing information using a generation device. This generation AI model automatically generates product descriptions and visual annotations. The generated listing information is then transmitted to the online trading platform via a communication device, completing the sale process for the item.
[0120] As a concrete example, let's assume a user takes a picture of a used laptop computer they have on hand using this system. The system recognizes the object and suggests its market price is, for example, 20,000 yen. If the user wishes to list it for sale, the system generates a product description and caption, and quickly lists it on an online trading platform. In this process, an example of a prompt message used with the generating AI model would be: "This is a used laptop computer, 3 years old and in good condition. We recommend listing it for 20,000 yen."
[0121] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0122] Step 1:
[0123] The user uses the camera on their portable information terminal to acquire video data of the object they wish to sell. Here, the input is the video data of the object, and the output is the saving of the video data to the terminal. This video data is then prepared to be sent to a server for subsequent processing.
[0124] Step 2:
[0125] The terminal transmits video data to the server. The input is the acquired video data, and the output is the transmission of the image to the server. Specifically, the terminal transfers data to the server using a network connection.
[0126] Step 3:
[0127] The server analyzes the received video data and automatically identifies specific objects using identification AI. The input is the transmitted video data, and the output is information about the identified objects. In this process, the server extracts and recognizes features such as the appearance, shape, and color of the objects.
[0128] Step 4:
[0129] The server uses identified object information to refer to a database and calculate the market selling price. The input is information about the identified object, and the output is the calculated market price. The server calculates the price based on past transaction history.
[0130] Step 5:
[0131] The server sends the calculated market price to the user's terminal. The input is the calculated market price, and the output is the transmission of price information to the terminal. This allows the user to view the price information on a visual display device.
[0132] Step 6:
[0133] The terminal displays the received market price on a visual display and uses a speech synthesis system to suggest listing the item. The input is price information sent from the server, and the output is the displayed price and a voice suggestion. The operation includes a process in which the voice assistant suggests, "Why don't you list this item for sale?"
[0134] Step 7:
[0135] When a user enters their intention to list an item into the terminal, the server generates listing information using a generator. The input is the user's intention to list an item, and the output is the generated listing information. This information includes a product description and image captions.
[0136] Step 8:
[0137] The server transmits the generated listing information to the online trading platform via a communication device. The input is the generated listing information, and the output is the transmission of information to the platform. This allows the user's product to be listed immediately.
[0138] Step 9:
[0139] The terminal notifies the user that the listing is complete. The input is confirmation information that the listing is complete, and the output is a screen display that notifies the user of the completion. Specifically, the terminal displays a message such as "Your item has been listed."
[0140] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0141] This invention combines a system that automatically identifies objects and suggests their market price with an emotion engine that recognizes user emotions, enabling more personalized and effective listing suggestions. The system includes a camera, a terminal, a server, an emotion engine, and an online marketplace application.
[0142] When a user uses the system, they first photograph the object they want to sell using the associated camera on their device. The device then sends the video data to the server. The server receives the data, and an identification AI identifies the object. This AI identifies the object based on its appearance, shape, color, and brand information, and then queries a database to calculate the market price for selling it.
[0143] Simultaneously, the emotion engine recognizes the user's emotional state from their voice tone and facial expressions. This information is used to gain a deeper understanding of the user's intentions and interest in listing items. For example, if the user is excited, the system will recommend listing items in a more positive tone, while if the user is hesitant, the system will adjust to a more cautious approach.
[0144] The terminal receives the calculated market price from the server and displays it to the user via a display device. It also provides voice recommendations, suggesting to the user that they list the item on the online marketplace. In this process, the user's emotional state, recorded by an emotion engine, is taken into consideration, and the content and wording of the recommendations are adjusted accordingly. For example, if the user is in a calm state, the recommendation might say, "The current market price is approximately 8,000 yen. Would you consider listing it?"
[0145] When a user wishes to list an item for sale, they notify the server of their intention via their device. The server uses a generator to create a product description and image captions, adjusting the content and tone to reflect the user's emotional state. The listing information is then sent to the online marketplace application, completing the listing process.
[0146] For example, when a user takes a picture of a used digital camera they have at home, if the emotion engine recognizes the user's happy expression, the system will quickly display a market value and suggest, "This camera is worth 10,000 yen! Let's list it for sale right away!" In this way, by making appropriate suggestions based on the user's emotions, the invention makes the user's selling experience smoother and more intuitive.
[0147] The following describes the processing flow.
[0148] Step 1:
[0149] The user points the camera of the object they intend to sell at the device, using the camera attached to the terminal. The terminal acquires video data of that object.
[0150] Step 2:
[0151] The device compresses the video data it acquires and sends it to a server via the internet. This data is then stored in the cloud.
[0152] Step 3:
[0153] The server analyzes the received video data and uses identification AI to identify objects. The server extracts characteristic information about the objects and determines the product category and brand.
[0154] Step 4:
[0155] The server accesses a database based on past trading history and calculates the market selling price of the recognized object. During this process, the server performs statistical analysis based on the historical data.
[0156] Step 5:
[0157] The emotion engine analyzes the audio and video of the user's surroundings to identify the user's emotional state. This includes voice tone, facial expressions, and other biosignals.
[0158] Step 6:
[0159] The server sends the calculated selling price and the user's emotional state, identified by the emotion engine, to the terminal.
[0160] Step 7:
[0161] The device displays the market price for selling the item on its display screen and initiates voice recommendations that take the user's emotions into account. For example, if the user is excited, it will suggest in an assertive tone, and if they are hesitant, it will suggest in a cautious tone, "Why not list this item for 10,000 yen?"
[0162] Step 8:
[0163] If a user wishes to list an item for sale, they notify the server of their intention to sell via their device. This information is then sent to the server.
[0164] Step 9:
[0165] The server uses a generator to automatically create product descriptions and image captions. The style and tone of the descriptions are adjusted according to the user's emotional state.
[0166] Step 10:
[0167] The server sends the generated listing information to the online marketplace application, completing the listing process.
[0168] Step 11:
[0169] The device notifies the user that the listing is complete and displays a confirmation of success.
[0170] (Example 2)
[0171] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0172] Conventional object recognition and sales systems fail to consider user emotions and interests when suggesting items to sell, resulting in a mechanical and unintuitive listing experience. Furthermore, the inability to provide information reflecting the user's emotional state makes it difficult to suggest the optimal timing and method for listing items. Consequently, these systems fail to attract user interest and promote sales more effectively.
[0173] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0174] In this invention, the server includes means for analyzing video data acquired by a camera and automatically identifying a specific object; means for calculating the market selling price of the identified object by referring to market transaction history based on the information of the identified object; and means for analyzing the user's voice and facial expressions to recognize emotions and adjust the content of the listing suggestions based on the corresponding emotional state. This makes it possible to provide personalized listing suggestions that take into account the user's emotional state in addition to the market value of the object.
[0175] A "photography device" is a device used to acquire image data of an object, and is a device that has the function of recording the appearance of an object in high resolution.
[0176] "Image data" refers to image information acquired by a camera, and is digital data that includes the shape, color, and other characteristics of an object.
[0177] A "server" is a computer system that receives data via a network and performs analysis and processing on it.
[0178] "Identification AI" refers to a technology that uses artificial intelligence algorithms to extract object features from acquired video data and identify objects.
[0179] "Market selling price" refers to the selling price calculated based on past market transaction data for the identified object.
[0180] An "emotion engine" is a technology that analyzes and recognizes a user's emotional state based on their voice tone and facial expressions.
[0181] A "user" refers to an individual or group that uses the system to identify and list an object for sale.
[0182] A "generative AI model" is an algorithm that automatically generates information about an object, and is an artificial intelligence used to create product descriptions and image captions.
[0183] A "prompt statement" is an instruction given to an AI model, serving as a guideline for how to perform a specific task.
[0184] An "online marketplace application" is an internet platform for buying and selling physical objects, providing a space for users to list items for sale.
[0185] A description of the embodiment for carrying out the invention will be provided.
[0186] This system automatically identifies objects owned by the user, displays the market price for the relevant item, and makes listing suggestions based on the user's emotional state. The system consists of a camera, terminal, server, emotion engine, and online marketplace application.
[0187] First, the user operates a terminal and uses a camera to photograph the object they wish to sell. The camera acquires high-definition video data. The terminal sends this video data to a server. The server uses identification AI to analyze the video data and extract the object's features. These features include the object's appearance, shape, color, and brand information.
[0188] The server refers to a database and identifies objects based on extracted features. During this process, it uses past market transaction data to calculate the market selling price of the object. Simultaneously, the terminal collects the user's voice and facial expression data and analyzes it using an emotion engine. The user's emotional state is recognized as joy, excitement, hesitation, etc.
[0189] Based on the analysis results from the emotion engine, the device displays the market selling price to the user and suggests listing items through voice recommendations. At this time, the recommendations are adjusted according to the user's emotions. For example, if the user is excited, the device will suggest, "This camera is worth 10,000 yen! Let's list it for sale right away!"
[0190] When a user wishes to list an item for sale, their device notifies the server of their intention. The server uses a generative AI model to create a detailed product description and image captions, fine-tuning the content according to the user's emotional state. The generated information is sent to the online marketplace application, completing the listing process.
[0191] As a concrete example, if a user takes a picture of a used digital camera at home, and the emotion engine recognizes a happy expression on the device's face, the system will quickly suggest listing the camera while displaying a market price. An example of a prompt message given to the generating AI model would be, "Recognize the user's emotions from their facial expression and tone of voice, and adjust the listing accordingly."
[0192] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0193] Step 1:
[0194] The user picks up the object they wish to sell and activates the camera by operating a terminal. The camera acquires high-resolution video data from the object. This video data becomes the input.
[0195] Step 2:
[0196] The device sends the captured video data to the server. The server passes the input video data to an identification AI, which then begins analyzing the object. During this process, data processing is performed to extract features such as the object's appearance, shape, color, and brand information, and the characteristic information of the object is output.
[0197] Step 3:
[0198] The server queries a database for characteristic information and references the object's market transaction history. Based on this data reference, it calculates the market selling price of the object and outputs the price information.
[0199] Step 4:
[0200] The device records the user's voice and facial expressions using its camera and microphone. This data becomes input and is sent to the emotion engine. The emotion engine recognizes the user's emotions from the voice and facial expressions and outputs the emotional state.
[0201] Step 5:
[0202] The device generates listing suggestions for the user based on the output of the emotion engine (emotional state) and the market selling price received from the server. The content and tone of the suggestions include actions that are adjusted according to the emotional state. For example, a voice recommendation such as "This camera is worth 10,000 yen! Let's list it for sale right away!" might be output.
[0203] Step 6:
[0204] The user inputs their intention to list an item into the terminal. The terminal notifies the server of this intention, and the server calls a generation AI model to generate a product description and image captions. At this time, the intention to list the item and emotional state are input to the generation AI model as prompts, and the listing information is output with the content adjusted accordingly.
[0205] Step 7:
[0206] The server sends the generated listing information to the online marketplace application, completing the listing process. Finally, the listing information is confirmed by the user, and the listing is completed.
[0207] (Application Example 2)
[0208] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0209] Traditional online listing systems focus on object recognition and market price display, but fail to consider user emotions. Therefore, there is a need for methods that alleviate user anxiety and hesitation regarding listing items and promote listings more effectively.
[0210] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0211] In this invention, the server includes means for analyzing video data to recognize and automatically identify objects, means for calculating market selling prices based on the identified objects, and means for analyzing user emotions and adjusting listing suggestions. This enables flexible and effective listing suggestions that respond to user emotions.
[0212] "Means of object recognition" refers to the process of analyzing video data acquired by a camera to automatically identify a specific object.
[0213] "Method for calculating market selling price" refers to the process of calculating the market value of an identified object by referring to a database of past sales history based on information about the object.
[0214] The "method of suggesting items to be listed via voice" is a process that uses voice to recommend to users that they list an item for sale, based on the calculated market price.
[0215] "Methods for analyzing emotions and adjusting listing suggestions" refers to a process that recognizes the user's emotional state from their facial expressions and voice data, and adjusts the content of listing suggestions based on that information.
[0216] An "emotion engine" is a technology that processes facial expressions and voice to analyze a user's emotions and recognize their state.
[0217] The system implementing this invention is realized by having the user photograph an object using a device such as a smartphone, and processing that information on a server. The system mainly uses the following hardware and software.
[0218] The terminal is equipped with a camera to capture video data of the object the user wants to sell. This video data is transmitted to a server via a communication function. On the server, software that executes an object recognition algorithm (e.g., TensorFlow) analyzes the video data and automatically identifies the object. For the identified object, the server refers to a database containing past sales history and calculates the market selling price.
[0219] Furthermore, for emotion analysis, the terminal acquires the user's facial expressions and voice tone and sends this data to the server. The server uses libraries for facial recognition and voice analysis (e.g., OpenCV, Google® Speech-to-Text) to analyze the user's emotional state. The emotion engine generates the analysis results and adjusts the listing suggestions according to the user's emotions.
[0220] Using a generative AI model (e.g., OpenAI® GPT), the system generates suggestion messages tailored to the user's emotional state. These suggestions are presented to the user via the device through voice output or text display. For example, if the user expresses positive emotions, a prompt message such as "The current market price is 8,000 yen. Why not list your item now?" might be displayed.
[0221] In this way, the system enables flexible transactions that take into account the emotions of users when listing items, providing users with a more personalized listing experience.
[0222] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0223] Step 1:
[0224] The user takes a picture of the object they want to sell using a smartphone or other device. The input is the image data of the captured object, and the output is the transmission of that image data to the server. Upon the user's action of taking a picture, the device uses its communication function to transfer the image data to the server.
[0225] Step 2:
[0226] The server uses received image data to identify objects using object recognition software (e.g., TensorFlow). The input is image data sent by the user, and the output is the object identification result. The server analyzes the video data and identifies objects by referring to a database based on their appearance, shape, color, etc.
[0227] Step 3:
[0228] The server calculates the market selling price based on the object identification result and by referencing past trading history from the database. The input is the object identification result, and the output is the calculated market selling price. The server searches market data and performs the calculation of the selling price.
[0229] Step 4:
[0230] Simultaneously, the device captures the user's facial expressions with its camera and records voice input with its microphone. The input consists of facial expression data and voice data, and the output is the transmission of this data to the server. This emotion-related data is transferred from the device to the server based on the user's actions.
[0231] Step 5:
[0232] The server analyzes the received emotion data using a facial recognition library (e.g., OpenCV) and a speech analysis library (e.g., Google Speech-to-Text). The input is facial expression data and speech data, and the output is the emotion recognition result. The server processes this data to determine the user's emotional state.
[0233] Step 6:
[0234] The server uses a generative AI model (e.g., OpenAI GPT) to generate listing suggestions based on sentiment recognition results. The input is the sentiment recognition result and the market selling price, and the output is a refined suggestion text. The generative AI model creates prompt text and selects the most suitable suggestion for the user.
[0235] Step 7:
[0236] The terminal displays the generated proposal text to the user and provides audio output. The input is the refined proposal text, and the output is a visual and audio proposal to the user. The terminal uses display and audio functions to communicate the market price and proposal to the user.
[0237] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0238] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0239] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0240] [Second Embodiment]
[0241] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0242] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0243] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0244] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0245] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0246] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0247] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0248] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0249] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0250] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0251] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0252] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0253] The system of this invention enables efficient listing on online marketplaces by allowing users to easily identify objects they own and quickly determine their market value. This system consists of a camera, a terminal, a server, and an online marketplace application.
[0254] When a user uses the system, they first use a camera associated with their device to photograph the object they intend to sell. This camera is typically a built-in camera in a smartphone or AR glasses. The device then acquires image data of the object and transmits it to a server via the internet.
[0255] The server analyzes the received video data. Using identification AI, it identifies objects from the video and extracts their characteristics. This includes the object's appearance, shape, and color. Based on the extracted information, the server queries a database containing past sales history to calculate the current market price for selling the object.
[0256] The device visually displays market price information transmitted from the server on its display device. This display can be the screen of AR glasses or a smartphone. The device also uses voice recommendation functionality to suggest that the user list an item on an online marketplace. This voice assistant might make suggestions such as, "You can sell this item for approximately 5,000 yen. Why not list it?"
[0257] If a user wishes to list an item for sale following the terminal's suggestions, they can indicate their intention to list the item through the terminal's interface. In this case, the server automatically generates listing information for the item using a generator. Specifically, a product description, title, and accompanying image captions are generated. This data is then compiled and sent to the online marketplace application.
[0258] Finally, the server completes the listing process on the online marketplace application. This entire process allows users to sell their items quickly and easily. This system promotes the effective use of users' goods and enhances convenience and sustainability.
[0259] For example, if a user takes a picture of an old smartphone at home, the system will identify the smartphone and indicate its current market value of approximately 15,000 yen. If the user then wishes to list it for sale, the system will quickly complete the listing using automatically generated information. This allows users to maximize their profits while saving time and effort.
[0260] The following describes the processing flow.
[0261] Step 1:
[0262] The user points the object they want to sell at the camera using a camera device associated with the terminal. The terminal acquires video data of this object in real time.
[0263] Step 2:
[0264] The terminal efficiently compresses the video data it acquires and sends it to the server via the internet. The server stores the received data.
[0265] Step 3:
[0266] The server analyzes the received video data and uses identification AI to identify objects. The server then extracts the object's characteristics (e.g., shape, color, brand, etc.).
[0267] Step 4:
[0268] The server uses the extracted object information to access a database that holds past trading history and calculates the market selling price of the identified object. In this process, the server utilizes statistical analysis and machine learning techniques.
[0269] Step 5:
[0270] The server sends the calculated market price information to the terminal. The terminal receives this information and presents it visually to the user through a display device.
[0271] Step 6:
[0272] The device uses its voice output function to offer the user a voice message suggesting they list the item for sale, such as, "This item can currently be sold for approximately 15,000 yen. Would you like to list it?"
[0273] Step 7:
[0274] When a user wishes to list an item for sale, they express their intention through voice commands or touch inputs via the terminal. The terminal notifies the server of this intention.
[0275] Step 8:
[0276] The server uses a generation device to automatically generate product descriptions and image captions, and constructs the listing information. This information is prepared for the online market application.
[0277] Step 9:
[0278] The server transmits the listing information it has constructed to the online market application via a communication device to complete the listing procedure.
[0279] Step 10:
[0280] The terminal notifies the user of the completion of the listing and reports to the user that the listing has been successfully carried out.
[0281] (Example 1)
[0282] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0283] In modern information society, it is still complex and time-consuming for an individual to list their owned items on the market quickly and at an appropriate price. Also, specialized knowledge and skills are required for item identification and market price calculation, which pose high barriers for general consumers to utilize. Therefore, there is an issue of how to simplify and efficiently perform these processes.
[0284] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0285] In this invention, the server includes means for analyzing video data acquired by imaging means and automatically identifying a specific object using image processing technology, means for calculating the value in the transaction of the identified object by referring to a recording medium having a past transaction history based on the visual features of the identified object, presenting the calculated value information through display means for displaying it, and means for recommending the transaction of the object by voice. As a result, even without specialized knowledge, the user can easily and quickly put an item on the market and conduct transactions at an appropriate price.
[0286] The "imaging means" is a device or function for acquiring video data of an object, and is a device including a camera for taking pictures or videos.
[0287] The "image processing technology" is a method for analyzing the acquired video data, and uses an algorithm or artificial intelligence technology for identifying an object in an image.
[0288] The "object" is an object that the user is considering selling, is identified by the system, and is treated as the subject of a transaction.
[0289] The "recording medium" is a data storage system for holding information such as past transaction histories, and includes a database and a digital storage device.
[0290] The "value" is the transaction price in the market calculated based on the identified object, and is estimated based on past transaction data.
[0291] The "display means" is a device for visually transmitting the calculated value and information to the user, and is a device including a display or a screen.
[0292] The "means for recommending by voice" is a device or system having a function of proposing to the user by voice to prompt the transaction of the object.
[0293] "Communication means" refers to devices or systems for transferring information over a network, including protocols for internet connectivity and data communication.
[0294] An "e-commerce platform" is a software infrastructure that enables the online trading of goods and is a system for conducting buying and selling transactions.
[0295] This invention is a system for efficiently listing items owned by users on the market, and it functions through the coordinated operation of multiple devices and software. The user photographs the items they wish to sell using a photographic means. This photographic means uses a camera built into a smartphone or augmented reality glasses. The captured video data is acquired by a terminal, which processes the image and transfers it to a server.
[0296] The server analyzes the acquired video data using image processing technology. Here, an identification AI model is used to identify objects and extract their features. The AI model used is created using TensorFlow or PyTorch. The extracted feature information is then compared with past transaction history stored on the recording medium. Based on this, the server calculates the market value of the object.
[0297] The calculated value information is transmitted to the terminal. The terminal uses a smartphone screen or augmented reality glasses display to provide the user with visual information. It also encourages the user to trade the item through voice prompts. For example, it might prompt, "You can sell this item for about 5,000 yen. Why not list it?"
[0298] When a user requests to list an item, the server automatically generates product descriptions and image annotations using a generative AI model, and then constructs the transaction information. This generated transaction information is then transferred to the e-commerce platform via communication. Finally, the server completes the listing process, allowing the user to easily and quickly list their items on the marketplace.
[0299] As a specific example, when a user holds an old smartphone at home in front of the system for shooting, the captured data is analyzed and shown to have a market value of about 15,000 yen. If the user wishes to list it for sale, transactions can be quickly enabled by the listing information automatically generated by the system. Through this process, the user can conduct optimal transactions while saving time and effort.
[0300] The flow of the specific process in Example 1 will be described using FIG. 11.
[0301] Step 1:
[0302] The user shoots the item under consideration for sale using a photographing device. Specifically, a camera built into a smartphone or augmented reality glasses is used to obtain video data of the object. This input video data is utilized for processing within the terminal.
[0303] Step 2:
[0304] The terminal formats the acquired video data into a predetermined format and transmits it to the server via the Internet. The output at this stage is video data formatted in a form that can be received by the server.
[0305] Step 3:
[0306] The server inputs the received video data into an identification AI model. Image processing for video analysis is performed to extract features such as the appearance and shape of the object. Through this process, the server identifies what the object is and obtains the feature information of the identified object as its output.
[0307] Step 4:
[0308] The server uses the extracted characteristic information to refer to past transaction history stored on the recording medium. Using database queries, it extracts historical data on transactions of similar objects and calculates the market value based on this data. The output is market price information representing the current market value of the object.
[0309] Step 5:
[0310] The server sends the calculated market price information to the terminal. The terminal visually displays the output market price on the smartphone screen or the display of augmented reality glasses. This allows the user to check the current market price of an object.
[0311] Step 6:
[0312] The device uses a voice assistant to suggest selling an item to the user. Specifically, it communicates the suggestion by voice and prompts the user to take the next action. The output of this step is that a valid selling suggestion is made to the user.
[0313] Step 7:
[0314] If a user wishes to sell based on the terminal's suggestions, they input their intention into the terminal. The terminal transmits this information to the server. Based on this input, the server uses a generative AI model to automatically generate product descriptions and image annotations, creating the listing information. This output contains all the listing information necessary for the transaction.
[0315] Step 8:
[0316] The server transmits the generated listing information to the e-commerce platform via communication means, completing the listing process. This final output means the item is listed on the marketplace and ready for immediate trading.
[0317] (Application Example 1)
[0318] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0319] Current online trading platforms present challenges in quickly and easily listing individual items, requiring users to manually input vast amounts of information and necessitating time-consuming market price searches. As a result, users often find the trading process burdensome, making it difficult to sell items quickly at appropriate market prices. Furthermore, the lack of mechanisms to ensure users can list items at fair prices can delay the listing process, potentially leading to missed sales opportunities.
[0320] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0321] In this invention, the server includes means for analyzing video data acquired by a camera to recognize an object and automatically identifying a specific object; means for calculating the market selling price of the identified object by referring to accumulated information of past transaction history based on the information of the identified object; and means for transmitting the generated listing information to an online trading platform via a communication device and completing the procedure for selling the object. This enables users to quickly grasp market prices and immediately list objects on the online platform at a fair price.
[0322] An "object" is something that the user intends to sell and that is recognized by the camera.
[0323] A "camera" is hardware used to acquire video data, and is mainly installed in mobile devices and other devices with camera functions.
[0324] "Video data" refers to visual information acquired by a camera or camera, and is fundamental information for identifying objects.
[0325] "Automatic identification means" refers to a combination of software or hardware that uses acquired video data to perform a process for identifying a specific object.
[0326] "Market selling price" is an estimate of the current market price of an object, calculated based on past transaction history.
[0327] A "generation device" refers to a software or hardware configuration for automatically generating listing information.
[0328] A "visual display device" is a digital screen or display used to present information to a user visually.
[0329] A "communication device" is a device used to send and receive data via a network connection.
[0330] An "online trading platform" refers to a website or application where goods and services are bought and sold via the internet.
[0331] A "generative AI model" is an artificial intelligence technology used to automatically generate text and information.
[0332] The system of this invention provides an efficient means for users to quickly put their owned objects up for market. The system mainly consists of a camera, a terminal, a server, and an online trading platform.
[0333] First, the user acquires video data of the object they wish to sell using the camera function of their portable information terminal. The camera function of a smartphone or tablet is used as the camera. This video data is then transmitted to a server via the internet through the terminal.
[0334] The server analyzes the received video data using identification AI to automatically identify specific objects. Information about the identified objects is then referenced in a database containing past transaction history to calculate the market price for those objects. This price information is transmitted to the terminal and displayed on the user's device. A speech synthesis system is also utilized, allowing users to receive voice-based listing suggestions.
[0335] When a user indicates their willingness to follow voice suggestions during the listing process, the server generates listing information using a generation device. This generation AI model automatically generates product descriptions and visual annotations. The generated listing information is then transmitted to the online trading platform via a communication device, completing the sale process for the item.
[0336] As a concrete example, let's assume a user takes a picture of a used laptop computer they have on hand using this system. The system recognizes the object and suggests its market price is, for example, 20,000 yen. If the user wishes to list it for sale, the system generates a product description and caption, and quickly lists it on an online trading platform. In this process, an example of a prompt message used with the generating AI model would be: "This is a used laptop computer, 3 years old and in good condition. We recommend listing it for 20,000 yen."
[0337] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0338] Step 1:
[0339] The user uses the camera on their portable information terminal to acquire video data of the object they wish to sell. Here, the input is the video data of the object, and the output is the saving of the video data to the terminal. This video data is then prepared to be sent to a server for subsequent processing.
[0340] Step 2:
[0341] The terminal transmits video data to the server. The input is the acquired video data, and the output is the transmission of the image to the server. Specifically, the terminal transfers data to the server using a network connection.
[0342] Step 3:
[0343] The server analyzes the received video data and automatically identifies specific objects using identification AI. The input is the transmitted video data, and the output is information about the identified objects. In this process, the server extracts and recognizes features such as the appearance, shape, and color of the objects.
[0344] Step 4:
[0345] The server uses identified object information to refer to a database and calculate the market selling price. The input is information about the identified object, and the output is the calculated market price. The server calculates the price based on past transaction history.
[0346] Step 5:
[0347] The server sends the calculated market price to the user's terminal. The input is the calculated market price, and the output is the transmission of price information to the terminal. This allows the user to view the price information on a visual display device.
[0348] Step 6:
[0349] The terminal displays the received market price on a visual display and uses a speech synthesis system to suggest listing the item. The input is price information sent from the server, and the output is the displayed price and a voice suggestion. The operation includes a process in which the voice assistant suggests, "Why don't you list this item for sale?"
[0350] Step 7:
[0351] When a user enters their intention to list an item into the terminal, the server generates listing information using a generator. The input is the user's intention to list an item, and the output is the generated listing information. This information includes a product description and image captions.
[0352] Step 8:
[0353] The server transmits the generated listing information to the online trading platform via a communication device. The input is the generated listing information, and the output is the transmission of information to the platform. This allows the user's product to be listed immediately.
[0354] Step 9:
[0355] The terminal notifies the user that the listing is complete. The input is confirmation information that the listing is complete, and the output is a screen display that notifies the user of the completion. Specifically, the terminal displays a message such as "Your item has been listed."
[0356] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0357] This invention combines a system that automatically identifies objects and suggests their market price with an emotion engine that recognizes user emotions, enabling more personalized and effective listing suggestions. The system includes a camera, a terminal, a server, an emotion engine, and an online marketplace application.
[0358] When a user uses the system, they first photograph the object they want to sell using the associated camera on their device. The device then sends the video data to the server. The server receives the data, and an identification AI identifies the object. This AI identifies the object based on its appearance, shape, color, and brand information, and then queries a database to calculate the market price for selling it.
[0359] Simultaneously, the emotion engine recognizes the user's emotional state from their voice tone and facial expressions. This information is used to gain a deeper understanding of the user's intentions and interest in listing items. For example, if the user is excited, the system will recommend listing items in a more positive tone, while if the user is hesitant, the system will adjust to a more cautious approach.
[0360] The terminal receives the calculated market price from the server and displays it to the user via a display device. It also provides voice recommendations, suggesting to the user that they list the item on the online marketplace. In this process, the user's emotional state, recorded by an emotion engine, is taken into consideration, and the content and wording of the recommendations are adjusted accordingly. For example, if the user is in a calm state, the recommendation might say, "The current market price is approximately 8,000 yen. Would you consider listing it?"
[0361] When a user wishes to list an item for sale, they notify the server of their intention via their device. The server uses a generator to create a product description and image captions, adjusting the content and tone to reflect the user's emotional state. The listing information is then sent to the online marketplace application, completing the listing process.
[0362] For example, when a user takes a picture of a used digital camera they have at home, if the emotion engine recognizes the user's happy expression, the system will quickly display a market value and suggest, "This camera is worth 10,000 yen! Let's list it for sale right away!" In this way, by making appropriate suggestions based on the user's emotions, the invention makes the user's selling experience smoother and more intuitive.
[0363] The following describes the processing flow.
[0364] Step 1:
[0365] The user points the camera of the object they intend to sell at the device, using the camera attached to the terminal. The terminal acquires video data of that object.
[0366] Step 2:
[0367] The device compresses the video data it acquires and sends it to a server via the internet. This data is then stored in the cloud.
[0368] Step 3:
[0369] The server analyzes the received video data and uses identification AI to identify objects. The server extracts characteristic information about the objects and determines the product category and brand.
[0370] Step 4:
[0371] The server accesses a database based on past trading history and calculates the market selling price of the recognized object. During this process, the server performs statistical analysis based on the historical data.
[0372] Step 5:
[0373] The emotion engine analyzes the audio and video of the user's surroundings to identify the user's emotional state. This includes voice tone, facial expressions, and other biosignals.
[0374] Step 6:
[0375] The server sends the calculated selling price and the user's emotional state, identified by the emotion engine, to the terminal.
[0376] Step 7:
[0377] The device displays the market price for selling the item on its display screen and initiates voice recommendations that take the user's emotions into account. For example, if the user is excited, it will suggest in an assertive tone, and if they are hesitant, it will suggest in a cautious tone, "Why not list this item for 10,000 yen?"
[0378] Step 8:
[0379] If a user wishes to list an item for sale, they notify the server of their intention to sell via their device. This information is then sent to the server.
[0380] Step 9:
[0381] The server uses a generator to automatically create product descriptions and image captions. The style and tone of the descriptions are adjusted according to the user's emotional state.
[0382] Step 10:
[0383] The server sends the generated listing information to the online marketplace application, completing the listing process.
[0384] Step 11:
[0385] The device notifies the user that the listing is complete and displays a confirmation of success.
[0386] (Example 2)
[0387] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0388] Conventional object recognition and sales systems fail to consider user emotions and interests when suggesting items to sell, resulting in a mechanical and unintuitive listing experience. Furthermore, the inability to provide information reflecting the user's emotional state makes it difficult to suggest the optimal timing and method for listing items. Consequently, these systems fail to attract user interest and promote sales more effectively.
[0389] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0390] In this invention, the server includes means for analyzing video data acquired by a camera and automatically identifying a specific object; means for calculating the market selling price of the identified object by referring to market transaction history based on the information of the identified object; and means for analyzing the user's voice and facial expressions to recognize emotions and adjust the content of the listing suggestions based on the corresponding emotional state. This makes it possible to provide personalized listing suggestions that take into account the user's emotional state in addition to the market value of the object.
[0391] A "photography device" is a device used to acquire image data of an object, and is a device that has the function of recording the appearance of an object in high resolution.
[0392] "Image data" refers to image information acquired by a camera, and is digital data that includes the shape, color, and other characteristics of an object.
[0393] A "server" is a computer system that receives data via a network and performs analysis and processing on it.
[0394] "Identification AI" refers to a technology that uses artificial intelligence algorithms to extract object features from acquired video data and identify objects.
[0395] "Market selling price" refers to the selling price calculated based on past market transaction data for the identified object.
[0396] An "emotion engine" is a technology that analyzes and recognizes a user's emotional state based on their voice tone and facial expressions.
[0397] A "user" refers to an individual or group that uses the system to identify and list an object for sale.
[0398] A "generative AI model" is an algorithm that automatically generates information about an object, and is an artificial intelligence used to create product descriptions and image captions.
[0399] A "prompt statement" is an instruction given to an AI model, serving as a guideline for how to perform a specific task.
[0400] An "online marketplace application" is an internet platform for buying and selling physical objects, providing a space for users to list items for sale.
[0401] A description of the embodiment for carrying out the invention will be provided.
[0402] This system automatically identifies objects owned by the user, displays the market price for the relevant item, and makes listing suggestions based on the user's emotional state. The system consists of a camera, terminal, server, emotion engine, and online marketplace application.
[0403] First, the user operates a terminal and uses a camera to photograph the object they wish to sell. The camera acquires high-definition video data. The terminal sends this video data to a server. The server uses identification AI to analyze the video data and extract the object's features. These features include the object's appearance, shape, color, and brand information.
[0404] The server refers to a database and identifies objects based on extracted features. During this process, it uses past market transaction data to calculate the market selling price of the object. Simultaneously, the terminal collects the user's voice and facial expression data and analyzes it using an emotion engine. The user's emotional state is recognized as joy, excitement, hesitation, etc.
[0405] Based on the analysis results from the emotion engine, the device displays the market selling price to the user and suggests listing items through voice recommendations. At this time, the recommendations are adjusted according to the user's emotions. For example, if the user is excited, the device will suggest, "This camera is worth 10,000 yen! Let's list it for sale right away!"
[0406] When a user wishes to list an item for sale, their device notifies the server of their intention. The server uses a generative AI model to create a detailed product description and image captions, fine-tuning the content according to the user's emotional state. The generated information is sent to the online marketplace application, completing the listing process.
[0407] As a concrete example, if a user takes a picture of a used digital camera at home, and the emotion engine recognizes a happy expression on the device's face, the system will quickly suggest listing the camera while displaying a market price. An example of a prompt message given to the generating AI model would be, "Recognize the user's emotions from their facial expression and tone of voice, and adjust the listing accordingly."
[0408] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0409] Step 1:
[0410] The user picks up the object they wish to sell and activates the camera by operating a terminal. The camera acquires high-resolution video data from the object. This video data becomes the input.
[0411] Step 2:
[0412] The device sends the captured video data to the server. The server passes the input video data to an identification AI, which then begins analyzing the object. During this process, data processing is performed to extract features such as the object's appearance, shape, color, and brand information, and the characteristic information of the object is output.
[0413] Step 3:
[0414] The server queries a database for characteristic information and references the object's market transaction history. Based on this data reference, it calculates the market selling price of the object and outputs the price information.
[0415] Step 4:
[0416] The device records the user's voice and facial expressions using its camera and microphone. This data becomes input and is sent to the emotion engine. The emotion engine recognizes the user's emotions from the voice and facial expressions and outputs the emotional state.
[0417] Step 5:
[0418] The device generates listing suggestions for the user based on the output of the emotion engine (emotional state) and the market selling price received from the server. The content and tone of the suggestions include actions that are adjusted according to the emotional state. For example, a voice recommendation such as "This camera is worth 10,000 yen! Let's list it for sale right away!" might be output.
[0419] Step 6:
[0420] The user inputs their intention to list an item into the terminal. The terminal notifies the server of this intention, and the server calls a generation AI model to generate a product description and image captions. At this time, the intention to list the item and emotional state are input to the generation AI model as prompts, and the listing information is output with the content adjusted accordingly.
[0421] Step 7:
[0422] The server sends the generated listing information to the online marketplace application, completing the listing process. Finally, the listing information is confirmed by the user, and the listing is completed.
[0423] (Application Example 2)
[0424] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0425] Traditional online listing systems focus on object recognition and market price display, but fail to consider user emotions. Therefore, there is a need for methods that alleviate user anxiety and hesitation regarding listing items and promote listings more effectively.
[0426] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0427] In this invention, the server includes means for analyzing video data to recognize and automatically identify objects, means for calculating market selling prices based on the identified objects, and means for analyzing user emotions and adjusting listing suggestions. This enables flexible and effective listing suggestions that respond to user emotions.
[0428] "Means of object recognition" refers to the process of analyzing video data acquired by a camera to automatically identify a specific object.
[0429] "Method for calculating market selling price" refers to the process of calculating the market value of an identified object by referring to a database of past sales history based on information about the object.
[0430] The "method of suggesting items to be listed via voice" is a process that uses voice to recommend to users that they list an item for sale, based on the calculated market price.
[0431] "Methods for analyzing emotions and adjusting listing suggestions" refers to a process that recognizes the user's emotional state from their facial expressions and voice data, and adjusts the content of listing suggestions based on that information.
[0432] An "emotion engine" is a technology that processes facial expressions and voice to analyze a user's emotions and recognize their state.
[0433] The system implementing this invention is realized by having the user photograph an object using a device such as a smartphone, and processing that information on a server. The system mainly uses the following hardware and software.
[0434] The terminal is equipped with a camera to capture video data of the object the user wants to sell. This video data is transmitted to a server via a communication function. On the server, software that executes an object recognition algorithm (e.g., TensorFlow) analyzes the video data and automatically identifies the object. For the identified object, the server refers to a database containing past sales history and calculates the market selling price.
[0435] Furthermore, for emotion analysis, the device acquires the user's facial expressions and voice tone and sends this data to the server. The server uses libraries for facial recognition and voice analysis (e.g., OpenCV, Google Speech-to-Text) to analyze the user's emotional state. The emotion engine generates the analysis results and adjusts the listing suggestions according to the user's emotions.
[0436] Using a generative AI model (e.g., OpenAI GPT), the system generates suggestion messages tailored to the user's emotional state. These suggestions are presented to the user via the device through voice output or text display. For example, if the user expresses positive emotions, a prompt message such as "The current market price is 8,000 yen. Why not list your item now?" might be displayed.
[0437] In this way, the system enables flexible transactions that take into account the emotions of users when listing items, providing users with a more personalized listing experience.
[0438] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0439] Step 1:
[0440] The user takes a picture of the object they want to sell using a smartphone or other device. The input is the image data of the captured object, and the output is the transmission of that image data to the server. Upon the user's action of taking a picture, the device uses its communication function to transfer the image data to the server.
[0441] Step 2:
[0442] The server uses received image data to identify objects using object recognition software (e.g., TensorFlow). The input is image data sent by the user, and the output is the object identification result. The server analyzes the video data and identifies objects by referring to a database based on their appearance, shape, color, etc.
[0443] Step 3:
[0444] The server calculates the market selling price based on the object identification result and by referencing past trading history from the database. The input is the object identification result, and the output is the calculated market selling price. The server searches market data and performs the calculation of the selling price.
[0445] Step 4:
[0446] Simultaneously, the device captures the user's facial expressions with its camera and records voice input with its microphone. The input consists of facial expression data and voice data, and the output is the transmission of this data to the server. This emotion-related data is transferred from the device to the server based on the user's actions.
[0447] Step 5:
[0448] The server analyzes the received emotion data using a facial recognition library (e.g., OpenCV) and a speech analysis library (e.g., Google Speech-to-Text). The input is facial expression data and speech data, and the output is the emotion recognition result. The server processes this data to determine the user's emotional state.
[0449] Step 6:
[0450] The server uses a generative AI model (e.g., OpenAI GPT) to generate listing suggestions based on sentiment recognition results. The input is the sentiment recognition result and the market selling price, and the output is a refined suggestion text. The generative AI model creates prompt text and selects the most suitable suggestion for the user.
[0451] Step 7:
[0452] The terminal displays the generated proposal text to the user and provides audio output. The input is the refined proposal text, and the output is a visual and audio proposal to the user. The terminal uses display and audio functions to communicate the market price and proposal to the user.
[0453] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0454] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0455] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0456] [Third Embodiment]
[0457] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0458] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0459] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0460] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0461] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0462] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0463] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0464] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0465] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0466] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0467] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0468] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0469] The system of this invention enables efficient listing on online marketplaces by allowing users to easily identify objects they own and quickly determine their market value. This system consists of a camera, a terminal, a server, and an online marketplace application.
[0470] When a user uses the system, they first use a camera associated with their device to photograph the object they intend to sell. This camera is typically a built-in camera in a smartphone or AR glasses. The device then acquires image data of the object and transmits it to a server via the internet.
[0471] The server analyzes the received video data. Using identification AI, it identifies objects from the video and extracts their characteristics. This includes the object's appearance, shape, and color. Based on the extracted information, the server queries a database containing past sales history to calculate the current market price for selling the object.
[0472] The device visually displays market price information transmitted from the server on its display device. This display can be the screen of AR glasses or a smartphone. The device also uses voice recommendation functionality to suggest that the user list an item on an online marketplace. This voice assistant might make suggestions such as, "You can sell this item for approximately 5,000 yen. Why not list it?"
[0473] If a user wishes to list an item for sale following the terminal's suggestions, they can indicate their intention to list the item through the terminal's interface. In this case, the server automatically generates listing information for the item using a generator. Specifically, a product description, title, and accompanying image captions are generated. This data is then compiled and sent to the online marketplace application.
[0474] Finally, the server completes the listing process on the online marketplace application. This entire process allows users to sell their items quickly and easily. This system promotes the effective use of users' goods and enhances convenience and sustainability.
[0475] For example, if a user takes a picture of an old smartphone at home, the system will identify the smartphone and indicate its current market value of approximately 15,000 yen. If the user then wishes to list it for sale, the system will quickly complete the listing using automatically generated information. This allows users to maximize their profits while saving time and effort.
[0476] The following describes the processing flow.
[0477] Step 1:
[0478] The user points the object they want to sell at the camera using a camera device associated with the terminal. The terminal acquires video data of this object in real time.
[0479] Step 2:
[0480] The terminal efficiently compresses the video data it acquires and sends it to the server via the internet. The server stores the received data.
[0481] Step 3:
[0482] The server analyzes the received video data and uses identification AI to identify objects. The server then extracts the object's characteristics (e.g., shape, color, brand, etc.).
[0483] Step 4:
[0484] The server uses the extracted object information to access a database that holds past trading history and calculates the market selling price of the identified object. In this process, the server utilizes statistical analysis and machine learning techniques.
[0485] Step 5:
[0486] The server sends the calculated market price information to the terminal. The terminal receives this information and presents it visually to the user through a display device.
[0487] Step 6:
[0488] The device uses its voice output function to offer the user a voice message suggesting they list the item for sale, such as, "This item can currently be sold for approximately 15,000 yen. Would you like to list it?"
[0489] Step 7:
[0490] If a user wishes to list an item for sale, they express their intention via voice command or touch input through their device. The device then notifies the server of this intention.
[0491] Step 8:
[0492] The server uses a generator to automatically produce product descriptions and image captions, thus consolidating the listing information. This information is then prepared for use in online marketplace applications.
[0493] Step 9:
[0494] The server sends the listing information it has configured to the online marketplace application via a communication device, completing the listing process.
[0495] Step 10:
[0496] The device notifies the user that the listing is complete and reports to the user that the listing was successfully completed.
[0497] (Example 1)
[0498] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0499] In today's information society, listing personal belongings on the market quickly and at a fair price remains a complex and time-consuming process. Furthermore, identifying items and calculating market prices requires specialized knowledge and skills, making it difficult for the average consumer to use. Therefore, there is a challenge in how to simplify and efficiently carry out these processes.
[0500] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0501] In this invention, the server includes means for analyzing video data acquired by a shooting means and automatically identifying a specific object using image processing technology; means for calculating the transaction value of the identified object by referring to a recording medium containing past transaction history based on the visual characteristics of the identified object; and means for presenting the calculated value information through a display means and recommending the transaction of the object by voice. As a result, users can easily and quickly list items on the market and trade them at a fair price without having specialized knowledge.
[0502] "Means of shooting" refers to a device or function for acquiring image data of an object, and includes devices such as cameras for taking images or videos.
[0503] "Image processing technology" refers to methods for analyzing acquired video data, and involves the use of algorithms and artificial intelligence technologies to identify objects within images.
[0504] An "object" is an object that a user intends to sell, which is identified by the system and treated as the subject of a transaction.
[0505] A "recording medium" is a data storage system used to hold information such as past transaction history, and includes databases and digital storage devices.
[0506] "Value" refers to the market transaction price calculated based on the identified object, and is estimated based on past transaction data.
[0507] "Display means" refers to a device used to visually communicate calculated value or information to the user, and includes devices such as displays and screens.
[0508] "Voice-based recommendation methods" refer to devices or systems that have the function of using voice to suggest to users how to encourage them to trade an item.
[0509] "Communication means" refers to devices or systems for transferring information over a network, including protocols for internet connectivity and data communication.
[0510] An "e-commerce platform" is a software infrastructure that enables the online trading of goods and is a system for conducting buying and selling transactions.
[0511] This invention is a system for efficiently listing items owned by users on the market, and it functions through the coordinated operation of multiple devices and software. The user photographs the items they wish to sell using a photographic means. This photographic means uses a camera built into a smartphone or augmented reality glasses. The captured video data is acquired by a terminal, which processes the image and transfers it to a server.
[0512] The server analyzes the acquired video data using image processing technology. Here, an identification AI model is used to identify objects and extract their features. The AI model used is created using TensorFlow or PyTorch. The extracted feature information is then compared with past transaction history stored on the recording medium. Based on this, the server calculates the market value of the object.
[0513] The calculated value information is transmitted to the terminal. The terminal uses a smartphone screen or augmented reality glasses display to provide the user with visual information. It also encourages the user to trade the item through voice prompts. For example, it might prompt, "You can sell this item for about 5,000 yen. Why not list it?"
[0514] When a user requests to list an item, the server automatically generates product descriptions and image annotations using a generative AI model, and then constructs the transaction information. This generated transaction information is then transferred to the e-commerce platform via communication. Finally, the server completes the listing process, allowing the user to easily and quickly list their items on the marketplace.
[0515] For example, if a user holds an old smartphone from their home up to the system and takes a picture, the system analyzes the captured data and indicates a market value of approximately 15,000 yen. If the user wishes to list the item for sale, the system automatically generates listing information, enabling a quick transaction. Through this process, users can make optimal transactions while saving time and effort.
[0516] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0517] Step 1:
[0518] The user photographs the items they are considering selling using a photographic device. Specifically, they use the camera built into their smartphone or augmented reality glasses to acquire video data of the objects. This video data is then used for processing within the device.
[0519] Step 2:
[0520] The terminal formats the acquired video data into a predetermined format and sends it to the server via the internet. At this stage, the output is video data formatted in a format that the server can receive.
[0521] Step 3:
[0522] The server inputs the received video data into an identification AI model. It performs image processing for video analysis, extracting features such as the appearance and shape of objects. Through this process, the server identifies what the objects are and obtains feature information of the identified objects as output.
[0523] Step 4:
[0524] The server uses the extracted characteristic information to refer to past transaction history stored on the recording medium. Using database queries, it extracts historical data on transactions of similar objects and calculates the market value based on this data. The output is market price information representing the current market value of the object.
[0525] Step 5:
[0526] The server sends the calculated market price information to the terminal. The terminal visually displays the output market price on the smartphone screen or the display of augmented reality glasses. This allows the user to check the current market price of an object.
[0527] Step 6:
[0528] The device uses a voice assistant to suggest selling an item to the user. Specifically, it communicates the suggestion by voice and prompts the user to take the next action. The output of this step is that a valid selling suggestion is made to the user.
[0529] Step 7:
[0530] If a user wishes to sell based on the terminal's suggestions, they input their intention into the terminal. The terminal transmits this information to the server. Based on this input, the server uses a generative AI model to automatically generate product descriptions and image annotations, creating the listing information. This output contains all the listing information necessary for the transaction.
[0531] Step 8:
[0532] The server transmits the generated listing information to the e-commerce platform via communication means, completing the listing process. This final output means the item is listed on the marketplace and ready for immediate trading.
[0533] (Application Example 1)
[0534] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0535] Current online trading platforms present challenges in quickly and easily listing individual items, requiring users to manually input vast amounts of information and necessitating time-consuming market price searches. As a result, users often find the trading process burdensome, making it difficult to sell items quickly at appropriate market prices. Furthermore, the lack of mechanisms to ensure users can list items at fair prices can delay the listing process, potentially leading to missed sales opportunities.
[0536] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0537] In this invention, the server includes means for analyzing video data acquired by a camera to recognize an object and automatically identifying a specific object; means for calculating the market selling price of the identified object by referring to accumulated information of past transaction history based on the information of the identified object; and means for transmitting the generated listing information to an online trading platform via a communication device and completing the procedure for selling the object. This enables users to quickly grasp market prices and immediately list objects on the online platform at a fair price.
[0538] An "object" is something that the user intends to sell and that is recognized by the camera.
[0539] A "camera" is hardware used to acquire video data, and is mainly installed in mobile devices and other devices with camera functions.
[0540] "Video data" refers to visual information acquired by a camera or camera, and is fundamental information for identifying objects.
[0541] "Automatic identification means" refers to a combination of software or hardware that uses acquired video data to perform a process for identifying a specific object.
[0542] "Market selling price" is an estimate of the current market price of an object, calculated based on past transaction history.
[0543] A "generation device" refers to a software or hardware configuration for automatically generating listing information.
[0544] A "visual display device" is a digital screen or display used to present information to a user visually.
[0545] A "communication device" is a device used to send and receive data via a network connection.
[0546] An "online trading platform" refers to a website or application where goods and services are bought and sold via the internet.
[0547] A "generative AI model" is an artificial intelligence technology used to automatically generate text and information.
[0548] The system of this invention provides an efficient means for users to quickly put their owned objects up for market. The system mainly consists of a camera, a terminal, a server, and an online trading platform.
[0549] First, the user acquires video data of the object they wish to sell using the camera function of their portable information terminal. The camera function of a smartphone or tablet is used as the camera. This video data is then transmitted to a server via the internet through the terminal.
[0550] The server analyzes the received video data using identification AI to automatically identify specific objects. Information about the identified objects is then referenced in a database containing past transaction history to calculate the market price for those objects. This price information is transmitted to the terminal and displayed on the user's device. A speech synthesis system is also utilized, allowing users to receive voice-based listing suggestions.
[0551] When a user indicates their willingness to follow voice suggestions during the listing process, the server generates listing information using a generation device. This generation AI model automatically generates product descriptions and visual annotations. The generated listing information is then transmitted to the online trading platform via a communication device, completing the sale process for the item.
[0552] As a concrete example, let's assume a user takes a picture of a used laptop computer they have on hand using this system. The system recognizes the object and suggests its market price is, for example, 20,000 yen. If the user wishes to list it for sale, the system generates a product description and caption, and quickly lists it on an online trading platform. In this process, an example of a prompt message used with the generating AI model would be: "This is a used laptop computer, 3 years old and in good condition. We recommend listing it for 20,000 yen."
[0553] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0554] Step 1:
[0555] The user uses the camera on their portable information terminal to acquire video data of the object they wish to sell. Here, the input is the video data of the object, and the output is the saving of the video data to the terminal. This video data is then prepared to be sent to a server for subsequent processing.
[0556] Step 2:
[0557] The terminal transmits video data to the server. The input is the acquired video data, and the output is the transmission of the image to the server. Specifically, the terminal transfers data to the server using a network connection.
[0558] Step 3:
[0559] The server analyzes the received video data and automatically identifies specific objects using identification AI. The input is the transmitted video data, and the output is information about the identified objects. In this process, the server extracts and recognizes features such as the appearance, shape, and color of the objects.
[0560] Step 4:
[0561] The server uses identified object information to refer to a database and calculate the market selling price. The input is information about the identified object, and the output is the calculated market price. The server calculates the price based on past transaction history.
[0562] Step 5:
[0563] The server sends the calculated market price to the user's terminal. The input is the calculated market price, and the output is the transmission of price information to the terminal. This allows the user to view the price information on a visual display device.
[0564] Step 6:
[0565] The terminal displays the received market price on a visual display and uses a speech synthesis system to suggest listing the item. The input is price information sent from the server, and the output is the displayed price and a voice suggestion. The operation includes a process in which the voice assistant suggests, "Why don't you list this item for sale?"
[0566] Step 7:
[0567] When a user enters their intention to list an item into the terminal, the server generates listing information using a generator. The input is the user's intention to list an item, and the output is the generated listing information. This information includes a product description and image captions.
[0568] Step 8:
[0569] The server transmits the generated listing information to the online trading platform via a communication device. The input is the generated listing information, and the output is the transmission of information to the platform. This allows the user's product to be listed immediately.
[0570] Step 9:
[0571] The terminal notifies the user that the listing is complete. The input is confirmation information that the listing is complete, and the output is a screen display that notifies the user of the completion. Specifically, the terminal displays a message such as "Your item has been listed."
[0572] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0573] This invention combines a system that automatically identifies objects and suggests their market price with an emotion engine that recognizes user emotions, enabling more personalized and effective listing suggestions. The system includes a camera, a terminal, a server, an emotion engine, and an online marketplace application.
[0574] When a user uses the system, they first photograph the object they want to sell using the associated camera on their device. The device then sends the video data to the server. The server receives the data, and an identification AI identifies the object. This AI identifies the object based on its appearance, shape, color, and brand information, and then queries a database to calculate the market price for selling it.
[0575] Simultaneously, the emotion engine recognizes the user's emotional state from their voice tone and facial expressions. This information is used to gain a deeper understanding of the user's intentions and interest in listing items. For example, if the user is excited, the system will recommend listing items in a more positive tone, while if the user is hesitant, the system will adjust to a more cautious approach.
[0576] The terminal receives the calculated market price from the server and displays it to the user via a display device. It also provides voice recommendations, suggesting to the user that they list the item on the online marketplace. In this process, the user's emotional state, recorded by an emotion engine, is taken into consideration, and the content and wording of the recommendations are adjusted accordingly. For example, if the user is in a calm state, the recommendation might say, "The current market price is approximately 8,000 yen. Would you consider listing it?"
[0577] When a user wishes to list an item for sale, they notify the server of their intention via their device. The server uses a generator to create a product description and image captions, adjusting the content and tone to reflect the user's emotional state. The listing information is then sent to the online marketplace application, completing the listing process.
[0578] For example, when a user takes a picture of a used digital camera they have at home, if the emotion engine recognizes the user's happy expression, the system will quickly display a market value and suggest, "This camera is worth 10,000 yen! Let's list it for sale right away!" In this way, by making appropriate suggestions based on the user's emotions, the invention makes the user's selling experience smoother and more intuitive.
[0579] The following describes the processing flow.
[0580] Step 1:
[0581] The user points the camera of the object they intend to sell at the device, using the camera attached to the terminal. The terminal acquires video data of that object.
[0582] Step 2:
[0583] The device compresses the video data it acquires and sends it to a server via the internet. This data is then stored in the cloud.
[0584] Step 3:
[0585] The server analyzes the received video data and uses identification AI to identify objects. The server extracts characteristic information about the objects and determines the product category and brand.
[0586] Step 4:
[0587] The server accesses a database based on past trading history and calculates the market selling price of the recognized object. During this process, the server performs statistical analysis based on the historical data.
[0588] Step 5:
[0589] The emotion engine analyzes the audio and video of the user's surroundings to identify the user's emotional state. This includes voice tone, facial expressions, and other biosignals.
[0590] Step 6:
[0591] The server sends the calculated selling price and the user's emotional state, identified by the emotion engine, to the terminal.
[0592] Step 7:
[0593] The device displays the market price for selling the item on its display screen and initiates voice recommendations that take the user's emotions into account. For example, if the user is excited, it will suggest in an assertive tone, and if they are hesitant, it will suggest in a cautious tone, "Why not list this item for 10,000 yen?"
[0594] Step 8:
[0595] If a user wishes to list an item for sale, they notify the server of their intention to sell via their device. This information is then sent to the server.
[0596] Step 9:
[0597] The server uses a generator to automatically create product descriptions and image captions. The style and tone of the descriptions are adjusted according to the user's emotional state.
[0598] Step 10:
[0599] The server sends the generated listing information to the online marketplace application, completing the listing process.
[0600] Step 11:
[0601] The device notifies the user that the listing is complete and displays a confirmation of success.
[0602] (Example 2)
[0603] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0604] Conventional object recognition and sales systems fail to consider user emotions and interests when suggesting items to sell, resulting in a mechanical and unintuitive listing experience. Furthermore, the inability to provide information reflecting the user's emotional state makes it difficult to suggest the optimal timing and method for listing items. Consequently, these systems fail to attract user interest and promote sales more effectively.
[0605] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0606] In this invention, the server includes means for analyzing video data acquired by a camera and automatically identifying a specific object; means for calculating the market selling price of the identified object by referring to market transaction history based on the information of the identified object; and means for analyzing the user's voice and facial expressions to recognize emotions and adjust the content of the listing suggestions based on the corresponding emotional state. This makes it possible to provide personalized listing suggestions that take into account the user's emotional state in addition to the market value of the object.
[0607] A "photography device" is a device used to acquire image data of an object, and is a device that has the function of recording the appearance of an object in high resolution.
[0608] "Image data" refers to image information acquired by a camera, and is digital data that includes the shape, color, and other characteristics of an object.
[0609] A "server" is a computer system that receives data via a network and performs analysis and processing on it.
[0610] "Identification AI" refers to a technology that uses artificial intelligence algorithms to extract object features from acquired video data and identify objects.
[0611] "Market selling price" refers to the selling price calculated based on past market transaction data for the identified object.
[0612] An "emotion engine" is a technology that analyzes and recognizes a user's emotional state based on their voice tone and facial expressions.
[0613] A "user" refers to an individual or group that uses the system to identify and list an object for sale.
[0614] A "generative AI model" is an algorithm that automatically generates information about an object, and is an artificial intelligence used to create product descriptions and image captions.
[0615] A "prompt statement" is an instruction given to an AI model, serving as a guideline for how to perform a specific task.
[0616] An "online marketplace application" is an internet platform for buying and selling physical objects, providing a space for users to list items for sale.
[0617] A description of the embodiment for carrying out the invention will be provided.
[0618] This system automatically identifies objects owned by the user, displays the market price for the relevant item, and makes listing suggestions based on the user's emotional state. The system consists of a camera, terminal, server, emotion engine, and online marketplace application.
[0619] First, the user operates a terminal and uses a camera to photograph the object they wish to sell. The camera acquires high-definition video data. The terminal sends this video data to a server. The server uses identification AI to analyze the video data and extract the object's features. These features include the object's appearance, shape, color, and brand information.
[0620] The server refers to a database and identifies objects based on extracted features. During this process, it uses past market transaction data to calculate the market selling price of the object. Simultaneously, the terminal collects the user's voice and facial expression data and analyzes it using an emotion engine. The user's emotional state is recognized as joy, excitement, hesitation, etc.
[0621] Based on the analysis results from the emotion engine, the device displays the market selling price to the user and suggests listing items through voice recommendations. At this time, the recommendations are adjusted according to the user's emotions. For example, if the user is excited, the device will suggest, "This camera is worth 10,000 yen! Let's list it for sale right away!"
[0622] When a user wishes to list an item for sale, their device notifies the server of their intention. The server uses a generative AI model to create a detailed product description and image captions, fine-tuning the content according to the user's emotional state. The generated information is sent to the online marketplace application, completing the listing process.
[0623] As a concrete example, if a user takes a picture of a used digital camera at home, and the emotion engine recognizes a happy expression on the device's face, the system will quickly suggest listing the camera while displaying a market price. An example of a prompt message given to the generating AI model would be, "Recognize the user's emotions from their facial expression and tone of voice, and adjust the listing accordingly."
[0624] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0625] Step 1:
[0626] The user picks up the object they wish to sell and activates the camera by operating a terminal. The camera acquires high-resolution video data from the object. This video data becomes the input.
[0627] Step 2:
[0628] The device sends the captured video data to the server. The server passes the input video data to an identification AI, which then begins analyzing the object. During this process, data processing is performed to extract features such as the object's appearance, shape, color, and brand information, and the characteristic information of the object is output.
[0629] Step 3:
[0630] The server queries a database for characteristic information and references the object's market transaction history. Based on this data reference, it calculates the market selling price of the object and outputs the price information.
[0631] Step 4:
[0632] The device records the user's voice and facial expressions using its camera and microphone. This data becomes input and is sent to the emotion engine. The emotion engine recognizes the user's emotions from the voice and facial expressions and outputs the emotional state.
[0633] Step 5:
[0634] The device generates listing suggestions for the user based on the output of the emotion engine (emotional state) and the market selling price received from the server. The content and tone of the suggestions include actions that are adjusted according to the emotional state. For example, a voice recommendation such as "This camera is worth 10,000 yen! Let's list it for sale right away!" might be output.
[0635] Step 6:
[0636] The user inputs their intention to list an item into the terminal. The terminal notifies the server of this intention, and the server calls a generation AI model to generate a product description and image captions. At this time, the intention to list the item and emotional state are input to the generation AI model as prompts, and the listing information is output with the content adjusted accordingly.
[0637] Step 7:
[0638] The server sends the generated listing information to the online marketplace application, completing the listing process. Finally, the listing information is confirmed by the user, and the listing is completed.
[0639] (Application Example 2)
[0640] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0641] Traditional online listing systems focus on object recognition and market price display, but fail to consider user emotions. Therefore, there is a need for methods that alleviate user anxiety and hesitation regarding listing items and promote listings more effectively.
[0642] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0643] In this invention, the server includes means for analyzing video data to recognize and automatically identify objects, means for calculating market selling prices based on the identified objects, and means for analyzing user emotions and adjusting listing suggestions. This enables flexible and effective listing suggestions that respond to user emotions.
[0644] "Means of object recognition" refers to the process of analyzing video data acquired by a camera to automatically identify a specific object.
[0645] "Method for calculating market selling price" refers to the process of calculating the market value of an identified object by referring to a database of past sales history based on information about the object.
[0646] The "method of suggesting items to be listed via voice" is a process that uses voice to recommend to users that they list an item for sale, based on the calculated market price.
[0647] "Methods for analyzing emotions and adjusting listing suggestions" refers to a process that recognizes the user's emotional state from their facial expressions and voice data, and adjusts the content of listing suggestions based on that information.
[0648] An "emotion engine" is a technology that processes facial expressions and voice to analyze a user's emotions and recognize their state.
[0649] The system implementing this invention is realized by having the user photograph an object using a device such as a smartphone, and processing that information on a server. The system mainly uses the following hardware and software.
[0650] The terminal is equipped with a camera to capture video data of the object the user wants to sell. This video data is transmitted to a server via a communication function. On the server, software that executes an object recognition algorithm (e.g., TensorFlow) analyzes the video data and automatically identifies the object. For the identified object, the server refers to a database containing past sales history and calculates the market selling price.
[0651] Furthermore, for emotion analysis, the device acquires the user's facial expressions and voice tone and sends this data to the server. The server uses libraries for facial recognition and voice analysis (e.g., OpenCV, Google Speech-to-Text) to analyze the user's emotional state. The emotion engine generates the analysis results and adjusts the listing suggestions according to the user's emotions.
[0652] Using a generative AI model (e.g., OpenAI GPT), the system generates suggestion messages tailored to the user's emotional state. These suggestions are presented to the user via the device through voice output or text display. For example, if the user expresses positive emotions, a prompt message such as "The current market price is 8,000 yen. Why not list your item now?" might be displayed.
[0653] In this way, the system enables flexible transactions that take into account the emotions of users when listing items, providing users with a more personalized listing experience.
[0654] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0655] Step 1:
[0656] The user takes a picture of the object they want to sell using a smartphone or other device. The input is the image data of the captured object, and the output is the transmission of that image data to the server. Upon the user's action of taking a picture, the device uses its communication function to transfer the image data to the server.
[0657] Step 2:
[0658] The server uses received image data to identify objects using object recognition software (e.g., TensorFlow). The input is image data sent by the user, and the output is the object identification result. The server analyzes the video data and identifies objects by referring to a database based on their appearance, shape, color, etc.
[0659] Step 3:
[0660] The server calculates the market selling price based on the object identification result and by referencing past trading history from the database. The input is the object identification result, and the output is the calculated market selling price. The server searches market data and performs the calculation of the selling price.
[0661] Step 4:
[0662] Simultaneously, the device captures the user's facial expressions with its camera and records voice input with its microphone. The input consists of facial expression data and voice data, and the output is the transmission of this data to the server. This emotion-related data is transferred from the device to the server based on the user's actions.
[0663] Step 5:
[0664] The server analyzes the received emotion data using a facial recognition library (e.g., OpenCV) and a speech analysis library (e.g., Google Speech-to-Text). The input is facial expression data and speech data, and the output is the emotion recognition result. The server processes this data to determine the user's emotional state.
[0665] Step 6:
[0666] The server uses a generative AI model (e.g., OpenAI GPT) to generate listing suggestions based on sentiment recognition results. The input is the sentiment recognition result and the market selling price, and the output is a refined suggestion text. The generative AI model creates prompt text and selects the most suitable suggestion for the user.
[0667] Step 7:
[0668] The terminal displays the generated proposal text to the user and provides audio output. The input is the refined proposal text, and the output is a visual and audio proposal to the user. The terminal uses display and audio functions to communicate the market price and proposal to the user.
[0669] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0670] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0671] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0672] [Fourth Embodiment]
[0673] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0674] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0675] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0676] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0677] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0678] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0679] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0680] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0681] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0682] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0683] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0684] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0685] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0686] The system of this invention enables efficient listing on online marketplaces by allowing users to easily identify objects they own and quickly determine their market value. This system consists of a camera, a terminal, a server, and an online marketplace application.
[0687] When a user uses the system, they first use a camera associated with their device to photograph the object they intend to sell. This camera is typically a built-in camera in a smartphone or AR glasses. The device then acquires image data of the object and transmits it to a server via the internet.
[0688] The server analyzes the received video data. Using identification AI, it identifies objects from the video and extracts their characteristics. This includes the object's appearance, shape, and color. Based on the extracted information, the server queries a database containing past sales history to calculate the current market price for selling the object.
[0689] The device visually displays market price information transmitted from the server on its display device. This display can be the screen of AR glasses or a smartphone. The device also uses voice recommendation functionality to suggest that the user list an item on an online marketplace. This voice assistant might make suggestions such as, "You can sell this item for approximately 5,000 yen. Why not list it?"
[0690] If a user wishes to list an item for sale following the terminal's suggestions, they can indicate their intention to list the item through the terminal's interface. In this case, the server automatically generates listing information for the item using a generator. Specifically, a product description, title, and accompanying image captions are generated. This data is then compiled and sent to the online marketplace application.
[0691] Finally, the server completes the listing process on the online marketplace application. This entire process allows users to sell their items quickly and easily. This system promotes the effective use of users' goods and enhances convenience and sustainability.
[0692] For example, if a user takes a picture of an old smartphone at home, the system will identify the smartphone and indicate its current market value of approximately 15,000 yen. If the user then wishes to list it for sale, the system will quickly complete the listing using automatically generated information. This allows users to maximize their profits while saving time and effort.
[0693] The following describes the processing flow.
[0694] Step 1:
[0695] The user points the object they want to sell at the camera using a camera device associated with the terminal. The terminal acquires video data of this object in real time.
[0696] Step 2:
[0697] The terminal efficiently compresses the video data it acquires and sends it to the server via the internet. The server stores the received data.
[0698] Step 3:
[0699] The server analyzes the received video data and uses identification AI to identify objects. The server then extracts the object's characteristics (e.g., shape, color, brand, etc.).
[0700] Step 4:
[0701] The server uses the extracted object information to access a database that holds past trading history and calculates the market selling price of the identified object. In this process, the server utilizes statistical analysis and machine learning techniques.
[0702] Step 5:
[0703] The server sends the calculated market price information to the terminal. The terminal receives this information and presents it visually to the user through a display device.
[0704] Step 6:
[0705] The device uses its voice output function to offer the user a voice message suggesting they list the item for sale, such as, "This item can currently be sold for approximately 15,000 yen. Would you like to list it?"
[0706] Step 7:
[0707] If a user wishes to list an item for sale, they express their intention via voice command or touch input through their device. The device then notifies the server of this intention.
[0708] Step 8:
[0709] The server uses a generator to automatically produce product descriptions and image captions, thus consolidating the listing information. This information is then prepared for use in online marketplace applications.
[0710] Step 9:
[0711] The server sends the listing information it has configured to the online marketplace application via a communication device, completing the listing process.
[0712] Step 10:
[0713] The device notifies the user that the listing is complete and reports to the user that the listing was successfully completed.
[0714] (Example 1)
[0715] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0716] In today's information society, listing personal belongings on the market quickly and at a fair price remains a complex and time-consuming process. Furthermore, identifying items and calculating market prices requires specialized knowledge and skills, making it difficult for the average consumer to use. Therefore, there is a challenge in how to simplify and efficiently carry out these processes.
[0717] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0718] In this invention, the server includes means for analyzing video data acquired by a shooting means and automatically identifying a specific object using image processing technology; means for calculating the transaction value of the identified object by referring to a recording medium containing past transaction history based on the visual characteristics of the identified object; and means for presenting the calculated value information through a display means and recommending the transaction of the object by voice. As a result, users can easily and quickly list items on the market and trade them at a fair price without having specialized knowledge.
[0719] "Means of shooting" refers to a device or function for acquiring image data of an object, and includes devices such as cameras for taking images or videos.
[0720] "Image processing technology" refers to methods for analyzing acquired video data, and involves the use of algorithms and artificial intelligence technologies to identify objects within images.
[0721] An "object" is an object that a user intends to sell, which is identified by the system and treated as the subject of a transaction.
[0722] A "recording medium" is a data storage system used to hold information such as past transaction history, and includes databases and digital storage devices.
[0723] "Value" refers to the market transaction price calculated based on the identified object, and is estimated based on past transaction data.
[0724] "Display means" refers to a device used to visually communicate calculated value or information to the user, and includes devices such as displays and screens.
[0725] "Voice-based recommendation methods" refer to devices or systems that have the function of using voice to suggest to users how to encourage them to trade an item.
[0726] "Communication means" refers to devices or systems for transferring information over a network, including protocols for internet connectivity and data communication.
[0727] An "e-commerce platform" is a software infrastructure that enables the online trading of goods and is a system for conducting buying and selling transactions.
[0728] This invention is a system for efficiently listing items owned by users on the market, and it functions through the coordinated operation of multiple devices and software. The user photographs the items they wish to sell using a photographic means. This photographic means uses a camera built into a smartphone or augmented reality glasses. The captured video data is acquired by a terminal, which processes the image and transfers it to a server.
[0729] The server analyzes the acquired video data using image processing technology. Here, an identification AI model is used to identify objects and extract their features. The AI model used is created using TensorFlow or PyTorch. The extracted feature information is then compared with past transaction history stored on the recording medium. Based on this, the server calculates the market value of the object.
[0730] The calculated value information is transmitted to the terminal. The terminal uses a smartphone screen or augmented reality glasses display to provide the user with visual information. It also encourages the user to trade the item through voice prompts. For example, it might prompt, "You can sell this item for about 5,000 yen. Why not list it?"
[0731] When a user requests to list an item, the server automatically generates product descriptions and image annotations using a generative AI model, and then constructs the transaction information. This generated transaction information is then transferred to the e-commerce platform via communication. Finally, the server completes the listing process, allowing the user to easily and quickly list their items on the marketplace.
[0732] For example, if a user holds an old smartphone from their home up to the system and takes a picture, the system analyzes the captured data and indicates a market value of approximately 15,000 yen. If the user wishes to list the item for sale, the system automatically generates listing information, enabling a quick transaction. Through this process, users can make optimal transactions while saving time and effort.
[0733] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0734] Step 1:
[0735] The user photographs the items they are considering selling using a photographic device. Specifically, they use the camera built into their smartphone or augmented reality glasses to acquire video data of the objects. This video data is then used for processing within the device.
[0736] Step 2:
[0737] The terminal formats the acquired video data into a predetermined format and sends it to the server via the internet. At this stage, the output is video data formatted in a format that the server can receive.
[0738] Step 3:
[0739] The server inputs the received video data into an identification AI model. It performs image processing for video analysis, extracting features such as the appearance and shape of objects. Through this process, the server identifies what the objects are and obtains feature information of the identified objects as output.
[0740] Step 4:
[0741] The server uses the extracted characteristic information to refer to past transaction history stored on the recording medium. Using database queries, it extracts historical data on transactions of similar objects and calculates the market value based on this data. The output is market price information representing the current market value of the object.
[0742] Step 5:
[0743] The server sends the calculated market price information to the terminal. The terminal visually displays the output market price on the smartphone screen or the display of augmented reality glasses. This allows the user to check the current market price of an object.
[0744] Step 6:
[0745] The device uses a voice assistant to suggest selling an item to the user. Specifically, it communicates the suggestion by voice and prompts the user to take the next action. The output of this step is that a valid selling suggestion is made to the user.
[0746] Step 7:
[0747] If a user wishes to sell based on the terminal's suggestions, they input their intention into the terminal. The terminal transmits this information to the server. Based on this input, the server uses a generative AI model to automatically generate product descriptions and image annotations, creating the listing information. This output contains all the listing information necessary for the transaction.
[0748] Step 8:
[0749] The server transmits the generated listing information to the e-commerce platform via communication means, completing the listing process. This final output means the item is listed on the marketplace and ready for immediate trading.
[0750] (Application Example 1)
[0751] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0752] Current online trading platforms present challenges in quickly and easily listing individual items, requiring users to manually input vast amounts of information and necessitating time-consuming market price searches. As a result, users often find the trading process burdensome, making it difficult to sell items quickly at appropriate market prices. Furthermore, the lack of mechanisms to ensure users can list items at fair prices can delay the listing process, potentially leading to missed sales opportunities.
[0753] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0754] In this invention, the server includes means for analyzing video data acquired by a camera to recognize an object and automatically identifying a specific object; means for calculating the market selling price of the identified object by referring to accumulated information of past transaction history based on the information of the identified object; and means for transmitting the generated listing information to an online trading platform via a communication device and completing the procedure for selling the object. This enables users to quickly grasp market prices and immediately list objects on the online platform at a fair price.
[0755] An "object" is something that the user intends to sell and that is recognized by the camera.
[0756] A "camera" is hardware used to acquire video data, and is mainly installed in mobile devices and other devices with camera functions.
[0757] "Video data" refers to visual information acquired by a camera or camera, and is fundamental information for identifying objects.
[0758] "Automatic identification means" refers to a combination of software or hardware that uses acquired video data to perform a process for identifying a specific object.
[0759] "Market selling price" is an estimate of the current market price of an object, calculated based on past transaction history.
[0760] A "generation device" refers to a software or hardware configuration for automatically generating listing information.
[0761] A "visual display device" is a digital screen or display used to present information to a user visually.
[0762] A "communication device" is a device used to send and receive data via a network connection.
[0763] An "online trading platform" refers to a website or application where goods and services are bought and sold via the internet.
[0764] A "generative AI model" is an artificial intelligence technology used to automatically generate text and information.
[0765] The system of this invention provides an efficient means for users to quickly put their owned objects up for market. The system mainly consists of a camera, a terminal, a server, and an online trading platform.
[0766] First, the user acquires video data of the object they wish to sell using the camera function of their portable information terminal. The camera function of a smartphone or tablet is used as the camera. This video data is then transmitted to a server via the internet through the terminal.
[0767] The server analyzes the received video data using identification AI to automatically identify specific objects. Information about the identified objects is then referenced in a database containing past transaction history to calculate the market price for those objects. This price information is transmitted to the terminal and displayed on the user's device. A speech synthesis system is also utilized, allowing users to receive voice-based listing suggestions.
[0768] When a user indicates their willingness to follow voice suggestions during the listing process, the server generates listing information using a generation device. This generation AI model automatically generates product descriptions and visual annotations. The generated listing information is then transmitted to the online trading platform via a communication device, completing the sale process for the item.
[0769] As a concrete example, let's assume a user takes a picture of a used laptop computer they have on hand using this system. The system recognizes the object and suggests its market price is, for example, 20,000 yen. If the user wishes to list it for sale, the system generates a product description and caption, and quickly lists it on an online trading platform. In this process, an example of a prompt message used with the generating AI model would be: "This is a used laptop computer, 3 years old and in good condition. We recommend listing it for 20,000 yen."
[0770] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0771] Step 1:
[0772] The user uses the camera on their portable information terminal to acquire video data of the object they wish to sell. Here, the input is the video data of the object, and the output is the saving of the video data to the terminal. This video data is then prepared to be sent to a server for subsequent processing.
[0773] Step 2:
[0774] The terminal transmits video data to the server. The input is the acquired video data, and the output is the transmission of the image to the server. Specifically, the terminal transfers data to the server using a network connection.
[0775] Step 3:
[0776] The server analyzes the received video data and automatically identifies specific objects using identification AI. The input is the transmitted video data, and the output is information about the identified objects. In this process, the server extracts and recognizes features such as the appearance, shape, and color of the objects.
[0777] Step 4:
[0778] The server uses identified object information to refer to a database and calculate the market selling price. The input is information about the identified object, and the output is the calculated market price. The server calculates the price based on past transaction history.
[0779] Step 5:
[0780] The server sends the calculated market price to the user's terminal. The input is the calculated market price, and the output is the transmission of price information to the terminal. This allows the user to view the price information on a visual display device.
[0781] Step 6:
[0782] The terminal displays the received market price on a visual display and uses a speech synthesis system to suggest listing the item. The input is price information sent from the server, and the output is the displayed price and a voice suggestion. The operation includes a process in which the voice assistant suggests, "Why don't you list this item for sale?"
[0783] Step 7:
[0784] When a user enters their intention to list an item into the terminal, the server generates listing information using a generator. The input is the user's intention to list an item, and the output is the generated listing information. This information includes a product description and image captions.
[0785] Step 8:
[0786] The server transmits the generated listing information to the online trading platform via a communication device. The input is the generated listing information, and the output is the transmission of information to the platform. This allows the user's product to be listed immediately.
[0787] Step 9:
[0788] The terminal notifies the user that the listing is complete. The input is confirmation information that the listing is complete, and the output is a screen display that notifies the user of the completion. Specifically, the terminal displays a message such as "Your item has been listed."
[0789] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0790] This invention combines a system that automatically identifies objects and suggests their market price with an emotion engine that recognizes user emotions, enabling more personalized and effective listing suggestions. The system includes a camera, a terminal, a server, an emotion engine, and an online marketplace application.
[0791] When a user uses the system, they first photograph the object they want to sell using the associated camera on their device. The device then sends the video data to the server. The server receives the data, and an identification AI identifies the object. This AI identifies the object based on its appearance, shape, color, and brand information, and then queries a database to calculate the market price for selling it.
[0792] Simultaneously, the emotion engine recognizes the user's emotional state from their voice tone and facial expressions. This information is used to gain a deeper understanding of the user's intentions and interest in listing items. For example, if the user is excited, the system will recommend listing items in a more positive tone, while if the user is hesitant, the system will adjust to a more cautious approach.
[0793] The terminal receives the calculated market price from the server and displays it to the user via a display device. It also provides voice recommendations, suggesting to the user that they list the item on the online marketplace. In this process, the user's emotional state, recorded by an emotion engine, is taken into consideration, and the content and wording of the recommendations are adjusted accordingly. For example, if the user is in a calm state, the recommendation might say, "The current market price is approximately 8,000 yen. Would you consider listing it?"
[0794] When a user wishes to list an item for sale, they notify the server of their intention via their device. The server uses a generator to create a product description and image captions, adjusting the content and tone to reflect the user's emotional state. The listing information is then sent to the online marketplace application, completing the listing process.
[0795] For example, when a user takes a picture of a used digital camera they have at home, if the emotion engine recognizes the user's happy expression, the system will quickly display a market value and suggest, "This camera is worth 10,000 yen! Let's list it for sale right away!" In this way, by making appropriate suggestions based on the user's emotions, the invention makes the user's selling experience smoother and more intuitive.
[0796] The following describes the processing flow.
[0797] Step 1:
[0798] The user points the camera of the object they intend to sell at the device, using the camera attached to the terminal. The terminal acquires video data of that object.
[0799] Step 2:
[0800] The device compresses the video data it acquires and sends it to a server via the internet. This data is then stored in the cloud.
[0801] Step 3:
[0802] The server analyzes the received video data and uses identification AI to identify objects. The server extracts characteristic information about the objects and determines the product category and brand.
[0803] Step 4:
[0804] The server accesses a database based on past trading history and calculates the market selling price of the recognized object. During this process, the server performs statistical analysis based on the historical data.
[0805] Step 5:
[0806] The emotion engine analyzes the audio and video of the user's surroundings to identify the user's emotional state. This includes voice tone, facial expressions, and other biosignals.
[0807] Step 6:
[0808] The server sends the calculated selling price and the user's emotional state, identified by the emotion engine, to the terminal.
[0809] Step 7:
[0810] The device displays the market price for selling the item on its display screen and initiates voice recommendations that take the user's emotions into account. For example, if the user is excited, it will suggest in an assertive tone, and if they are hesitant, it will suggest in a cautious tone, "Why not list this item for 10,000 yen?"
[0811] Step 8:
[0812] If a user wishes to list an item for sale, they notify the server of their intention to sell via their device. This information is then sent to the server.
[0813] Step 9:
[0814] The server uses a generator to automatically create product descriptions and image captions. The style and tone of the descriptions are adjusted according to the user's emotional state.
[0815] Step 10:
[0816] The server sends the generated listing information to the online marketplace application, completing the listing process.
[0817] Step 11:
[0818] The device notifies the user that the listing is complete and displays a confirmation of success.
[0819] (Example 2)
[0820] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0821] Conventional object recognition and sales systems fail to consider user emotions and interests when suggesting items to sell, resulting in a mechanical and unintuitive listing experience. Furthermore, the inability to provide information reflecting the user's emotional state makes it difficult to suggest the optimal timing and method for listing items. Consequently, these systems fail to attract user interest and promote sales more effectively.
[0822] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0823] In this invention, the server includes means for analyzing video data acquired by a camera and automatically identifying a specific object; means for calculating the market selling price of the identified object by referring to market transaction history based on the information of the identified object; and means for analyzing the user's voice and facial expressions to recognize emotions and adjust the content of the listing suggestions based on the corresponding emotional state. This makes it possible to provide personalized listing suggestions that take into account the user's emotional state in addition to the market value of the object.
[0824] A "photography device" is a device used to acquire image data of an object, and is a device that has the function of recording the appearance of an object in high resolution.
[0825] "Image data" refers to image information acquired by a camera, and is digital data that includes the shape, color, and other characteristics of an object.
[0826] A "server" is a computer system that receives data via a network and performs analysis and processing on it.
[0827] "Identification AI" refers to a technology that uses artificial intelligence algorithms to extract object features from acquired video data and identify objects.
[0828] "Market selling price" refers to the selling price calculated based on past market transaction data for the identified object.
[0829] An "emotion engine" is a technology that analyzes and recognizes a user's emotional state based on their voice tone and facial expressions.
[0830] A "user" refers to an individual or group that uses the system to identify and list an object for sale.
[0831] A "generative AI model" is an algorithm that automatically generates information about an object, and is an artificial intelligence used to create product descriptions and image captions.
[0832] A "prompt statement" is an instruction given to an AI model, serving as a guideline for how to perform a specific task.
[0833] An "online marketplace application" is an internet platform for buying and selling physical objects, providing a space for users to list items for sale.
[0834] A description of the embodiment for carrying out the invention will be provided.
[0835] This system automatically identifies objects owned by the user, displays the market price for the relevant item, and makes listing suggestions based on the user's emotional state. The system consists of a camera, terminal, server, emotion engine, and online marketplace application.
[0836] First, the user operates a terminal and uses a camera to photograph the object they wish to sell. The camera acquires high-definition video data. The terminal sends this video data to a server. The server uses identification AI to analyze the video data and extract the object's features. These features include the object's appearance, shape, color, and brand information.
[0837] The server refers to a database and identifies objects based on extracted features. During this process, it uses past market transaction data to calculate the market selling price of the object. Simultaneously, the terminal collects the user's voice and facial expression data and analyzes it using an emotion engine. The user's emotional state is recognized as joy, excitement, hesitation, etc.
[0838] Based on the analysis results from the emotion engine, the device displays the market selling price to the user and suggests listing items through voice recommendations. At this time, the recommendations are adjusted according to the user's emotions. For example, if the user is excited, the device will suggest, "This camera is worth 10,000 yen! Let's list it for sale right away!"
[0839] When a user wishes to list an item for sale, their device notifies the server of their intention. The server uses a generative AI model to create a detailed product description and image captions, fine-tuning the content according to the user's emotional state. The generated information is sent to the online marketplace application, completing the listing process.
[0840] As a concrete example, if a user takes a picture of a used digital camera at home, and the emotion engine recognizes a happy expression on the device's face, the system will quickly suggest listing the camera while displaying a market price. An example of a prompt message given to the generating AI model would be, "Recognize the user's emotions from their facial expression and tone of voice, and adjust the listing accordingly."
[0841] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0842] Step 1:
[0843] The user picks up the object they wish to sell and activates the camera by operating a terminal. The camera acquires high-resolution video data from the object. This video data becomes the input.
[0844] Step 2:
[0845] The device sends the captured video data to the server. The server passes the input video data to an identification AI, which then begins analyzing the object. During this process, data processing is performed to extract features such as the object's appearance, shape, color, and brand information, and the characteristic information of the object is output.
[0846] Step 3:
[0847] The server queries a database for characteristic information and references the object's market transaction history. Based on this data reference, it calculates the market selling price of the object and outputs the price information.
[0848] Step 4:
[0849] The device records the user's voice and facial expressions using its camera and microphone. This data becomes input and is sent to the emotion engine. The emotion engine recognizes the user's emotions from the voice and facial expressions and outputs the emotional state.
[0850] Step 5:
[0851] The device generates listing suggestions for the user based on the output of the emotion engine (emotional state) and the market selling price received from the server. The content and tone of the suggestions include actions that are adjusted according to the emotional state. For example, a voice recommendation such as "This camera is worth 10,000 yen! Let's list it for sale right away!" might be output.
[0852] Step 6:
[0853] The user inputs their intention to list an item into the terminal. The terminal notifies the server of this intention, and the server calls a generation AI model to generate a product description and image captions. At this time, the intention to list the item and emotional state are input to the generation AI model as prompts, and the listing information is output with the content adjusted accordingly.
[0854] Step 7:
[0855] The server sends the generated listing information to the online marketplace application, completing the listing process. Finally, the listing information is confirmed by the user, and the listing is completed.
[0856] (Application Example 2)
[0857] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0858] Traditional online listing systems focus on object recognition and market price display, but fail to consider user emotions. Therefore, there is a need for methods that alleviate user anxiety and hesitation regarding listing items and promote listings more effectively.
[0859] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0860] In this invention, the server includes means for analyzing video data to recognize and automatically identify objects, means for calculating market selling prices based on the identified objects, and means for analyzing user emotions and adjusting listing suggestions. This enables flexible and effective listing suggestions that respond to user emotions.
[0861] "Means of object recognition" refers to the process of analyzing video data acquired by a camera to automatically identify a specific object.
[0862] "Method for calculating market selling price" refers to the process of calculating the market value of an identified object by referring to a database of past sales history based on information about the object.
[0863] The "method of suggesting items to be listed via voice" is a process that uses voice to recommend to users that they list an item for sale, based on the calculated market price.
[0864] "Methods for analyzing emotions and adjusting listing suggestions" refers to a process that recognizes the user's emotional state from their facial expressions and voice data, and adjusts the content of listing suggestions based on that information.
[0865] An "emotion engine" is a technology that processes facial expressions and voice to analyze a user's emotions and recognize their state.
[0866] The system implementing this invention is realized by having the user photograph an object using a device such as a smartphone, and processing that information on a server. The system mainly uses the following hardware and software.
[0867] The terminal is equipped with a camera to capture video data of the object the user wants to sell. This video data is transmitted to a server via a communication function. On the server, software that executes an object recognition algorithm (e.g., TensorFlow) analyzes the video data and automatically identifies the object. For the identified object, the server refers to a database containing past sales history and calculates the market selling price.
[0868] Furthermore, for emotion analysis, the device acquires the user's facial expressions and voice tone and sends this data to the server. The server uses libraries for facial recognition and voice analysis (e.g., OpenCV, Google Speech-to-Text) to analyze the user's emotional state. The emotion engine generates the analysis results and adjusts the listing suggestions according to the user's emotions.
[0869] Using a generative AI model (e.g., OpenAI GPT), the system generates suggestion messages tailored to the user's emotional state. These suggestions are presented to the user via the device through voice output or text display. For example, if the user expresses positive emotions, a prompt message such as "The current market price is 8,000 yen. Why not list your item now?" might be displayed.
[0870] In this way, the system enables flexible transactions that take into account the emotions of users when listing items, providing users with a more personalized listing experience.
[0871] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0872] Step 1:
[0873] The user takes a picture of the object they want to sell using a smartphone or other device. The input is the image data of the captured object, and the output is the transmission of that image data to the server. Upon the user's action of taking a picture, the device uses its communication function to transfer the image data to the server.
[0874] Step 2:
[0875] The server uses received image data to identify objects using object recognition software (e.g., TensorFlow). The input is image data sent by the user, and the output is the object identification result. The server analyzes the video data and identifies objects by referring to a database based on their appearance, shape, color, etc.
[0876] Step 3:
[0877] The server calculates the market selling price based on the object identification result and by referencing past trading history from the database. The input is the object identification result, and the output is the calculated market selling price. The server searches market data and performs the calculation of the selling price.
[0878] Step 4:
[0879] Simultaneously, the device captures the user's facial expressions with its camera and records voice input with its microphone. The input consists of facial expression data and voice data, and the output is the transmission of this data to the server. This emotion-related data is transferred from the device to the server based on the user's actions.
[0880] Step 5:
[0881] The server analyzes the received emotion data using a facial recognition library (e.g., OpenCV) and a speech analysis library (e.g., Google Speech-to-Text). The input is facial expression data and speech data, and the output is the emotion recognition result. The server processes this data to determine the user's emotional state.
[0882] Step 6:
[0883] The server uses a generative AI model (e.g., OpenAI GPT) to generate listing suggestions based on sentiment recognition results. The input is the sentiment recognition result and the market selling price, and the output is a refined suggestion text. The generative AI model creates prompt text and selects the most suitable suggestion for the user.
[0884] Step 7:
[0885] The terminal displays the generated proposal text to the user and provides audio output. The input is the refined proposal text, and the output is a visual and audio proposal to the user. The terminal uses display and audio functions to communicate the market price and proposal to the user.
[0886] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0887] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0888] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0889] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0890] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0891] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0892] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0893] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0894] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0895] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0896] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0897] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0898] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0899] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0900] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0901] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0902] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0903] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0904] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0905] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0906] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0907] The following is further disclosed regarding the embodiments described above.
[0908] (Claim 1)
[0909] In order to recognize an object, a means for analyzing video data acquired by a camera and automatically identifying a specific object,
[0910] A means for calculating the market selling price of an identified object by referring to a database of past sales history based on information about the identified object,
[0911] The calculated selling price is displayed on a display device, and a means is provided to suggest listing the item for sale via voice.
[0912] A system that includes this.
[0913] (Claim 2)
[0914] The system according to claim 1, comprising means for automatically creating a product description and image caption using a generation device in order to generate listing information for an object based on the user's expression of intent.
[0915] (Claim 3)
[0916] The system according to claim 1, comprising means for transmitting generated listing information to an online marketplace application via a communication device to complete the procedure for listing the object.
[0917] "Example 1"
[0918] (Claim 1)
[0919] A means for analyzing video data acquired by a shooting means and automatically identifying a specific object using image processing technology,
[0920] A means for calculating the transaction value of an identified object by referring to a recording medium containing past transaction history based on the visual characteristics of the identified object,
[0921] The calculated value information is presented through a display means, and the transaction of the object is recommended by voice.
[0922] A system that includes this.
[0923] (Claim 2)
[0924] The system according to claim 1, comprising means for automatically generating descriptive text and image annotations using an information creation device in order to create transaction information of an object in accordance with the user's decision.
[0925] (Claim 3)
[0926] The system according to claim 1, comprising means for transferring generated transaction information to an e-commerce platform via communication means and carrying out the procedure for listing the item on the market.
[0927] "Application Example 1"
[0928] (Claim 1)
[0929] In order to recognize an object, a means for analyzing video data acquired by a camera and automatically identifying a specific object,
[0930] A means for calculating the market selling price of an identified object by referring to accumulated information on past transaction history based on information about the identified object,
[0931] A visual display device that shows the calculated market price for sale, and a means for audibly suggesting the listing of an object using an information presentation means,
[0932] In order to generate information about the listing based on the user's expression of intent, a means for automatically creating product descriptions and visual information annotations using a generation device,
[0933] A means of transmitting the generated listing information to an online trading platform via a communication device and completing the procedure for selling the item,
[0934] A system that includes this.
[0935] (Claim 2)
[0936] The system according to claim 1, which includes means for displaying the market price of generated listing information in real time on the user's portable information terminal to facilitate the listing procedure.
[0937] (Claim 3)
[0938] The system according to claim 1, comprising means for proposing the sale of a specific object in real time and for automatically reconstructing listing information using a generated AI model based on the presented information.
[0939] "Example 2 of combining an emotion engine"
[0940] (Claim 1)
[0941] A means for analyzing video data acquired by a camera and automatically identifying a specific object,
[0942] A means for calculating the market selling price of an identified object by referring to its transaction history based on information about the identified object,
[0943] A means of analyzing the user's voice and facial expressions to recognize emotions and adjusting the content of the listing suggestions based on the corresponding emotional state,
[0944] A means of displaying the calculated selling price on a display device and suggesting the listing of an item by voice,
[0945] A system that includes this.
[0946] (Claim 2)
[0947] The system according to claim 1, comprising means for automatically creating product descriptions and image captions using a generation device in order to generate listing information for an object based on the user's expression of intent, and adjusting them according to the user's emotional state.
[0948] (Claim 3)
[0949] The system according to claim 1, comprising means for transmitting generated listing information to an online marketplace application via a communication device to complete the procedure for listing the object.
[0950] "Application example 2 when combining with an emotional engine"
[0951] (Claim 1)
[0952] In order to recognize an object, a means for analyzing video data acquired by a camera and automatically identifying a specific object,
[0953] A means for calculating the market selling price of an identified object by referring to a database of past sales history based on information about the identified object,
[0954] The calculated selling price is displayed on a display device, and a means is provided to suggest listing the item for sale via voice.
[0955] A means to analyze user emotions and adjust listing suggestions according to their emotional state,
[0956] A means equipped with an emotion engine for recognizing emotions from the user's facial expressions and voice data,
[0957] A system that includes this.
[0958] (Claim 2)
[0959] The system according to claim 1, which includes means for automatically creating product descriptions and image captions using a generation device in order to generate listing information for an object based on the user's expression of intent, and for adjusting the wording and tone according to the user's emotional state.
[0960] (Claim 3)
[0961] The system according to claim 1, comprising means for transmitting generated listing information to an online marketplace application via a communication device to complete the procedure for listing the object. [Explanation of Symbols]
[0962] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. In order to recognize an object, a means for analyzing video data acquired by a camera and automatically identifying a specific object, A means for calculating the market selling price of an identified object by referring to a database of past sales history based on information about the identified object, The calculated selling price is displayed on a display device, and a means is provided to suggest listing the item for sale via voice. A system that includes this.
2. The system according to claim 1, comprising means for automatically creating a product description and image caption using a generation device in order to generate listing information for an object based on the user's expression of intent.
3. The system according to claim 1, comprising means for transmitting generated listing information to an online marketplace application via a communication device to complete the procedure for listing the object.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A