Information processing device

The information processing device automates the classification and generation of unified product names using clustering and language models, addressing the inefficiencies of manual product registration in retail stores, thereby improving data management efficiency and accuracy.

JP2026058028APending Publication Date: 2026-04-03TOSHIBA TEC KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

The manual process of registering new products in a retail store's product master system is labor-intensive, requiring manual arrangement of product names and codes, which is inefficient and time-consuming.

Method used

An information processing device with an identification unit and a name generation unit that automates the classification and generation of unified product names using methods such as clustering and large-scale language models to streamline the product registration process.

Benefits of technology

Automated generation of unified product names reduces labor and time required for product registration, enhancing efficiency and accuracy in managing product data across multiple stores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026058028000001_ABST
    Figure 2026058028000001_ABST
Patent Text Reader

Abstract

This enables the generation of product names and other information in accordance with appropriate policies. [Solution] The information processing device comprises an identification unit and a name generation unit. The identification unit classifies the information contained in multiple POS data collected from multiple stores. For each cluster classified by the identification unit, the name generation unit generates a unified product name, which is a representative name for the product names contained in the information. In one embodiment, the name generation unit decomposes the product names contained in each piece of information in the cluster into individual words and uses the word with the most matches as the unified product name.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to an information processing apparatus.

Background Art

[0002] Normally, in a store that sells goods and services (hereinafter collectively referred to as "goods"), a receipt is issued for evidence that the purchaser has received money as the price of the goods. The receipt describes "when", "where", "what", "how many", and "how much" the goods were sold. At the time of product registration, all the operator enters into the POS (Point Of Sales) terminal are only the product code and the number of products. Other information such as the product name and the selling price are referred to from the product master held by the POS terminal or the backbone system (store server) (see, for example, Patent Document 1). The product master of a retail store includes not only the information necessary for the receipt but also the information necessary for operation such as the product category to which the product belongs, the "quantity" such as how many in a pack, and the "capacity" such as how many mL.

[0003] When a retail store starts handling a new product, the person in charge of management of the retail store registers the information of the new product in the product master.

[0004] When registering the information of this new product in the product master, the person in charge of management arranges the information according to an appropriate policy. For example, the person in charge of management determines the product name (hereinafter referred to as "unified product name") that is uniformly used in the retail store and associates the information with the product code. Since this process is all done manually, a great deal of labor has been generated.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] The problem that the embodiments of the present invention aim to solve is to provide an information processing device capable of generating information such as product names in accordance with an appropriate policy. [Means for solving the problem]

[0007] In one embodiment, the information processing device comprises an identification unit and a name generation unit. The identification unit classifies the information contained in multiple POS data collected from multiple stores. For each cluster classified by the identification unit, the name generation unit generates a unified product name, which is a representative name for the product names contained in the information. In one form, the name generation unit decomposes the product names contained in each piece of information in the cluster into individual words and uses the word with the most matches as the unified product name. In another form, the name generation unit uses the product name with the most characters from among the product names contained in each piece of information in the cluster as the unified product name. In yet another form, the name generation unit generates a unified product name by processing the product names contained in each piece of information in the cluster using a large-scale language model. [Brief explanation of the drawing]

[0008] [Figure 1] Figure 1 is a block diagram showing the overall configuration of a product master creation system including an information processing device according to one embodiment. [Figure 2] Figure 2 is a block diagram showing the main circuit configuration and program overview of a store server as an information processing device according to one embodiment. [Figure 3] Figure 3 is a sequence diagram showing an overview of the operation of the product master creation system. [Figure 4] Figure 4 is a sequence diagram showing an overview of the operation of the product master creation system. [Figure 5] Figure 5 is a flowchart illustrating the main information processing steps performed by the store server's processor. [Figure 6] Figure 6 is a schematic diagram showing an example of classification using the first classification method. [Figure 7] Figure 7 is a flowchart showing the detailed steps of the classification performed in ACT2 in Figure 5, using the second classification method for classification when silhouette analysis is employed. [Figure 8] Figure 8 is a schematic diagram showing an example of classification using the second classification method. [Figure 9] Figure 9 illustrates the method for determining the number of clusters using silhouette coefficients. [Figure 10] Figure 10 is a diagram illustrating the first name generation method. [Figure 11] Figure 11 shows an example of name generation using the first name generation method. [Figure 12] Figure 12 is a diagram illustrating the second name generation method. [Figure 13] Figure 13 shows an example of name generation using the second name generation method. [Figure 14] Figure 14 is a diagram illustrating the third name generation method. [Figure 15] Figure 15 shows an example of name generation using the third name generation method. [Figure 16] Figure 16 is a diagram illustrating the fourth name generation method. [Figure 17] Figure 17 is a schematic diagram illustrating an example of name determination in the fourth name generation method. [Figure 18] Figure 18 shows an example of an instruction statement in the fifth name generation method. [Figure 19] Figure 19 shows an example of a question in the fifth name generation method. [Figure 20] Figure 20 shows an example used for One-Shot learning in the fifth name generation method. [Figure 21] Figure 21 shows an example used for few-shot learning in the fifth name generation method. [Figure 22] Figure 22 shows an example of post-processing. [Modes for carrying out the invention]

[0009] Hereinafter, embodiments will be described with reference to the drawings.

[0010] FIG. 1 is a block diagram showing the overall configuration of a product master creation system including an information processing apparatus according to an embodiment. The product master creation system includes a plurality of store servers 1 arranged in each of a plurality of stores S, and a data collection server 2 arranged in a POS data collection center C. The store server 1 is an example of an information processing apparatus according to an embodiment.

[0011] Each store S includes, in addition to the store server 1, one or more POS terminals P and a manager terminal SM. The store server 1, the POS terminal P, and the manager terminal SM are communicably connected via an in-store network such as a LAN (Local Area Network). Also, the store server 1 can be connected to a network NW that connects the store S and the POS data collection center C via an external communication interface such as a router connected to the in-store network. Here, the network NW is, for example, the Internet or various public networks.

[0012] The POS terminal P is a fixed terminal arranged at a face-to-face cash register or self-checkout where a cashier who is a store clerk in charge of accounting or a shopper who is a buyer himself / herself registers purchased goods and conducts settlement. The POS terminal P calculates the settlement amount of the registered purchased goods, and the shopper pays the amount equivalent to the settlement amount. The payment can be made with cash, credit card, electronic money, points, gift certificates, or other vouchers. The POS terminal P is an example of a sales data processing apparatus that processes the registration and accounting of purchased goods. When the accounting is completed, the POS terminal P generates sales data including transaction data representing the transaction details of the goods to be settled, settlement data related to the settlement of the transaction, etc., and issues a paper receipt by printing the sales data with a printer. Note that the number of POS terminals P provided in the store S is not particularly limited.

[0013] The administrator terminal SM is the user interface terminal for store server 1. The person in charge of store S and the administrator who manages the systems within store S operate the administrator terminal SM to input various commands to store server 1, and the administrator terminal SM is used to check the processing results on store server 1. There is no particular limit to the number of administrator terminal SMs that can be placed in store S; there may be multiple units.

[0014] Store server 1 centrally manages the sales data processing performed by POS terminal P. For example, store server 1 stores and manages the sales data generated by the sales data processing of POS terminal P, and also stores and manages the product master data used in sales data processing.

[0015] The data collection server 2 is located at the POS data collection center C. The POS data collection center C is a third-party organization separate from the companies of each store S. For example, it contracts with the companies of each store S to acquire sales data managed by the store server 1, and then creates and manages POS data by removing personal information such as customer (member) names and member codes from the acquired sales data.

[0016] The data collection server 2 creates POS data by removing personal information from the sales data transmitted from the store server 1 and stores it in the POS database (database is abbreviated as "DB" in the diagram) 21. The data collection server 2 provides the POS data stored in the POS database 21 to the store server 1 of the store S, and to servers operated by companies such as product manufacturers (not shown in the diagram).

[0017] Figure 2 is a block diagram showing the main circuit configuration and program overview of the store server 1. The store server 1 comprises a processor 11, main memory 12, auxiliary storage device 13, network interface 14, and system transmission path 15. The system transmission path 15 includes an address bus, data bus, control signal lines, etc. The store server 1 connects the processor 11, main memory 12, auxiliary storage device 13, and network interface 14 to the system transmission path 15. The store server 1 constitutes a computer with the processor 11, main memory 12, auxiliary storage device 13, and the system transmission path 15 connecting them.

[0018] The processor 11 corresponds to the central part of the computer described above. The processor 11 controls various parts to realize various functions as a store server 1 according to the operating system and application programs. The processor 11 is, for example, a CPU (Central Processing Unit), but is not limited to this. The processor 11 may be multi-core / multi-threaded and capable of executing multiple processes in parallel. The processor 11 may also be an MPU (Micro Processing Unit). Furthermore, the processor 11 may be implemented in various other forms, including integrated circuits such as ASIC (Application Specific Integrated Circuit), GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), DSP (Digital Signal Processor), SoC (System on a Chip), and PLD (Programmable Logic Device). The processor 11 may also be a combination of several of these.

[0019] Main memory 12 corresponds to the main memory portion of the computer described above. Main memory 12 includes non-volatile memory areas and volatile memory areas. In the non-volatile memory area, main memory 12 stores the operating system and application programs. Main memory 12 may also store data necessary for the processor 11 to perform processing to control various parts in non-volatile or volatile memory areas. Main memory 12 uses the volatile memory area as a work area where data is rewritten as needed by the processor 11. For example, the non-volatile memory area is ROM (Read Only Memory). The volatile memory area is RAM (Random Access Memory).

[0020] The auxiliary storage device 13 corresponds to the auxiliary storage portion of the computer described above. For example, an EEPROM (Electric Erasable Programmable Read-Only Memory), an HDD (Hard Disk Drive), or an SSD (Solid State Drive) may be used as the auxiliary storage device 13. The auxiliary storage device 13 stores data used by the processor 11 for various processing tasks, as well as data created by the processing performed by the processor 11. The auxiliary storage device 13 may also store the application programs described above.

[0021] The network interface 14 transmits and receives data to and from POS terminals P and administrator terminals SM connected via the store network NWL, such as a LAN, in accordance with the communication protocol. Furthermore, the network interface 14 transmits and receives data to and from data collection server 2, connected via the network NW, through an external communication interface INT, such as a router, connected to the store network NWL, in accordance with the communication protocol.

[0022] The store server 1 includes a product master database 131, a store member master database 132, a sales data database 133, and a POS database 134 on an auxiliary storage device 13.

[0023] The product master database 131 is updated at any time as needed by the store administrator via an update operation from the administrator terminal SM. This product master database 131 stores information about each product sold at the store. The product master database 131 stores data such as product code, product name, price, and category name, with each product as one record. The product code is a unique code, such as a JAN code, set for each product to identify it individually. The information in the product master database 131 may also be sent to each POS terminal P at each store S when the terminal is started, and stored as a product data file on the POS terminal P for use.

[0024] The store member master database 132 stores data about customers who have registered as members, such as point members. The store member master database has one record for each member, and for example, it includes data such as the member ID. It may also include personal information such as the member's gender, age, name, address, and contact information. The member ID is a unique code set for each member to identify them individually. Members possess a recording medium on which their member ID is recorded. The recording medium may be, for example, a magnetic card, a contact IC (Integrated Circuit) card, a contactless IC card, or a smartphone. Purchasers can enjoy the benefits of being a member by entering their member code into the POS terminal P at some point before completing the transaction at the POS terminal P. There are no particular restrictions on the method of entering the member code, and methods such as reading from a medium using a scanner or reader by the cashier or the purchaser can be used.

[0025] The sales data database 133 stores sales data generated for each transaction, which is transmitted from each POS terminal P at the end of the transaction. The sales data database 133 stores transaction data representing the transaction details of the goods to be paid for and payment data related to the settlement of that transaction, with each transaction being treated as one record. The transaction data includes the product name (product code) and price of the goods to be paid for, and the total amount of each product. The payment data includes the amount paid by the buyer and the change amount if payment was made in cash, and information on the payment method and result if payment was made by credit card, electronic money, points, gift certificates, etc. In addition, the sales data may include elements such as member ID, company code (company name) indicating the company that operates the store, store code (store name), store telephone number, address, transaction date and time, transaction number (receipt number), register number, and name of person in charge. Note that the elements included in the sales data are not limited to these examples and may include other elements such as industry / business type codes, or elements such as telephone number and address may be removed. POS terminal P, for example, reads a code symbol printed or attached to a product using a scanner to obtain a product code, and for the product identified by that product code, it refers to the product master database 131 on the store server 1, or reads product data from a product data file stored on the terminal itself, and registers the product data as transaction data, thereby updating the sales data, which includes transaction data and settlement data. When the accounting process is completed, that is, when a transaction is completed, POS terminal P prints the sales data on paper to issue a paper receipt for that transaction, and also sends the sales data for that transaction to the store server 1. The sales data database 133 stores the sales data thus sent from each POS terminal P.

[0026] The POS database 134 is a copy of the POS database 21 located on the data collection server 2, and updates and stores the data from the POS database 21 that has been sent from the data collection server 2, for example, through nightly batch processing.

[0027] Although not specifically illustrated, the auxiliary storage device 13 may also include an electronic receipt member master database. The electronic receipt member master database stores data about members who have registered for the electronic receipt service. The electronic receipt member master database includes data such as the member ID, with each member being treated as one record. The member ID is a unique code set for each member to individually identify each member of the electronic receipt service. Members possess an information terminal such as a smartphone on which their member ID is recorded. The data stored in this electronic receipt member master database can be provided from an electronic receipt server (not shown) connected via a network NW. Members of the electronic receipt service who wish to receive an electronic receipt instead of a paper receipt can enter their member code into the POS terminal P at any point before performing payment at the POS terminal P, allowing the POS terminal P to issue an electronic receipt instead of a paper receipt.

[0028] Next, we will describe the various components implemented in the processor 11 of the store server 1. The processor 11 implements, for example, a sales data acquisition unit 111, a sales data transmission unit 112, a POS data reception unit 113, and a product master generation unit 114. Each component implemented in the processor 11 can also be called a functional module. Each component implemented in the processor 11 can also be called a program implemented in the control unit, which includes the processor 11 and the main memory 12.

[0029] The sales data acquisition unit 111 receives sales data for one transaction transmitted from the POS terminal P at the end of each transaction via the network interface 14. The sales data acquisition unit 111 stores this received sales data in the sales data database 133.

[0030] The sales data transmission unit 112, for example at night, uses batch processing to send a day's worth of sales data stored in the sales data database 133 to the data collection server 2 via the external communication interface INT and the network NW. The data collection server 2 creates POS data from the sales data transmitted from each store server 1 and stores it in the POS database 21.

[0031] The POS data receiving unit 113 receives POS data stored in the POS database 21, which is transmitted from the data collection server 2 via batch processing, for example at night, and updates the contents of the POS database 134 with this received POS data.

[0032] The product master generation unit 114 generates data such as the product code, product name, price, and classification name of the new product when a store S handles a new product, and registers this data as a single record in the product master database 131. The product master generation unit 114 implements a pre-processing unit 1141, a class classification unit 1142, a name generation unit 1143, a post-processing unit 1144, and a master registration unit 1145.

[0033] The preprocessing unit 1141 extracts information including product names from the POS data stored in the POS database 134 and supplies the extracted information to the class classification unit. As mentioned above, the POS data includes transaction data such as the product name, product code, amount, and total amount of the products subject to payment; payment data such as the amount received, the amount of change, or information on the payment method and result; and elements such as the company code, store code, store telephone number, address, transaction date and time, transaction number, register number, and name of the person in charge. The preprocessing unit 1141 extracts, for example, the product name, product code, and company code or store code from the POS data. Note that the preprocessing unit 1141 does not necessarily target all the POS data stored in the POS database 134, but may, for example, extract information only from the POS data of products with product codes that match the product code of the target product to be newly registered in the product master database 131.

[0034] The class classification unit 1142 classifies the information extracted by the preprocessing unit 1141. The class classification unit 1142 is an example of an identification unit that classifies information contained in multiple POS data collected from multiple stores. The detailed operation of this class classification unit 1142 will be described later.

[0035] The name generation unit 1143 generates a unified product name, which is a representative name for the product names included in the information extracted by the preprocessing unit 1141, for each cluster classified by the class classification unit 1142. The detailed operation of this name generation unit 1143 will be described later.

[0036] The post-processing unit 1144 converts the information containing the unified product name generated by the name generation unit 1143 into information to be registered in the product master database 131. Specifically, the post-processing unit 1144 removes duplicate information and data that does not need to be registered in the product master database 131 from the information containing the unified product name generated by the name generation unit 1143. The post-processing unit 1144 then adds data such as the price and classification name of newly handled products, which have been separately entered from the administrator terminal SM, to this deleted data to generate information to be registered in the product master database 131.

[0037] The master registration unit 1145 registers the information generated by the post-processing unit 1144 into the product master database 131.

[0038] The store server 1 can be implemented, for example, by using a general-purpose computer device for servers as hardware and writing the application program to the main memory 12 or auxiliary storage device 13. The application program may be stored in the main memory 12 or auxiliary storage device 13 when the store server 1 is transferred, or it may be transferred separately from the general-purpose computer device mentioned above. In the latter case, the application program may be recorded on a removable recording medium such as a magnetic disk, magneto-optical disk, optical disk, or semiconductor memory, or transferred via a network.

[0039] Next, we will explain the operation of the product master creation system configured as described above. Note that the various processes described below are examples, and various other processes that can achieve similar results can be used as appropriate.

[0040] Figures 3 and 4 are sequence diagrams illustrating the operation overview of the product master creation system. As shown in Figures 3 and 4, the processes performed by the product master creation system include process PRO1, which is performed repeatedly for each transaction at store S; process PRO2, which is performed once a day, such as in a night batch; and process PRO3, which is performed in response to the input of product codes for newly handled products from the administrator terminal SM.

[0041] First, let's explain process PRO1. Transactions and settlements are carried out at POS terminal P (step S101). At this time, POS terminal P queries store server 1 for product data about the product using the product code obtained from the product to be registered as a purchased item. Store server 1 reads the product data for the product from the product master database 131 using the product code and provides it to the querying POS terminal P (step S102). When the accounting process is completed, that is, when a transaction is completed, POS terminal P sends the sales data to store server 1 (step S103). Store server 1 receives the sales data with a sales data acquisition unit 111 and stores it in the sales data database 133 (step S104).

[0042] Next, let's explain processing PRO2. In the store server 1, when a predetermined time such as at night arrives, the sales data transmission unit 112 reads sales data for a specified period, for example, for the current day, stored in the sales data database 133, and sends the read sales data to the data collection server 2 (step S111).

[0043] The data collection server 2 receives sales data transmitted from its store server 1, extracts POS data, and stores it in the POS database 21 (step S201). The data collection server 2 performs this process in step S201 for each of the contracted stores S. In this way, the data collection server 2 collects and stores POS data based on sales data from each store S, for example, once a day. Once the POS data for all stores S has been stored, the data collection server 2 distributes the POS data stored in the POS database 21 to the store server 1 of each store S (step S202). The store server 1 of each store S receives this POS data from the data collection server 2 and updates the POS database 134 with that POS data (step S112).

[0044] Next, we will explain process PRO3. This process PRO3 can be performed at any time on the store server 1. On the store server 1, the product master generation unit 114 receives the product code of the newly handled product entered from the administrator terminal SM (step S131). The administrator terminal SM also receives other data, such as price and category name, excluding the product name. If there are multiple types of newly handled products, data for each product is entered on the administrator terminal SM, and the product master generation unit 114 receives multiple product codes. Once all product codes have been entered and a registration start command is issued on the administrator terminal SM, the product master generation unit 114 generates a unified product name for each of the received product codes based on the POS data stored in the POS database 134, generates information about the newly handled product including the generated unified product name, and registers it in the product master database 131 (step S132).

[0045] The operation of the store server 1 will be explained in more detail below. When the store server 1 is in a normal operating state, the processor 11 performs information processing based on application programs stored in the main memory 12 or auxiliary storage device 13. Figure 5 is a flowchart of the information processing performed by the processor 11. The process shown in this flowchart represents the process of registering information about newly handled products to the product master database 131. The processor 11 can perform various operations by executing processes in parallel according to other flowcharts, such as the flowchart for registering sales data. Unless otherwise specified, the process of the processor 11 shown in Figure 5 is assumed to transition from ACTn (where n is a natural number) to ACT(n+1). Here, for each newly registered product entered from the administrator terminal SM, data excluding the product name, such as product code, price, and classification name, is received and stored, for example, in the work area of ​​the main memory 12.

[0046] As ACT1, the preprocessing unit 1141 implemented in the processor 11 performs preprocessing. Specifically, the preprocessing unit 1141 extracts, for example, the product name, product code, and company code or store code from the POS data stored in the POS database 134. The preprocessing unit 1141 stores the extracted information, for example, in the work area of ​​the main memory 12.

[0047] As ACT2, the class classification unit 1142 implemented in the processor 11 classifies the extracted information. The class classification unit 1142 stores the classification results, for example, in the work area of ​​the main memory 12.

[0048] As ACT3, the name generation unit 1143 implemented in the processor 11 initializes the value of an internal counter n (not shown) provided in the processor 11 to "1".

[0049] As ACT4, the name generation unit 1143 generates a unified product name, which is a representative name for the product names included in the extracted information, for cluster n, which is indicated by the value of the internal counter n, among the multiple classified clusters. The name generation unit 1143 stores the generated unified product name in, for example, the work area of ​​the main memory 12.

[0050] In ACT5, the name generation unit 1143 checks whether the value of the internal counter n is less than the number of clusters in the classified clusters. If the value of the internal counter n is less than the number of clusters, the name generation unit 1143 determines it is YES and proceeds to ACT6. If the value of the internal counter n is not less than the number of clusters, the name generation unit 1143 determines it is NO and proceeds to ACT7.

[0051] As ACT6, the name generation unit 1143 increments the value of the internal counter n by "+1". After that, the name generation unit 1143 proceeds to ACT4.

[0052] As ACT7, the post-processing unit 1144 implemented in the processor 11 performs post-processing. Specifically, the post-processing unit 1144 deletes duplicate information and data that does not need to be registered in the product master database 131 from the information, including the generated unified product name, stored in the work area of ​​the main memory 12. The post-processing unit 1144 also adds data, such as the price and classification name of newly handled products, which are separately input from the administrator terminal SM and stored in the work area of ​​the main memory 12, to this deleted data, and generates information to be registered in the product master database 131. The post-processing unit 1144 stores the generated information in the work area of ​​the main memory 12, for example.

[0053] As ACT8, the master registration unit 1145 implemented in the processor 11 registers information to be registered in the product master database 131, for example, information stored in the work area of ​​the main memory 12, into the product master database 131.

[0054] Then, the processor 11 completes the process of registering information about the newly handled product into the product master database 131, as shown in this flowchart.

[0055] Next, we will explain the details of the process for registering information about this new product into the product master database 131.

[0056] The classification unit 1142 implemented in the processor 11 of ACT2 can perform classification processing using, for example, a first classification method that performs clustering by product code, or a second classification method that performs clustering by unsupervised learning.

[0057] Figure 6 is a schematic diagram showing an example of classification using the first classification method. The upper part of Figure 6 shows the product name, product code, and company code extracted from POS data, for example, stored in the work area of ​​the main memory 12. In the diagram, "Company" is the "company code," "Jancode" is the "product code," and "Itemname" is the "product name." As shown in the upper part of Figure 6, even for products with the same product code, the product name registered in the product master database 131 of the company's store S differs from company to company. For this information, the classification unit 1142 can assign the same cluster ID to products with the same product code by performing clustering by product code. As shown in the lower part of Figure 6, the classification unit 1142 adds an item called "ClusterID" to the POS data stored in the work area to store the cluster ID, and writes the cluster ID assigned to the product in each record there.

[0058] As a second classification method that performs clustering using unsupervised learning, non-hierarchical clustering methods such as KMeans and GMM may be used, or hierarchical clustering methods such as Ward's method and Diana's method may be used. When using non-hierarchical clustering methods, the number of clusters can be determined by silhouette analysis.

[0059] Figure 7 is a flowchart showing the detailed procedure for the classification performed in ACT2 in Figure 5 using the second classification method when silhouette analysis is used. Figure 8 is a schematic diagram showing an example of classification using the second classification method.

[0060] As ACT21, the class classification unit 1142 implemented in the processor 11 selects one product to be processed from POS data stored, for example, in the work area of ​​the main memory 12, and obtains the product name of that product. That is, as shown in the upper diagram in Figure 8, the class classification unit 1142 treats one product as one record, and from POS data consisting of multiple records, it selects one record to be processed from top to bottom and reads the product name.

[0061] As ACT22, the class classification unit 1142 vectorizes the acquired product names. The class classification unit 1142 adds the obtained vectors to the POS data of the products to be processed, stored in the work area of ​​the main memory 12, for example. That is, as shown in the upper diagram of Figure 8, the class classification unit 1142 adds an item called "Vector" to store the vectors and writes the vectors of the products to be processed to it.

[0062] As ACT23, the class classification unit 1142 determines whether all products in the POS data stored in the work area of ​​the main memory 12 have been vectorized, that is, whether there are any records in the POS data stored in the work area for which vectors have not been stored. If all products have not been vectorized, the class classification unit 1142 determines NO and proceeds to ACT21. If all products have been vectorized, the class classification unit 1142 determines YES and proceeds to ACT41.

[0063] As ACT24, the class classification unit 1142 initializes the value of an internal counter i (not shown) provided in the processor 11 to "1".

[0064] As ACT25, the class classification unit 1142 clusters the vector of POS data stored in the work area using the number of clusters i, which is indicated by the value of the internal counter i.

[0065] As ACT26, the classification unit 1142 calculates the silhouette coefficient in the clustering result. The silhouette coefficient is a coefficient that can be calculated for each data point, and the classification unit 1142 stores, for example, the average value of the calculated silhouette coefficients in association with the number of clusters i, in a work area of ​​the main memory 12.

[0066] As ACT27, the class classification unit 1142 checks whether the value of the internal counter i is less than the number of POS data records stored in the work area. If the value of the internal counter i is less than the number of records, the class classification unit 1142 determines it is YES and proceeds to ACT28. If the value of the internal counter i is not less than the number of records, the class classification unit 1142 determines it is NO and proceeds to ACT29.

[0067] As ACT28, the classification unit 1142 increments the value of the internal counter i by "+1". After that, the classification unit 1142 proceeds to ACT25.

[0068] As ACT29, the classification unit 1142 determines the number of clusters to be clustered by setting i to the value that maximizes the silhouette coefficient stored in the work area of ​​the main memory 12. Figure 9 is a diagram illustrating the method for determining the number of clusters using the silhouette coefficient. Figure 9 is a graph with the number of clusters i on the horizontal axis and the silhouette coefficient on the vertical axis. In the example shown in Figure 9, the classification unit 1142 determines i=2, where the silhouette coefficient is maximized, as the number of clusters to be clustered.

[0069] As ACT210, the classification unit 1142 clusters the vectors of the POS data stored in the work area using the determined number of clusters i. Then, as shown in the lower diagram in Figure 8, the classification unit 1142 adds an item called "ClusterID" to the POS data stored in the work area to store the cluster ID, and writes the cluster ID, which is the result of clustering the vectors of each record, to this item.

[0070] Furthermore, the name generation process executed by the name generation unit 1143 implemented in the processor 11 of ACT4 includes, for example, a first name generation method that uses a common word among multiple product names within the cluster as the unified product name, and a second name generation method that uses the longest word among multiple product names within the cluster as the unified product name. In addition, a large-scale language model (LLM) may be used as the name generation process. Examples of methods using an LLM include, for example, a third name generation method that summarizes multiple product names within the cluster using an LLM to create a unified product name, a fourth name generation method that vectorizes multiple product names within the cluster using an LLM and uses the product name closest to the centroid of the vector as the unified product name, and a fifth name generation method that learns multiple product names within the cluster using an LLM to generate a unified product name.

[0071] Figure 10 illustrates a first name generation method that uses a common word among multiple product names within the same cluster as the unified product name. In this first name generation method, multiple product names within the cluster are broken down into individual words, and the word with the most matches, which is the common word, is used as the unified product name. As shown in Figure 10, if the product names "Cake," "Chocolate cake," and "Choco chunk cake" are considered products within the same cluster, then "Cake," which is the common word among these product names, is designated as the unified item name. Similarly, if the product names "Iced coffee," "Chilled coffee," and "Coffee" are considered products within the same cluster, then "Coffee," which is the common word among these product names, is designated as the unified item name.

[0072] FIG. 11 is a diagram showing an example of name generation by such a first name generation method. For example, for product names such as "AAAAA AABB11 Plane 400G (where "AAAAA" and "Plane" are half-width characters)", "AAAAAA AA400", "AAAAA Yogurt", "AAAAA AA BB11 Plane", "AAAAA BB11 Pr", "AAAAA AA BB11 P", "AAAAA AA", and "AAAAA AA BB11", the word "AAAAA AA" with the most matches is used as the unified product name.

[0073] FIG. 12 is a diagram for explaining a second name generation method that uses the longest word among multiple product names within the same cluster as the unified product name. In this second name generation method, the product name with the most characters, that is, the longest word, among multiple product names within the same cluster is used as the unified product name. Regarding the number of characters, both half-width and full-width characters are counted as one character. As shown in FIG. 12, for product names such as "Cake", "Chocolate cake", and "Choco chunk cake" that are products within the same cluster, the name with the most characters among them, "Choco chunk cake", is used as the unified product name. Also, for product names such as "Iced coffee", "Chilled coffee", and "Coffee" that are products within the same cluster, the name with the most characters among them, "Chilled coffee", is used as the unified product name.

[0074] FIG. 13 is a diagram showing an example of name generation by such a second name generation method. For example, for product names such as "AAAAAA AABB11 Plane 400G (where "AAAAAA" and "Plane" are half-width characters)", "AAAAAA AA400", "AAAAAA Yogurt", "AAAAAA AA BB11 Plane", "AAAAAA BB11 Pr", "AAAAAA AA BB11 P", "AAAAAA AA", and "AAAAAA AA BB11", the name with the most characters among them, "AAAAAA AABB11 Plane 400G (where "AAAAAA" and "Plane" are half-width characters)", is used as the unified product name.

[0075] FIG. 14 is a diagram for explaining a third name generation method in which a plurality of product names within the same cluster are summarized by an LLM to obtain a unified product name. In this third name generation method, a summary of a plurality of product names within the same cluster is used as the unified product name. As shown in FIG. 14, for product names such as "Cake", "Chocolate cake", and "Choco chunk cake" that are products within the same cluster, by summarizing these product names with an LLM, "Flavoured cake" is generated as the unified product name. Also, for product names such as "Iced coffee", "Chilled coffee", and "Coffee" that are products within the same cluster, by summarizing these product names with an LLM, "Cold coffee" is generated as the unified product name. Thus, by using an LLM, it is possible to generate a unified product name using new words not included in the product names.

[0076] FIG. 15 is a diagram showing an example of name generation by such a third name generation method. For example, for product names such as "AAAAA AABB11 Plane 400G (where "AAAAA" and "Plane" are half-width characters)", "AAAAAA AA400", "AAAAA Yogurt", "AAAAA AA BB11 Plane", "AAAAA BB11 Pr", "AAAAA AA BB11 Pr", "AAAAA AA", and "AAAAA AA BB11", the unified product name "AAAAA AA BB11 Plane 400" is generated by summarizing those product names.

[0077] Figure 16 illustrates a fourth name generation method, in which multiple product names within the same cluster are vectorized using LLM, and the product name closest to the centroid of the vector is adopted as the unified product name. In this fourth name generation method, multiple product names within the same cluster are vectorized, and the product name that is the central word with the smallest sum of Euclidean distances to other product names is adopted as the unified product name. As shown in Figure 16, when the product names "Cake," "Chocolate cake," and "Choco chunk cake," which are products within the same cluster, are vectorized using LLM, the vectors "4.3373e-01,3.1948e-01,...", "4.1633e-01,3.75e-01,...", and "4.2173e-01,3.0048e-01,..." are obtained, respectively. Furthermore, when the product names "Iced coffee," "Chilled coffee," and "Coffee," which are considered products within the same cluster, are vectorized using LLM, the vectors obtained are "44.8373e-01,3.2548e-01,…", "4.4598e-01,3.0053e-01,…", and "4.678e-01,2.953e-01,…", respectively. Here, the sum of the Euclidean distances between each product name is calculated. Figure 17 is a schematic diagram showing an example of name determination in the fourth name generation method. When the Euclidean distance between each product name and the other product names is calculated, in the example in Figure 16, there are three products in the same cluster, so a 3x3 matrix is ​​obtained as shown in Figure 17. For example, for the vector "4.3373e-01,3.1948e-01,…" for the product name "Cake", the Euclidean distances between it and the vectors "4.3373e-01,3.1948e-01,…" for the product name "Cake", "4.1633e-01,3.75e-01,…" for the product name "Chocolate cake", and "4.2173e-01,3.0048e-01,…" for the product name "Choco chunk cake" are "0.0", "0.2", and "0.4", respectively. The Euclidean distance between items with the same name is zero. If we calculate the sum of these "0.0", "0.2", and "0.4", the "Sum" is "0.6".Similarly, the vector "4.1633e-01,3.75e-01,..." for the product name "Chocolate cake" yields a sum of "1.0", and the vector "4.2173e-01,3.0048e-01,..." for the product name "Choco chunk cake" yields a sum of "1.2". Therefore, the central word with the smallest sum, "Cake", is adopted as the unified product name.

[0078] The fifth name generation method, which uses LLM to learn multiple product names within the same cluster and generate a unified product name, includes three types of generation methods: a name generation method using Zero_Shot learning, a name generation method using One_Shot learning, and a name generation method using Few_Shot learning.

[0079] The Zero_Shot learning method for generating a unified product name involves providing the LLM with only instructions and a question, without any examples to learn from, and having the LLM generate a unified product name. Figure 18 shows an example of an instruction (Prompt_head) PH in the fifth name generation method, which generates a unified product name by having the LLM learn multiple product names within the same cluster. Figure 19 shows an example of a question Q in the fifth name generation method. In the Zero_Shot learning method for generating a unified product name, the LLM is given an instruction PH instructing it to generate a representative name in natural language, as shown in Figure 18, and a question Q containing multiple product names within the same cluster for which a unified product name is to be generated, as shown in Figure 19, thereby causing the LLM to generate a unified product name.

[0080] The name generation method using one-shot learning involves providing the LLM with an instruction statement PH and a question Q, along with one example to be trained, and having the LLM generate a unified product name. Figure 20 shows the first example EX1, which is used for one-shot learning in the fifth name generation method. The example used for training is a question that includes multiple product names and a unified product name for a product that is different from multiple product names within the same cluster for which a unified product name is to be generated.

[0081] The name generation method using Few_Shot learning involves providing the LLM with instruction PH and question Q, along with multiple examples to learn from, and having the LLM generate a unified product name. Figure 21 shows the second example EX2 and the third example EX3, which are examples used for Few-Shot learning in the fifth name generation method. In name generation using Few_Shot learning, at least two of the first to third examples EX1 to EX3 are given to the LLM. That is, the examples given to the LLM may be two (first example EX1 and second example EX2), two (first example EX1 and third example EX3), two (second example EX2 and third example EX3), or three (first example EX1 to third example EX3). Of course, there may also be other examples from the fourth example onwards.

[0082] Furthermore, in the post-processing performed by the post-processing unit 1144 implemented in the processor 11 of the AC74, as mentioned above, duplicate information and data that does not need to be registered in the product master database 131 are deleted from the information including the generated unified product name. Figure 22 shows an example of this post-processing. Through the name generation process, a unified product name is generated for multiple product names within multiple clusters, as shown in the upper part of Figure 22. In this example, for example, there are four records for a product with cluster ID "1" and product code "490270501xxxx". In post-processing, three of these four records are deleted, consolidating them into one record, as shown in the lower part of Figure 22. In addition, in post-processing, the company code and cluster ID, which are information that does not need to be stored in the product master database 131, are deleted, and product names that are no longer needed because a unified product name is used are deleted.

[0083] As described above, according to this embodiment, the store server 1 as an information processing device includes a class classification unit 1142, which is an identification unit that classifies the information contained in multiple POS data collected from multiple stores S, and a name generation unit 1143 that generates a unified product name, which is a representative name for the product names contained in the information contained in the multiple POS data, for each cluster classified by the class classification unit 1142. The name generation unit 1143 decomposes the product names contained in the information contained in each of the multiple POS data of the cluster into individual words, and uses the word with the most matches as the unified product name. In this way, when the store server 1 intends to register a new product that it will start handling at the store S in the product master database 131, it generates a unified product name, which will be the name of the product to be registered, based on the product name of the product used in multiple stores S other than its own store. Therefore, according to this embodiment, it is possible to provide an information processing device that can generate information to be registered in the product master database 131, including product names, in accordance with an appropriate policy of using product names similar to those of stores S other than its own store S.

[0084] Furthermore, as described above, according to this embodiment, the store server 1 as an information processing device includes a class classification unit 1142, which is an identification unit that classifies the information contained in multiple POS data collected from multiple stores S, and a name generation unit 1143 that generates a unified product name, which is a representative name for the product names contained in the information contained in the multiple POS data, for each cluster classified by the class classification unit 1142. The name generation unit 1143 uses the product name with the most characters from among the product names contained in the information contained in each of the multiple POS data of the cluster as the unified product name. In this way, when the store server 1 intends to register a new product that it will start handling at the store S in the product master database 131, it generates a unified product name, which will be the product name to be registered, based on the product name of the product used in multiple stores S other than its own store. Therefore, according to this embodiment, it is possible to provide an information processing device that can generate information to be registered in the product master database 131, including product names, in accordance with an appropriate policy of using product names similar to those of stores S other than its own store S.

[0085] Furthermore, as described above, according to this embodiment, the store server 1 as an information processing device includes a class classification unit 1142, which is an identification unit that classifies the information contained in multiple POS data collected from multiple stores S, and a name generation unit 1143 that generates a unified product name, which is a representative name for the product names contained in the information contained in the multiple POS data, for each cluster classified by the class classification unit 1142. The name generation unit 1143 generates the unified product name by processing the product names contained in each piece of information contained in the multiple POS data of the cluster using LLM. In this way, when the store server 1 intends to register a new product that it will start handling at the store S in the product master database 131, it generates a unified product name, which will be the name of the product to be registered, based on the product name of the product used in multiple stores S other than its own store. Therefore, according to this embodiment, it is possible to provide an information processing device that can generate information to be registered in the product master database 131, including product names, in accordance with an appropriate policy of using product names similar to those of stores S other than its own store S.

[0086] Furthermore, according to this embodiment, the store server 1 as an information processing device further includes a preprocessing unit 1141 that extracts information including product names from POS data and supplies the extracted information to the class classification unit 1142. Therefore, according to this embodiment, by extracting only the necessary information from multiple POS data collected from multiple stores S, the amount of information used in the following processing can be reduced, and processing speed can be increased.

[0087] Furthermore, according to this embodiment, the store server 1, as an information processing device, further includes a post-processing unit 1144 that converts the information containing the unified product name generated by the name generation unit 1143 into information to be registered in the product master database 131. Therefore, according to this embodiment, it is possible to delete unnecessary information and generate information to be registered in the product master database 131.

[0088] The name generation unit 1143 uses the LLM to summarize the product names and uses that as the unified product name. Alternatively, the name generation unit 1143 vectorizes the product names using the LLM and uses the product name closest to the centroid of the vector as the unified product name. Or, the name generation unit 1143 generates the unified product name by providing the LLM with an instruction statement and a product name, with or without examples to learn, to instruct the generation of the unified product name. Thus, according to this embodiment, the LLM can be used in various ways to generate the unified product name.

[0089] The above-described embodiment can be modified in various ways as follows.

[0090] For example, if multiple stores S are managed by a company's headquarters, the information processing device may be a headquarters server located at the headquarters, rather than a store server 1. In this case, the product master database is configured on the headquarters server and, at the appropriate time, is sent to the store servers 1 of each store S managed by the headquarters, updating the contents of the product master database 131 on each store server 1.

[0091] Furthermore, the data collection server 2 may function as an electronic receipt server. Alternatively, the data collection server 2 may acquire sales data or POS data from the electronic receipt server.

[0092] While several embodiments of the present invention have been described, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments can be carried out in a variety of other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of symbols]

[0093] 1...Store server, 2...Data collection server, 11...Processor, 12...Main memory, 13...Auxiliary storage device, 14...Network interface, 15...System transmission path, 21,134...POS database, 111...Sales data acquisition unit, 112...Sales data transmission unit, 113...POS data reception unit, 114...Product master generation unit, 1141...Pre-processing unit, 1142...Class classification unit, 1143...Name generation unit, 1144...Post-processing unit, 1145...Master registration unit, 131...Product master database, 132...Store member master database, 133...Sales data database, C...POS data collection center, EX1...First example, EX2...Second example, EX3...Third example, INT...External communication interface, NW...Network, P...POS terminal, PH...Instruction statement, Q...Question, S...Store, SM...Administrator terminal.

Claims

1. An identification unit that classifies information contained in multiple POS data collected from multiple stores, For each cluster classified by the aforementioned identification unit, a name generation unit generates a unified product name which is a representative name for the product name included in the information, Equipped with, The name generation unit is an information processing device that breaks down the product names contained in each of the pieces of information in the cluster into individual words and uses the word with the most matches as the unified product name.

2. An identification unit that classifies information contained in multiple POS data collected from multiple stores, For each cluster classified by the aforementioned identification unit, a name generation unit generates a unified product name which is a representative name for the product name included in the information, Equipped with, The name generation unit is an information processing device that selects the product name with the most characters from among the product names included in each of the pieces of information in the cluster as the unified product name.

3. An identification unit that classifies information contained in multiple POS data collected from multiple stores, For each cluster classified by the aforementioned identification unit, a name generation unit generates a unified product name which is a representative name for the product name included in the information, Equipped with, The name generation unit is an information processing device that generates a unified product name by processing the product names contained in each of the pieces of information in the cluster using a large-scale language model.

4. The information processing apparatus according to any one of claims 1 to 3, further comprising a preprocessing unit that extracts information including the product name from the POS data and supplies the extracted information to the identification unit.

5. The information processing apparatus according to any one of claims 1 to 3, further comprising a post-processing unit that converts information including the unified product name generated by the name generation unit into information to be registered in the product master.

Citation Information

Patent Citations

  • Information processing device and program

    JP2022000799A