Product name matching system, method for generating product name matching master, and program

The product name matching system addresses inefficiencies in product data transmission by unifying and integrating data across organizations, enhancing data management and supporting marketing analysis.

JP7784688B2Active Publication Date: 2025-12-12LAZULI CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2021117730
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-16
Publication Date
2025-12-12
Estimated Expiration
2041-07-16

AI Technical Summary

Technical Problem

Conventional product information transmission methods are time-consuming and labor-intensive due to variations in data formats across organizations, leading to data transcription errors and inefficiencies in data maintenance, especially when handling multiple products across different organizations.

Method used

A product name matching system that includes a product data acquisition unit, an identical product determination unit, an integration generation unit, and an information addition unit to unify and integrate product data, generate feature/evaluation information, and associate additional information with the same product, reducing the effort and time required for data transmission and organization.

Benefits of technology

The system reduces the time and effort needed for transmitting and organizing product information, enabling efficient data management and providing necessary information for marketing analysis such as demand forecasting and product recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007784688000001
    Figure 0007784688000001
  • Figure 0007784688000002
    Figure 0007784688000002
  • Figure 0007784688000003
    Figure 0007784688000003
Patent Text Reader

Abstract

To provide information related to products that makes it possible to save time and effort in product information conveyance and data maintenance, and / or that is necessary for marketing analysis such as demand forecast, product development, or product recommendation, by product information or additional information.SOLUTION: A product name identification system 1 includes: a product data acquisition unit 20 that acquires a plurality of pieces of product data; a same product determination unit 30 that analyzes the plurality of pieces of product data, and identifies a plurality of pieces of product data related to the same product; an integration generation unit 40 that analyzes the plurality of pieces of product data, classifies product information included in the product data according to each category of the product information, integrates, for each product, product information included in the plurality of pieces of product data related to the same product for each category of the product information, and generates a name identification product master; a feature / evaluation information generation unit 50 that generates product feature / evaluation information from data indicating product features included in the product data; and an information addition unit 70 that associates additional information including the feature / evaluation information related to the same product with the same product.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a product name collation system, a method for generating a name collated product master, and a program. [Background technology]

[0002] Products are transported and distributed through product transactions between multiple organizations, such as between a manufacturer and a wholesaler, between a wholesaler and a retailer, or between multiple departments within a single company. Within each organization, product data including product information is managed using spreadsheet software or the like. When a product is traded, not only is the product transported, but product data is also transmitted between organizations. The conventional method of transmitting product data is to transmit product information about the products involved in the transaction by sending it to the trading partner organization via email. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 6427850 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the conventional product information transmission method described above poses the problem of time-consuming and labor-intensive handling of product information. For example, product data for the same product may be written differently or in different formats depending on the organization. Therefore, in order to use product information contained in product data obtained from other organizations within one's own organization, it may be necessary to re-enter the data into a spreadsheet program specifically for that organization and manage it accordingly. This can lead to data transcription errors or omissions of data that should be entered. Furthermore, this data maintenance work must be performed every time the product information contained in the product data is updated. In addition, one organization may handle multiple products with multiple organizations, such as between a manufacturer and a wholesaler, between a wholesaler and a retailer, or between multiple departments within a single company. In such cases, the data maintenance work described above can become enormous. As such, the transmission of product information and data maintenance have been problematic in that they are time-consuming and labor-intensive.

[0005] Furthermore, product data may contain information indicating product features / ratings. However, there is a problem in that simply collecting product data does not allow for the information to be utilized.

[0006] The present invention has been made to solve such problems, and one of its objectives is to provide a product name matching system, a method for generating a name matching product master, and a program that can reduce the effort and time required for transmitting product information and organizing data.

[0007] Another object of the present invention is to provide a product name consolidation system, a method for generating a name consolidation product master, and a program that can provide product-related information necessary for marketing analysis such as demand forecasting, product development, or product recommendations using product information or additional information. [Means for solving the problem]

[0008] A product name merging system according to one aspect of the present invention comprises a product data acquisition unit that acquires multiple product data; an identical product determination unit that analyzes the multiple product data and identifies the multiple product data related to the same product; an integration generation unit that analyzes the multiple product data, classifies product information included in the product data by category of the product information, and for each of the products, integrates product information included in the multiple product data related to the same product by category of the product information to generate an integrated product master; a feature / evaluation information generation unit that generates feature / evaluation information for the product from data indicating the product features included in the product data; and an information addition unit that associates additional information including the feature / evaluation information related to the same product with the same product.

[0009] A method for generating a name-matched product master as one aspect of the present invention is characterized by comprising: a product data acquisition step of acquiring multiple product data; an identical product determination step of analyzing the multiple product data and identifying the multiple product data related to the same product; an integration generation step of analyzing the multiple product data, classifying the product information contained in the product data by category of the product information, and for each of the products, integrating the product information contained in the multiple product data related to the same product by category of the product information to generate a name-matched product master; a feature / evaluation information generation step of generating feature / evaluation information of the product from data indicating the product features contained in the product data; and an information addition step of associating additional information including the feature / evaluation information related to the same product with the same product.

[0010] A program according to one aspect of the present invention is a program for causing a computer to execute the above method. [Effects of the Invention]

[0011] According to the present invention, it is possible to reduce the time and effort required for transmitting product information and organizing data, and / or to provide product-related information necessary for marketing analysis such as demand forecasting, product development, or product recommendations using product information or additional information. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a diagram illustrating a hardware configuration of a product name matching system according to an embodiment. [Figure 2] 1 is a diagram illustrating a configuration of a product name matching system according to an embodiment. [Figure 3] This is an example of product data, specifically, an example of web data on an EC site. [Figure 4] FIG. 10 is a diagram illustrating an example of a detailed configuration of an identical product determination unit. [Figure 5] FIG. 10 is a diagram for explaining a process related to identical product determination. [Figure 6] 10 is a diagram showing an example in which products 0 to 3 are determined to be the same product by the same product determination unit. FIG. [Figure 7] FIG. 10 is a diagram showing an example in which the identical product determination unit determines that products 0, 4 to 6 are different products. [Figure 8] FIG. 10 is a diagram showing an example in which the identical product determination unit determines that products 10 to 13, which are different from product 0, are the same product. [Figure 9] 10 is a diagram showing an example in which the identical product determination unit determines that the products 10, 14 to 16 are different products. [Figure 10] 10 is a diagram showing an example of identical product determination of products 20 to 23 by an identical product determination unit. FIG. [Figure 11] 10 is a diagram showing an example of identical product determination of products 30 to 33 by an identical product determination unit. FIG. [Figure 12] FIG. 2 illustrates an example of a detailed configuration of an integration generation unit. [Figure 13] FIG. 10 is a diagram illustrating an example of a merged product master. [Figure 14] FIG. 2 is a diagram illustrating an example of a detailed configuration of a feature / evaluation information generating unit. [Figure 15] FIG. 10 is a diagram illustrating an example of a detailed configuration of a product graph generation unit. [Figure 16] FIG. 10 is a diagram illustrating an example of a product graph. [Figure 17] 10 is an example of an operation flowchart of the product name matching system according to the embodiment. [Figure 18] 10 is an example of an operation flowchart of the name-collected product master data generation device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, an embodiment of a product name merging system and a product master data generation device according to the present invention will be described in detail with reference to the drawings. For the sake of convenience, more detailed explanations than necessary may be omitted in this specification. For example, detailed explanations of already well-known matters and redundant explanations of substantially identical configurations may be omitted.

[0014] [1. Configuration] [1-1. Overview] The product name consolidation system of this embodiment acquires multiple product data that are data related to products, and if these product data relate to the same product, classifies and integrates the product information contained in the product data by category of the product information to generate a name consolidation product master (see Figure 13 described below).

[0015] Product data is data related to products, and includes product information such as product names, product specifications, and information indicating product features / ratings. Product data may also include product logistics information (e.g., inventory location, allowable number of days for shipping and receiving, etc.), transaction information (e.g., purchase price, sales price, etc.), customer information, and purchasing information. Products covered by product data may also be pharmaceuticals.

[0016] Even when product data pertains to the same product, there may be variations in product information and the product information may not be unified. For example, in the case of product names, there may be product data that includes the manufacturer's official product name, while there may be product data that does not include the official product name but includes the product's abbreviation or nickname. In the case of product specification information, the units of content volume and size of the product may vary. In this way, the product name consolidation system of this embodiment organizes and integrates the unified product data for each product to generate a unified product master, which is a name-consolidated product master.

[0017] The product name matching system also generates feature / rating information for a product from data indicating the features of the product included in the product data. The product name matching system may generate a product graph based on the feature / rating information. The feature / rating information is information indicating the features and / or ratings of a product, and does not include product name or specification information. The feature / rating information will be described in detail later. The product graph is a graph indicating the relationships between products, and will be described in detail later.

[0018] Such a product name matching system, a product master, feature / rating information, and product graphs can be provided by cloud computing.

[0019] [1-2. Hardware configuration] FIG. 1 is a diagram showing the hardware configuration of a product name collation system according to this embodiment. The product name collation system 1 includes a processor 1a, a storage device 1b, an input device 1c, an output device 1d, and a communication device 1e. The components 1a to 1e are connected by a bus 1f. Note that an interface may be interposed between the bus 1f and the components 1a to 1e as needed. The product name collation system 1 may be configured to include computers such as desktop computers, tablet computers, and laptop computers, and does not have to be configured as a single physical device but may be configured as multiple physical devices.

[0020] The processor 1a controls the overall operation of the product name matching system 1. The processor 1a is, for example, an electronic circuit such as a CPU or MPU. The processor 1a performs various processes by reading and executing programs and data stored in the storage device 1b. The processor 1a may be composed of multiple processors.

[0021] The storage device 1b includes RAM 1b-1, which is a volatile memory, and ROM 1b-2, which is a nonvolatile memory. The storage device 1b may also include external memory 1b-3. RAM 1b-1 functions as the main memory and / or work area of ​​the processor 1a. The processor 1a loads programs and the like required for executing processes from ROM 1b-2 or external memory 1b-3 into RAM 1b-1 and executes the loaded programs to perform various operations. The ROM 1b-2 and external memory 1b-3 store the BIOS and OS, which are the control programs of the processor 1a, as well as various programs, data, tables, and the like required to realize the functions executed by the computer. The external memory 1b-3 may include, for example, a flash memory, a hard disk, a DVD-RAM, a USB memory, an SSD, and the like.

[0022] The input device 1c accepts operation instructions and inputs from a user, etc. The input device 1c is a user interface such as an input button, a keyboard, a mouse, a touch panel, a touch pad, a wireless remote control, a microphone, a camera, etc. Note that the touch panel functions as both the input device 1c and the output device 1d.

[0023] The output device 1d outputs data processed by the processor 1a and data that is stored and / or has been stored in the storage device 1b. Examples of the output device 1d include display devices such as CRT displays, liquid crystal displays, organic EL displays, and plasma displays, audio devices such as speakers that emit sound, and printing devices such as printers.

[0024] The communication device 1e is an interface that connects to and communicates with an external device via a network or directly, and may be, for example, a serial interface, a LAN interface, or the like.

[0025] Each part of the product name matching system 1 is realized by various programs stored in the ROM 1b-2 and the external memory 1b-3 using each of the components 1a to 1f as resources.

[0026] [1-3. Detailed configuration] Fig. 2 is a diagram showing the configuration of the product name merging system 1 according to this embodiment. As shown in Fig. 2, the product name merging system 1 includes a product data acquisition unit 20, a identical product determination unit 30, an integration and generation unit 40, a feature / evaluation information generation unit 50, a product graph generation unit 60, and an information addition unit 70. The identical product determination unit 30 and the integration and generation unit 40 may be included to configure the merged product master data generation device 10. The merged product master data generation device 10 may include the product data acquisition unit 20, and the hardware configuration of the merged product master data generation device 10 is similar to that of the product name merging system 1 shown in Fig. 1, so a description thereof will be omitted.

[0027] The product data acquisition unit 20 acquires multiple product data. Examples of product data include web data related to products, such as HTML, and product master data used within an organization. Web data is data related to products on the Internet, such as web pages related to the manufacturer's products on the manufacturer's website, web pages related to the products on an EC (electronic commerce) website, etc.

[0028] In one example, the product data acquisition unit 20 crawls the Internet using a crawler, which is a program stored in the storage device 1b, to acquire web data related to products. This web data can be, for example, HTML data related to products on an EC (electronic commerce) site or HTML data related to products on a manufacturer's site.

[0029] In another example, the product data acquiring unit 20 acquires product master data used by two or more organizations (for example, different companies or different departments within the same company). For example, the product data acquiring unit 20 may acquire the product master data from a device owned by the organization via the communication device 1e in a wired or wireless manner, or may acquire the product master data from an external memory 1b-3 in which the product master data is stored.

[0030] In another example, the product data acquisition unit 20 may acquire web data and a product master data held by an organization as product data. That is, the product data acquisition unit 20 can acquire any data related to products, not limited to open data such as web data that can be accessed by unspecified persons, or closed data such as an internal product master data.

[0031] The product data acquisition unit 20 stores the acquired product data in the storage device 1b, and the storage device 1b can hold the acquired product data as a database of product data.

[0032] Product data is data related to a product, and includes product information such as the product name, product specification information, and information indicating the product's features / ratings. The product data may also include a product identification code, such as a JAN code, assigned to each product. The product name may include the product's name, abbreviation, or nickname. The specification information varies depending on the type of product, but for example, if the product is bottled tea, it may include objective information about the product, such as the product's content volume, container type, size, weight, and number of pieces. The specification information may include numerical information about the product. Numerical information is information including numbers and their units. The information indicating the product's features / ratings may include a product description, product introduction, product sales pitch, catchphrase, product word-of-mouth, and consumer reviews of the product. One example, in the case of web data related to a product, it is the product description included in the data.

[0033] The product data includes product identification potential information. The product identification potential information can include information that can potentially identify the product in addition to information that can identify the product. The product identification potential information can include, for example, the product name.

[0034] FIG. 3 shows an example of product data, specifically, an example of web data on an e-commerce site. In the example shown in FIG. 3, the web data includes a product image G0, product identification latent information G1, specification information G2, and information G3 indicating product features / ratings. In the example shown in FIG. 3, the product identification latent information G1 is "Yokohama Foods Instant Noodle Type Pork Kim Ramen, Spicy Flavor, 3-Meal Pack" listed at the top of the webpage. The product identification latent information G1 includes the product name "Pork Kim Ramen." Specification information G2 includes information such as the product brand "Pork Kim Ramen," the manufacturer name "Yokohama Foods," the product size, weight, and ingredients. The information G3 indicating product features / ratings is product description information, specifically, "We boast a deliciously spicy kimchi and sesame oil soup. Kimchi and eggs are included as added ingredients. Adding chives makes it devilishly delicious...! You'll be hooked!"

[0035] The identical product determination unit 30 analyzes the multiple pieces of product data acquired by the product data acquisition unit 20, and identifies multiple pieces of product data relating to the same product. Figure 4 is a diagram showing an example of the detailed configuration of the identical product determination unit 30.

[0036] As shown in FIG. 4, the identical product determination unit 30 includes a pattern generation unit 31, a vector information generation unit 32, a similarity calculation unit 33, and a determination unit .

[0037] (Pattern generation section) The pattern generation unit 31 generates multiple product identification latent patterns for each of multiple product data items based on the product identification latent information. Specifically, the pattern generation unit 31 extracts the product identification latent information included in the product data items. The pattern generation unit 31 identifies and extracts the product identification latent information using a program 311 according to the type of product data. The pattern generation unit 31 then identifies one or more unit pieces of information included in the product identification latent information and generates multiple product identification latent patterns based on the unit pieces of information. The program 311 may use artificial intelligence (AI). In this specification, AI may be specialized artificial intelligence specialized for a specific application, or general-purpose artificial intelligence capable of dealing with various situations and challenges. Furthermore, the AI ​​may use a known machine-learned model, such as a neural network, a decision tree, a random forest, an SVM (support vector machine), or a k-nearest neighbor algorithm. The trained model may be updated after training.

[0038] (Product identification latent information) Product identification potential information is potential information that can identify a product that includes one or more unit information. Unit information is information that is cohesive. Unit information includes at least the product name. In addition to the product name, the unit information may also include the product manufacturer, content volume, number of pieces, and surplus information. Surplus information is information that is not necessary to identify the product, such as "free shipping," "bulk purchase," "refill," or "credit card payment only." The product name may include the product name, abbreviation, or nickname.

[0039] In one example, if the product data is HTML data acquired from an e-commerce site, the product identification latent information is the product title. The product title is information including the product name, which is included in the location or range of the title tag or name tag in the HTML data. In the example shown in FIG. 3, the product identification latent information (product title) is product identification latent information G1, which includes the product name "Pork Kimchi Ramen." The product title is typically written at the top of the webpage together with an image of the product, as shown in FIG. 3. In another example, if the product data is a product master, the product identification latent information is the product name included in the product master. The notation of the product name is not particularly limited, and may be the official product name published by the manufacturer, an abbreviation or nickname of the product, or a katakana notation.

[0040] When the product data is web data, the pattern generation unit 31 can identify the product identification latent information by identifying the above tags using the program 311, the trained model, or AI and identifying the locations containing the product identification latent information. When the product data is a product master, the pattern generation unit 31 can identify the product identification latent information using the program 311, the trained model, or AI. These program 311, the trained model, and the AI ​​are stored in the storage device 1b, read out by the pattern generation unit 31, and executed by the processor 1a.

[0041] The pattern generation unit 31 eliminates or converts predetermined unit information based on the dictionary 313 or AI. The dictionary 313 or AI includes ones for elimination and ones for conversion, and is stored in the storage device 1b. For example, the exclusion dictionary 313a defines words and phrases that are unit information to be eliminated. The pattern generation unit 31 references the exclusion dictionary 313a from the storage device 1b and eliminates the words and phrases from the extracted product identification latent information. The predetermined unit information to be eliminated is, for example, redundant information. The conversion dictionary 313b defines unit information to be converted, and the pattern generation unit 31 converts the unit information defined in the conversion dictionary 313b into corresponding unit information. The unit information to be converted can be, for example, the manufacturer name or product name of the product. Conversions can include conversion from half-width to full-width, full-width to half-width, English-Japanese, Japanese-English, and conversion of manufacturer name, product name abbreviation, and nickname to official name. In one example, the pattern generation unit 31 converts the abbreviation or nickname of a product, which is unit information included in the product identification latent information, into the official name of the product based on the conversion dictionary 313b. The exclusion AI is a model that excludes unit information to be excluded from the extracted product identification latent information. The conversion AI is a model that identifies unit information to be converted from the extracted product identification latent information and converts it.

[0042] After converting the unit information as described above, if there is duplicated unit information, the pattern generating section 31 may eliminate one of the duplicated unit information.

[0043] In this way, the pattern generation unit 31 performs pre-processing of elimination or conversion before generating a product identification latent pattern, thereby improving the accuracy of determining whether the products are the same.

[0044] After the preprocessing is performed, the pattern generation unit 31 generates a commodity identification latent pattern based on the converted unit information and other unit information included in the commodity identification latent information. The other unit information here refers to unit information that has not been excluded or converted in the preprocessing.

[0045] (Product identification latent pattern) A commodity identification latent pattern is a character string pattern generated by removing, converting, or extracting unit information from commodity identification latent information. Here, the multiple commodity identification latent patterns include any one of the first to seventh patterns, or a combination of two or more of these patterns.

[0046] The first pattern is a character string pattern generated by excluding or converting predetermined unit information from the product identification latent information and excluding specification information included in the product identification latent information. The predetermined unit information that is excluded is redundant information. The specification information is, for example, content volume, size, weight, container type, and number of pieces. If the specification information is content volume "500 ml," the "500 ml" included in the product identification latent information is excluded in the first pattern. The first pattern is, for example, a character string pattern that includes a manufacturer name and a product name.

[0047] The second pattern is a character string pattern obtained by converting the first pattern into another expression format. Conversion to another notation format refers to conversion into a notation format different from the original notation format, such as kanji, hiragana, katakana, or alphabetic characters. In one example of the present embodiment, when the product identification latent information includes kanji, hiragana, katakana, or a combination thereof, the second pattern is a kana notation character string obtained by converting these character strings into kana notation. The second pattern may also be a conversion between full-width and half-width characters.

[0048] The third pattern is a character string pattern obtained by excluding the manufacturer name from the first pattern. The third pattern is, for example, a character string pattern including a product name or a brand name.

[0049] The fourth pattern is a character string pattern generated by excluding or converting predetermined unit information from the product identification latent information and excluding numerical information. The predetermined unit information that is excluded or converted is the same as that of the first pattern. The fourth pattern is a character string pattern that includes, for example, a manufacturer name, a product name, and a brand name. The numerical information is information that includes a numerical value and its unit.

[0050] The fifth pattern is a character string pattern consisting of only the manufacturer name included in the product identification latent information. The fifth pattern may be, for example, the official name of the manufacturer, and may be converted to the official name by the conversion dictionary 313b or the like when the product identification latent information includes the manufacturer's abbreviation or nickname. Furthermore, when the conversion dictionary 313b defines the product name (the official name, abbreviation, or nickname of the product) in association with the manufacturer name, the fifth pattern may be generated by the pattern generation unit 31 converting the product name included in the product identification latent information.

[0051] The sixth pattern is a character string pattern of only the product name (brand name) included in the product identification latent information. The sixth pattern may be, for example, the official name of the product, and when the product identification latent information includes an abbreviation or nickname of the product name (brand name), the abbreviation or nickname may be converted into the official name using the conversion dictionary 313b or the like.

[0052] The seventh pattern is a character string pattern containing only the product specification information included in the product identification latent information. The seventh pattern can be, for example, a character string pattern containing only the content volume, size, weight, container type, or quantity. Multiple seventh patterns may be generated for each type of specification information, such as the content volume, size, weight, container type, and quantity.

[0053] The commodity identification latent pattern is not limited to the first to seventh patterns, and can be any pattern as long as it is based on the commodity identification latent information.

[0054] FIG. 5 is a diagram illustrating a process for determining whether a product is identical. In the example shown in FIG. 5, the pattern generation unit 31 identifies and extracts product identification latent information G1, "Free Shipping Showa Goemon Genmaicha (Brown Rice Tea) PET Bottles, 2000ml x 6 Bottles," as product identification latent information from the product data. The pattern generation unit 31 identifies unit information from the product identification latent information G1. Here, for example, the pattern generation unit 31 identifies unit information such as "Shipping Fee," "Free," "Showa," "Goemon," "Brown Rice Tea," "PET Bottles," "2000ml," and "6 Bottles." The pattern generation unit 31 then excludes "Shipping Fee" and "Free" defined in the exclusion dictionary 313a from the product identification latent information G1, and identifies specification information (here, "PET Bottles" and "2000ml x 6 Bottles") using a trained model, AI, etc., and deletes it from the product identification latent information G1 to generate "Showa Goemon Genmaicha," which is the first pattern P1.

[0055] The pattern generation unit 31 converts the first pattern P1 into kana notation and half-width format to generate a second pattern P2. The pattern generation unit 31 removes the manufacturer name "Showa" from the first pattern P1 using a dictionary 313 that defines manufacturer names, AI, etc., to generate a third pattern P3, "Goemon Genmaicha."

[0056] As with the first pattern P1, the pattern generation unit 31 eliminates "Free Shipping" and "2000ml x 6 bottles" from the product identification latent information G1 to generate a fourth pattern P4, "Showa Goemon Brown Rice Tea Pet." The pattern generation unit 31 generates a fifth pattern P5, which includes only the manufacturer name "Showa," from the product identification latent information G1 using a dictionary 313 that defines manufacturer names, AI, etc. The pattern generation unit 31 generates a sixth pattern P6, which includes only the product name or brand name "Goemon," using a dictionary 313 that defines product names or brand names, AI, etc.

[0057] The pattern generation unit 31 generates a seventh pattern P7 containing only specification information using the program 311, AI, etc. Here, three seventh patterns P7 are generated: a seventh pattern P7a of "2000" indicating the content volume, a seventh pattern P7b of "pet" indicating the container type, and a seventh pattern P7c of "6" indicating the quantity.

[0058] In this way, the pattern generation unit 31 generates multiple commodity identification latent patterns from multiple perspectives by breaking down the commodity identification latent information into unit information as a single unit and eliminating, converting, or extracting the unit information. Therefore, each commodity identification latent pattern made up of a combination of unit information and / or converted unit information can be made to focus on the unit information that humans pay attention to when identifying a commodity.

[0059] As described above, the pattern generation unit 31 removes or converts predetermined unit information included in the commodity identification latent information, and generates the commodity identification latent pattern based on the converted unit information and other unit information included in the commodity identification latent information.

[0060] If the product identification latent information of the reference product data does not include the manufacturer name or its abbreviation, nickname, or specification information, the pattern generation unit 31 may identify this information from the product data and generate the fifth and seventh patterns based on the identified information. The reference product data is product data that serves as the basis for determining whether a product is identical, and can be, for example, web data related to the product on the manufacturer's website. The reference product data is also referred to as reference product data.

[0061] Referring again to FIG. 4, the vector information generation unit 32 generates vector information for each commodity identifying latent pattern for each commodity data. For example, when first to fifth commodity identifying latent patterns are generated for each of two commodity data, the vector information generation unit 32 generates vector information for the first to seventh patterns for one of the commodity data, and generates vector information for the first to seventh patterns for the other commodity data. In the example shown in FIG. 5, the vector information generation unit 32 generates vector information V1 to V7 for the first pattern P1 to the seventh pattern P7 for each commodity data. The vector information for the seventh patterns P7a to P7c is vector information V7a to V7c.

[0062] To generate vector information, a known tool that converts character strings into vectors (N-dimensional numerical values) can be used. For example, Bert provided by Google and fast text provided by Facebook can be used, but are not limited to these.

[0063] The vector information generating unit 32 may generate vector information for each pattern, such as the fifth pattern, the seventh pattern, etc., generated based on the reference product data.

[0064] The similarity calculation unit 33 calculates the similarity between two products for each product identification latent pattern based on the vector information of the product data of the two products. For example, the similarity calculation unit 33 calculates the similarity based on the vector information of the ith pattern of one product data and the vector information of the ith pattern of the other product data (i is a natural number between 1 and the number of product identification latent patterns). That is, the similarity calculation unit 33 does not add or multiply the vector information of different patterns, but calculates the similarity for each pattern independently. The similarity can be, for example, the inner product of the vector information of the ith pattern of one product data and the vector information of the ith pattern of the other product data, cosine similarity, or Euclidean distance. In this embodiment, the similarity is cosine similarity. In one example, the closer the similarity is to 0, the less similar it is, and the closer it is to 1, the more similar it is.

[0065] The determination unit 34 determines, based on the similarity calculated by the similarity calculation unit 33, whether the products indicated by the two product data are the same or not.

[0066] In one example, the determination unit 34 determines that the products indicated by the two product data are the same product if the similarity in one or more patterns among the multiple product identification latent patterns is equal to or greater than a predetermined threshold, and determines that the products indicated by the two product data are different products if the similarity in all patterns is less than the predetermined threshold.

[0067] In another example, the determination unit 34 determines that the products indicated by the two product data are the same product when the similarity of at least one of the first to fourth patterns among the multiple product identification latent patterns is equal to or greater than a predetermined threshold, and determines that the products indicated by the two product data are different products when this is not the case. This makes it possible to determine whether the products are the same or not at the granularity of the product name.

[0068] In yet another example, the determination unit 34 determines that the products indicated by the two product data are the same product if the similarity of all of the first to seventh patterns among the multiple product identification latent patterns is equal to or greater than a predetermined threshold, and determines that the products indicated by the two product data are different products if not. This makes it possible to determine whether the products are the same at the granularity of a product identification code level such as a JAN code.

[0069] The predetermined threshold may be different for each commodity identification latent pattern, and may be appropriately modified to 0.6, 0.7, 0.8, 0.9, etc. Also, although the condition for the commodity being the same is that the similarity is equal to or greater than a predetermined threshold, the condition may also be that the similarity is greater than the predetermined threshold for any of the commodity identification patterns.

[0070] In this way, it is possible to adjust the granularity at which it is determined that products are the same, depending on the type and number of product identification latent patterns.

[0071] When reference product data (hereinafter also referred to as reference product data) exists, the identical product determination unit 30 can perform identical product determination using the respective units 31 to 34 based on the reference product data and product data that has not yet been determined as to whether it is identical to the product in the reference product data (hereinafter also referred to as undetermined product data). The reference product data can be web data of the product on the manufacturer's website. Furthermore, regardless of the presence or absence of reference product data, even if two product data are undetermined product data, the respective units 31 to 34 can perform identical product determination.

[0072] Here, with reference to FIGS. 6 to 11, examples of identical product determination using web data and an organization's product master as product data will be described.

[0073] (Example using web data) FIG. 6 is a diagram showing an example in which products 0 to 3 are determined to be the same product by the same product determination unit 30. Product identification latent information G1 indicates that product 1 is "Showa Orange Drink 900ml 1 case (12 bottles)" (however, "Showa Orange Drink" is written in half-width katakana), product 2 is "Free shipping Heisei Orange Drink 900ml PET bottle x 12 bottles", product 3 is "Showa Heisei Orange Drink 900ml x 12 bottles 1 case KK", and product 0 is "Heisei Orange Drink". The product identification latent information G1 for products 1 to 3 was obtained from an e-commerce site, and the product identification latent information G1 for product 0 was obtained from the manufacturer's site for product 0. In other words, the product data for product 0 is reference product data.

[0074] The exclusion dictionary 313a defines "1 case," "free shipping," "KK," etc. The conversion dictionary 313b defines conversions such as "Showa," "orange drink," etc.

[0075] The first pattern P1 to the seventh pattern P7 for each of the products 0 to 3 are generated by the pattern generation unit 31 as shown in FIG. 6. As shown in FIG. 6, redundant information and numerical information are eliminated from the first to sixth patterns, and half-width katakana characters are converted to full-width characters. However, blank spaces in the product identification latent patterns in FIG. 6 (e.g., the sixth pattern for product 1, the fifth pattern for product 2, and the seventh pattern Pb for product 3) indicate that the product identification latent information G1 does not contain any unit information. Furthermore, the fifth pattern P5 and the seventh pattern P7 for product 0 are generated by the pattern generation unit 31 by identifying the manufacturer name and each specification information based on the reference product data. The seventh pattern P7c for product 0 indicates that there is one case of product 0.

[0076] The similarity for each pattern of each product 0 to 3 is shown at the bottom of FIG. 6. The similarity for each pattern of each product is generated by the vector information generation unit 32. The similarity for each pattern of products 0 to 3 is a value calculated by the similarity calculation unit 33 based on the vector information for each pattern of products 0 to 3 and the vector information for each pattern of product 0. The similarity for each pattern of product 0 is based only on its own vector information, so it is all 1.0 for each pattern. Note that a pattern with a blank similarity or 0 indicates that the unit information corresponding to that pattern is not included in the product identification latent information G1.

[0077] In the example shown in Figure 6, there are patterns in which one or two similarities for products 1 to 3 are blank or 0, but at least the first pattern P1 to the sixth pattern P6 have five or more similarities of 0.6 or higher, and the judgment unit 34 judges that products 1 to 3 are the same product as product 0.

[0078] FIG. 7 is a diagram showing an example in which the identical product determination unit 30 determines that products 0, 4, and 6 are different products. Product identification latent information G1 indicates that product 4 is "[Outlet] Elber Delicious Vitamin C Orange, 1 box (24 bottles)," product 5 is "Nagoya Settlement Orange & Cam Cam, 500 mL PET bottle [24 bottles]," and product 6 is "Free Shipping Heisei Orange Drink, 200 mL paper pack, 24 bottles [Showa]." The product identification latent information G1 for products 4 to 6 was obtained from an e-commerce site. Product 0 is the same as product 0 in FIG. 6.

[0079] The exclusion dictionary 313a defines "outlet," "1 box," "free shipping," etc. The conversion dictionary 313b defines conversions such as {mL, ml}, {PET bottle, pet}, and {paper carton, pack}.

[0080] The first pattern P1 to the seventh pattern P7 of each of the products 4 to 6 are generated by the pattern generation unit 31 as shown in FIG. 7. As shown in FIG. 7, redundant information and numerical information are eliminated from the first to sixth patterns, and units are converted to lowercase. However, blanks in the product identification latent patterns in FIG. 7 (for example, the sixth pattern P6, seventh patterns P7a and P7b of the product 4, and the fifth pattern P5 of the product 5) indicate that the product identification latent information G1 does not include each unit of information.

[0081] The similarity of each of the products 0, 4 to 6 is shown at the bottom of FIG. 7. The similarity of each pattern of the products 0, 4 to 6 is a value calculated by the similarity calculation unit 33 based on the vector information of each pattern of the products 0, 4 to 6 and the vector information of each pattern of the product 0. The similarity of each pattern of the product 0 is based only on its own vector information, so it is 1.0 for all patterns. Note that a pattern with a blank similarity or 0 indicates that the unit information corresponding to that pattern is not included in the product identification latent information G1.

[0082] In the example shown in FIG. 7 , many patterns for product 4 and product 5 have blank or 0 similarity scores, and the calculated similarity scores for each pattern are all less than 0.7 with some exceptions. Therefore, the determination unit 34 determines that product 4 and product 5 are different from product 0. Product 6 has similarity scores of 0.7 or higher for the first pattern P1 to the sixth pattern P6, and can be determined to be the same product at the granularity of brand name. However, the similarity score for the seventh pattern P7 for the specification information is low, less than 0.7, and the determination unit 34 determines that product 6 is different from product 0. In fact, this determination is appropriate because product 0 is a PET bottle product, while product 6 is a paper carton product. In other words, by using as the criterion for determining whether products are the same not only the similarity scores for the first to sixth patterns, but also whether all patterns, including the similarity score for the seventh pattern for the specification information, are equal to or greater than a predetermined threshold, it is possible to determine whether products are the same at a granularity as fine as the JAN code level.

[0083] FIG. 8 is a diagram illustrating an example in which the identical product determination unit 30 determines that products 10 to 13, which are different from product 0, are identical products. While the product category of products 0 to 6 is beverages, the product category of products 10 to 13 is food. Product identification latent information G1 for product 11 is "Yokohama Beef Ramen Bowl 12 servings x 1 case [Credit card payment only]," for product 12 is "Yokohama Foods Yokohama Beef Ramen Bowl (12 bowls) Disaster Preparedness," for product 13 is "<<Case>> Yokohama Foods Beef Ramen Bowl (85g) x 12 Cup Noodles," and for product 10 is "Beef Ramen Bowl." The product identification latent information G1 for products 11 to 13 was obtained from an e-commerce site, and the product identification latent information G1 for product 10 was obtained from the manufacturer's site for product 10. In other words, the product data for product 10 is reference product data.

[0084] The exclusion dictionary 313a defines "credit card payment," "only," "1 case," "case," "<<case>>," etc. The conversion dictionary 313b defines {Yokohama, Yokohama Foods}, etc.

[0085] The first pattern P1 to the seventh pattern P7 for each of the products 10 to 13 are generated by the pattern generation unit 31 as shown in FIG. 8. As shown in FIG. 8, redundant information and numerical information are eliminated from the first to sixth patterns, and the manufacturer name is converted to the official name "Yokohama Foods." However, the blanks in the product identification latent patterns in FIG. 8 (e.g., the seventh patterns P7b and P7c for product 11, the seventh patterns P7b and P7c for product 12, and the seventh pattern P7b for product 13) indicate that the product identification latent information G1 does not include each unit of information. Furthermore, the fifth pattern P5 and the seventh pattern P7 for product 10 are generated by the pattern generation unit 31 by identifying the manufacturer name and each specification information based on the reference product data. The seventh pattern P7c for product 10 indicates that product 10 is one case.

[0086] The similarity between each pattern of each of the products 10 to 13 is shown at the bottom of Fig. 8. The similarity between each pattern of each product is generated by the vector information generation unit 32. The similarity between each pattern of the products 10 to 13 is a value calculated by the similarity calculation unit 33 based on the vector information of each pattern of the products 10 to 13 and the vector information of each pattern of the product 10. The similarity between each pattern of the product 10 is based only on its own vector information, so it is all 1.0 for each pattern.

[0087] In the example shown in Figure 8, there are patterns in which one or two similarities are blank or 0 for products 11 to 13, but at least six or more of the first pattern P1 to the sixth pattern P6 have similarities of 0.6 or higher, and the judgment unit 34 judges that products 11 to 13 are the same product as product 10.

[0088] FIG. 9 is a diagram showing an example in which the identical product determination unit 30 determines that products 10, 14, and 16 are different products. Product identification latent information G1 indicates that product 14 is "Oyama Company Child Star Ramen Beef 39g [Confectionery 1 case 24 bags]," product 15 is "Yokohama Beef Ramen Cabbage Salad Light Soy Sauce Flavor 3-pack 120g," and product 16 is "S&X Oyatsu Ramen (Beef Flavor) 12 servings." The product identification latent information G1 for products 14 to 16 was obtained from an e-commerce site. Product 10 is the same as product 10 in FIG. 8.

[0089] The exclusion dictionary 313a defines "confectionery," "one case," etc. The conversion dictionary 313b defines conversions such as {Yokohama, Yokohama Foods}.

[0090] The first pattern P1 to the seventh pattern P7 for each of the products 10, 14 to 16 are generated by the pattern generation unit 31 as shown in Fig. 9. As shown in Fig. 9, redundant information and numerical information are eliminated from the first to sixth patterns, and the manufacturer name is converted to the official name "Yokohama Foods." However, the blanks in the product identification latent patterns in Fig. 9 (for example, the seventh pattern P7c for product 14, the seventh pattern P7c for product 15, and the fifth pattern P5, seventh patterns P7b, and P7c for product 16) indicate that the product identification latent information G1 does not include each unit of information.

[0091] The similarity for each pattern of each product 10, 14 to 16 is shown at the bottom of Figure 9. The similarity for each pattern of each product is generated by the vector information generation unit 32. The similarity for each pattern of products 10, 14 to 16 is a value calculated by the similarity calculation unit 33 based on the vector information for each pattern of the products 10, 14 to 16 and the vector information for each pattern of the product 10. The similarity for each pattern of the product 10 is based only on its own vector information, so it is all 1.0 for each pattern. Note that a pattern with a blank similarity or 0 indicates that the unit information corresponding to that pattern is not included in the product identification latent information G1.

[0092] In the example shown in FIG. 9, the products 14 to 16 are determined by the determining unit 34 to be different from the product 10 because the calculated similarity of each pattern is less than 0.7 with some exceptions.

[0093] (Example using organization product master) FIG. 10 is a diagram showing an example of identical product determination for products 20 to 23 by the identical product determination unit 30. Product identification latent information G1 for product 21 is "Reiwa Probao R-1 Drink Type 112ml," for product 22 is "Reiwa (reiwa) 'R-1' 112g," for product 23 is "Reiwa 'R-1' Low Fat 112g," and for product 20 is "Reiwa Probao Yogurt R-1 112g." The product identification latent information G1 for products 21 to 23 is obtained from the product master of the organization (e.g., wholesaler, retailer, etc.) that handles the products 21 to 23, and the product identification latent information G1 for product 20 is obtained from the manufacturer's website or the manufacturer's product master for product 20. In other words, the product data for product 20 is reference product data.

[0094] The exclusion dictionary 313a defines various redundant information. The conversion dictionary 313b defines conversions such as {R-1, Probao Yogurt R-1}, {reiwa, reiwa}, {reiwa, Reiwa}, {R-1, R-1}, and {R-1, Probao Yogurt R-1}.

[0095] The first patterns P1 to seventh patterns P7 of each of the products 20 to 23 are generated by the pattern generation unit 31 as shown in FIG. 10. As shown in FIG. 10, in the first to sixth patterns, surplus information and numerical information are excluded, and the abbreviations "R-1", "R-1" are converted to the official product name "Probio Yogurt R-1", and the full-width alphabetic notation "reiwa" is converted to the kanji notation "令和". However, the blanks in the product identification potential patterns in FIG. 10 (for example, the seventh pattern P7a of product 21, the seventh patterns P7b, P7c of product 22, the seventh patterns P7b, P7c of product 23) indicate that each unit information is not included in the product identification potential information G1. Also, the fifth pattern P5 and the seventh pattern P7 of product 20 are generated by the pattern generation unit 31 by specifying each specification information based on the reference product data. Note that the seventh pattern P7b of product 20 indicates that the unit form of its content volume is the ml form, and the official unit form of the content volume is the g form, so the corresponding unit information is not included in the product identification potential information G1. The seventh pattern P7c of product 20 indicates that product 20 is in the cup form. That is, product 20 is a yogurt with the product name "Probio Yogurt R-1" in one cup with a content volume of 112 g, and is a type of yogurt to be eaten.

[0096] Below FIG. 10, the similarity in each pattern of each of the products 20 to 23 is shown. The similarity of each pattern of each product is generated by the vector information generation unit 32. The similarity of each pattern of products 20 to 23 is a value calculated by the similarity calculation unit 33 based on the vector information of each pattern of the products 20 to 23 and the vector information of each pattern of product 20. The similarity of each pattern of product 20 is 1.0 in each pattern because it is based only on its own vector information.

[0097] In the example shown in FIG. 10, the similarity of product 22 is 1.0 for all patterns except for seventh patterns P7b and P7c, and the determination unit 34 determines that product 22 is the same product as product 20. That is, product 22 is a yogurt to be eaten, just like product 20. On the other hand, product 21 has several patterns among the similarity patterns with similarity of less than 0.8 (for example, first pattern P1 to sixth pattern P6), and the determination unit 34 determines that product 21 is a different product from product 20. In fact, product 21 is of the same brand as product 20, but is a drink type, which is different from product 20, which is a drink type. In other words, the similarity is reduced because the character string "drink type" is included in the first to fourth patterns P1 to P4, and product 22 is determined to be a different product. Furthermore, product 23 has a similarity of 0.8 or more in many patterns, and is similar to product 20, but since the similarity in third pattern P3 is 0.3, which is less than the threshold value of 0.8, product 23 is determined by the determination unit 34 to be a different product from product 20. In fact, product 23 is of the same brand as product 20, but is a low-fat version, and the inclusion of "low-fat" in third pattern P3 is a factor that reduces the similarity and causes product 23 to be determined to be a different product.

[0098] As described above, it can be seen that even when the product data is the product master of an organization and the product abbreviation is included in the product identification latent information G1, it is possible to correctly determine whether the products are the same.

[0099] FIG. 11 is a diagram showing an example of identical product determination for products 30 to 33 by the identical product determination unit 30. Product identification latent information G1 is written in half-width characters as follows: product 31 is "Good Puren Cleanse Care Shampoo," product 32 is "Good Puren Natural Shampoo 340ml Refill," product 33 is "Hanako Kako Good Puren Natural Shampoo Pump 425ml," and product 30 is "Good Puren Natural Cleanse Care Shampoo 425ml." The product identification latent information G1 for products 31 to 33 is obtained from the product master of the organization (e.g., wholesaler, retailer, etc.) that handles the products 31 to 33, and the product identification latent information G1 for product 30 is obtained from the manufacturer's website or product master of product 30. In other words, the product data for product 30 is reference product data.

[0100] Various redundant information is defined in the exclusion dictionary 313a. Conversions such as {kako, Hanako} and {good, Hanako} are defined in the conversion dictionary 313b.

[0101] The first pattern P1 to the seventh pattern P7 for each of the products 30 to 33 are generated by the pattern generation unit 31 as shown in FIG. 11. As shown in FIG. 11, redundant information and numerical information are eliminated from the first to sixth patterns, and the manufacturer's name in English, "kako," is converted to the kanji notation "Hanako." For products 31 and 32, the manufacturer's name "Hanako" is not included in the product identification latent information G1. However, the fifth pattern P5 for products 31 and 32 was generated by converting the product name "Good" to the manufacturer's name "Hanako" using the conversion dictionary 313b. However, blanks in the product identification latent patterns in FIG. 11 (e.g., the seventh pattern P7 for product 31 and the seventh pattern P7b for product 32) indicate that the product identification latent information G1 does not include any unit information. Furthermore, the fifth pattern P5 and the seventh pattern P7 for product 30 were generated by the pattern generation unit 31 by identifying each piece of specification information based on the reference product data. Product 30 is a bottle (pump) type shampoo with a content of 425 ml.

[0102] The similarity for each pattern of each product 30 to 33 is shown at the bottom of Fig. 11. The similarity for each pattern of each product is generated by the vector information generation unit 32. The similarity for each pattern of products 30 to 33 is a value calculated by the similarity calculation unit 33 based on the vector information for each pattern of the products 30 to 33 and the vector information for each pattern of product 30. The similarity for each pattern of product 30 is based only on its own vector information, so it is all 1.0 for each pattern.

[0103] In the example shown in FIG. 11 , the similarity of product 33 is equal to or greater than the threshold value of 0.7 for all patterns except for the seventh pattern P7b for the container type (packaging type), and the determination unit 34 determines that product 33 is the same product as product 30. That is, product 33 is a bottle-type shampoo, just like product 30. The similarity of the seventh pattern P7b for the container type (packaging type) is 0.6, which is lower than the other patterns. This is because the seventh pattern P7b for the reference product 30 contains two units of information: "bottle, pump." Meanwhile, product 31 has a similarity of 0.8 for the first through sixth patterns, exceeding the threshold value of 0.7, but does not have a similarity for the seventh pattern, and its content volume and container type (packaging type) are unknown. Therefore, while it can be determined that the two products are the same at the brand name level, it is impossible to determine whether they are the same or different at the JAN code level. Product 32 has a similarity of 0.6 for approximately half of the patterns, which is below the threshold value of 0.7, and the determination unit 34 determines that it is a different product from product 30. That is, product 32 is a "refill" shampoo, which is different from bottled shampoo, and therefore it is thought that the similarity was calculated to be low in half of the patterns.

[0104] (Integrated generation part) 12 is a diagram showing an example of the detailed configuration of the integrating / generating unit 40. The integrating / generating unit 40 analyzes multiple product data, classifies product information included in the product data by category of the product information, and for each product, integrates product information included in multiple product data related to the same product by category of the product information to generate a merged product master.

[0105] 12, the integrating and generating unit 40 estimates, classifies, and / or divides the categories of product information included in the product data. For example, the integrating and generating unit 40 estimates, classifies, and / or divides the categories of product information included in the product data using AI 41. Then, the integrating and generating unit 40 generates a merged product master by integrating product information of product data related to the same product for each divided category.

[0106] The integrating and generating unit 40 estimates the category of a product from the product information in the product data. The product category can be any classification of product categories, such as food, beverages, confectionery, alcoholic beverages, bags, cosmetics, etc. The product category may be estimated by classifying the product into multiple divisions, such as major divisions, medium divisions, and minor divisions. In one example, if the product is beer, the integrating and generating unit 40 estimates the product category by classifying food as the major division, beverages as the medium division, and alcoholic beverages or alcoholic beverages as the minor division. The product category can be estimated using, for example, a dictionary and / or AI. The integrating and generating unit 40 generates a product information category for the product category and adds the generated product category to the estimated product category, thereby adding the product category to the aggregated product master as one of the product information categories.

[0107] The integration / generation unit 40 may eliminate overlapping product information for each divided category during integration. For example, the integration / generation unit 40 associates the same product name from two pieces of product data with a product name category, but eliminates one of them.

[0108] During integration, the integration generation unit 40 may normalize the representation of the product information for each divided category. Normalization may involve, for example, standardizing units included in the product information within a category and standardizing variations in the words and phrases that make up the product information. A normalization dictionary 42 and a normalization AI 43 may be used for normalization. The normalization dictionary 42 and the normalization AI 43 are dictionaries or AIs that convert units, such as converting l to ml and m to cm, or dictionaries or AIs that standardize variations in the spelling of words and phrases that make up the product information, such as converting full-width character strings to half-width character strings. The normalization dictionary 42 and the normalization AI 43 may be provided independently for unit standardization and spelling standardization.

[0109] (Name-matching product master) The merged product master is a product master that is composed of the product information of two or more product data items integrated together. The merged product master can be a table in which each product is listed vertically and the categories of product information for the products are listed horizontally. In other words, the merged product master is a table in which product information for one product is stored in classified columns in each row. Each column is a classified or classified category. The merged product master includes columns for product information categories that are present in only one of the product data items related to two or more identical products. In other words, product information from product data items related to two or more identical products is aggregated into the merged product master.

[0110] FIG. 13 is a diagram showing an example of the merged product master data M. In the example shown in FIG. 13, product information for three products is aggregated in each row. The merged product master data M has at least product information categories (columns) C10 to C17 for each product. Columns C10 to C17 respectively represent the product name, product identifier (here, JAN code), product category, capacity, energy, product size, feature / rating information, and vector information. The feature / rating information and vector information are generated by the feature / rating information generation unit 50 and the product graph generation unit 60, as will be described later, and added to the merged product master data M by the information addition unit 70. Note that although the merged product master data M in FIG. 13 shows three products, this is not a limitation. The merged product master data M can contain as many products as there are available.

[0111] (Feature / Evaluation Information Generation Unit) 14 is a diagram showing an example of a detailed configuration of the feature / evaluation information generation unit. The feature / evaluation information generation unit 50 generates feature / evaluation information of a product from data indicating the feature of the product included in the product data. Specifically, the feature / evaluation information generation unit 50 has an identification unit 51, a generation unit 52, and a category estimation unit 53.

[0112] The identification unit 51 identifies data indicating product characteristics contained in the product data. Here, the data indicating product characteristics is text data describing the product, but it may also be image data or audio data as long as it indicates product characteristics. The identification unit 51 identifies words and phrases indicating the product characteristics from the identified text data. This identification can be performed by a dictionary 51a, an AI, or both. Alternatively, the identification may be performed by a program depending on the type of product data.

[0113] In one example, when the product data is HTML data acquired from an e-commerce site, the identification unit 51 identifies and extracts a product description included in a portion or range of the data tagged with a "description" tag. For example, the portion or range is the portion or range marked with the symbol G3 in FIG. 3. The identification unit 51 then performs natural language analysis on the extracted product description using AI to break it down into words and phrases, and identifies words and phrases that indicate product features. Morphological analysis can be used in the process of breaking down the product description into words. Examples of tools that can be used for morphological analysis include, but are not limited to, MeCab, Juman++, and Janome.

[0114] In the example shown in Figure 3, the identification unit 51 identifies words such as "kimchi," "sesame oil," "spicy kimchi," "spicy soup," "egg," "devilish," and "addictive" from information G3 indicating the product's characteristics / evaluation: "We're proud of our delicious spicy kimchi and sesame oil soup. Kimchi and eggs are included as added ingredients. If you add chives, it's devilishly delicious...! You'll be addicted, guaranteed."

[0115] The generation unit 52 generates characteristic / evaluation information in association with the identified words and phrases. Specifically, a column of characteristic / evaluation information is generated as one of the product information categories (columns) in the name-matched product master, and the characteristic / evaluation information is generated by associating the column with the identified words and phrases.

[0116] In the example shown in Figure 3, the generation unit 52 generates a characteristic / evaluation information column and generates characteristic / evaluation information by associating the characteristic / evaluation information column with words such as "kimchi," "sesame oil," "spicy," "spicy soup," "egg," "devilish," and "addictive."

[0117] In the example shown in FIG. 13, the characteristic / evaluation information for the product "ABC Coffee" in the top row of the aggregated product master M is "black," "Mandheling," "refreshing," and "fruit." The characteristic / evaluation information for the product "DEF Sports Drink" in the middle row of the aggregated product master M is "sports," "large volume," and "grapefruit flavor." The characteristic / evaluation information for the product "GHI Snack" in the bottom row of the aggregated product master M is "corn," "crunchy," "BBQ," and "limited edition." These words and phrases are associated with the characteristic / evaluation information column C16 for each product to generate characteristic / evaluation information.

[0118] The generation unit 52 may also associate the identified words and phrases with a product information category estimated or generated by the category estimation unit 53. This product information category is a category classification or category division assigned to the feature / evaluation information column or the identified words and phrases included in the feature / evaluation information column. That is, the generation unit 52 may associate the identified words and phrases with the feature / evaluation information column regardless of the meaning of the identified words and phrases. The generation unit 52 may also associate the identified words and phrases with a feature / evaluation information column according to the meaning of the identified words and phrases. Alternatively, the generation unit 52 may associate the identified words and phrases with the semantic content of the identified words and phrases in the feature / evaluation information column. In the example shown in FIG. 13 , the category estimation unit 53 may generate an ingredient column from “Mandheling” and “Corn” and associate these words and phrases with the column, as described below. Furthermore, the generation unit 52 may generate an impression column from “refreshing” and “crispy” and associate these words and phrases with the column. In this case, the characteristic / evaluation information column C16 includes an ingredient column including "Mandheling" and "corn" and an impression column including "refreshing" and "crispy."

[0119] The category estimation unit 53 estimates or generates a product information category corresponding to the identified word or phrase. This estimation or generation can be performed using a dictionary 53a, a natural language processing library 53b, AI, or two or more of these. In other words, the category estimation unit 53 assigns meaning to the identified word or phrase. Assigning meaning means estimating or generating a product information category (column) corresponding to the word or phrase.

[0120] In one example, the category estimation unit 53 determines the part of speech of a word or phrase and associates the determined part of speech with the word or phrase. Parts of speech include nouns, adjectives, adjectives, adverbs, etc. Adjectives are words that indicate the shape of a product. The part of speech can be determined using a dictionary 53a, a natural language processing library 53b, AI, or two or more of these. The category estimation unit 53 estimates or generates a product information category based on the determined part of speech, or generates a product information category corresponding to the determined part of speech.

[0121] In another example, the category estimation unit 53 determines the meaning of a word or phrase, and estimates or generates a product information category that encompasses the meaning. That is, the category estimation unit 53 estimates or generates one of the product information categories (i.e., columns) from the meaning of the word or phrase. This determination can be made using the dictionary 53a, the natural language processing library 53b, AI, or two or more of them.

[0122] Specifically, when the meaning of a word or phrase corresponds to the subjective meaning of a product, the category estimation unit 53 estimates or generates a product information category that represents the subjective meaning of the product. In one example, the category estimation unit 53 determines whether the word or phrase corresponds to one of the product information categories (i.e., columns) such as impression, atmosphere, taste, texture, quality, and use. The product information category corresponding to the determined word or phrase is not limited to these, and may be estimated or generated arbitrarily according to the product. For example, words and phrases such as "refreshing," "soft," and "fluffy" correspond to the impression column. The product information category to which meaning is assigned is not limited to subjective ones, but may also be objective ones such as the product's raw materials and ingredients. For the estimation or generation, the dictionary 53a, the natural language processing library 53b, AI, or two or more of these may be used.

[0123] In yet another example, the category estimation unit 53 generates combinations of words or phrases having dependency relationships. Dependencies refer to relationships in which different words or phrases are connected in meaning, such as between a subject and a predicate, a modifier and a modified word, or a sentence that indicates what is being referred to. Examples of dependency relationships include, but are not limited to, "vivid color" and "stylish atmosphere," and can be suited to a product based on information indicating the product's features / ratings. The combinations can be generated using the dictionary 53a, the natural language processing library 53b, AI, or two or more of these.

[0124] Subjective characteristics / evaluation information such as impressions, and dependency characteristics / evaluation information are likely to be entered as search keywords, and are therefore highly useful when searching the database of matched product master data, and are highly useful as meta information for e-commerce sites, for example.

[0125] The feature / evaluation information generation unit 50 generates feature / evaluation information that compiles all of the generated feature / evaluation information. In this specification, this feature / evaluation information is referred to as overall feature / evaluation information, and feature / evaluation information related to the above words, phrases, or combinations of words and phrases may be referred to as individual feature / evaluation information. The overall feature / evaluation information is a string of characters formed by concatenating the words and phrases of all of the individual feature / evaluation information, and is associated with an overall feature / evaluation information column, which is one of the feature / evaluation information columns generated by the generation unit 52.

[0126] (Features / Evaluation Information) The feature / rating information is a string of characters indicating the features of a product, and is also called a meta tag. In this embodiment, the feature / rating information is a word, phrase, or combination of words or phrases having dependencies that indicate the product features. The feature / rating information associates the word, phrase, and combination with the corresponding product information category. The feature / rating information is associated with the name-matched product master for the corresponding product by the information addition unit 70. The feature / rating information is one of the additional information added to the name-matched product master.

[0127] (Product graph generation part) 15 is a diagram showing an example of the detailed configuration of the product graph generation unit. The product graph generation unit 60 generates a product graph showing the relationships between products based on feature / rating information. Specifically, the product graph generation unit 60 has a vector information calculation unit 61, a distance calculation unit 62, and a graph generation unit 63.

[0128] The vector information calculation unit 61 calculates vector information based on the feature / evaluation information. For example, the vector information calculation unit 61 converts a character string of the feature / evaluation information into vector information. A known tool that converts a character string into a vector (N-dimensional numerical value) can be used to calculate (convert) this vector information. For example, Bert provided by Google Inc. or FastText provided by Facebook Inc. can be used, but is not limited to these.

[0129] The vector information calculation unit 61 calculates vector information for all feature / evaluation information, i.e., all individual feature / evaluation information and overall feature / evaluation information. In this specification, the vector information for individual feature / evaluation information may be referred to as individual vector information, and the overall feature / evaluation information may be referred to as overall vector information. Each calculated vector information is stored in the storage device 1b in association with the corresponding product and the aggregated product master. This association may be performed by, for example, the information addition unit 70.

[0130] The distance calculation unit 62 calculates the distance between products based on the vector information. This distance can be, for example, the inner product of the vector information or the Euclidean distance.

[0131] The distance between products can be broadly divided into the distance between individual vector information (also referred to as the "distance between feature / rating information") and the distance between overall vector information (also referred to as the "distance between products"). The distance between feature / rating information includes the distance between individual vector information of different products in the same product category and the distance between individual vector information of different products in different product categories. The distance between products includes the distance between overall vector information of different products in the same product category and the distance between overall vector information of different products in different product categories.

[0132] The graph generation unit 63 (product graph generation unit 60) generates a product graph for the same product category and / or a product graph between multiple product categories. This generation is based on the distance calculated by the distance calculation unit 62. For example, the graph generation unit 63 generates a product graph that includes products whose calculated distance is within a predetermined distance. The product graph can be generated using a known method such as a social graph creation method.

[0133] FIG. 16 is a diagram illustrating an example of a product graph. The product graph in FIG. 16 is a product graph for convenience store sweets from companies A, B, and C. Specifically, this product graph was obtained by extracting products from the aggregated product master data whose manufacturer names are "Company A," "Company B," or "Company C" and whose characteristic / evaluation information column is "convenience store sweets" and plotting them on a graph. In the product graph, cohesive regions are circled, and each region is labeled with a word indicated by the characteristic / evaluation information, as shown in FIG. 16. For example, a region for a product with ingredient characteristic / evaluation information is labeled with the ingredient (e.g., "strawberry," "blueberry," etc.), a region for a product with product category characteristic / evaluation information is labeled with the category (e.g., "cake," "pudding," etc.), and a region for a product with impression characteristic / evaluation information is labeled with the impression (e.g., "smooth," "chewy," "moist," etc.). This product graph allows us to understand the competitive relationships between each company.

[0134] (Product graph) Product graphs for the same product category include (1) a product graph showing the relationship between products in the same product category based on the distance between a single individual feature / rating information, (2) a product graph showing the relationship between products in the same product category based on the distance between multiple individual features / rating information, and (3) a product graph showing the relationship between products in the same product category based on the distance between overall features / rating information. The product graph in (1) above is a graph (map) for products that share a product category and a single individual feature / rating information. The product graph in (2) above is a graph (map) for products that share a product category and multiple individual features / rating information. The product graph in (3) above is a graph (map) for products that share a common product category.

[0135] Product graphs across multiple product categories include (4) a product graph showing the relationships between products within multiple product categories based on the distance between a single individual feature / rating information, (5) a product graph showing the relationships between products within multiple product categories based on the distance between multiple individual features / rating information, and (6) a product graph showing the relationships between products within multiple product categories based on the distance between overall features / rating information. The product graph in (4) above is a graph (map) for products that share a single individual feature / rating information. The product graph in (5) above is a graph (map) for products that share multiple individual features / rating information. The product graph in (6) above is a graph (map) for products in multiple product categories.

[0136] The product graphs (1) to (6) above can provide relationships between products from different perspectives depending on the type and number of feature / evaluation information and the number of product categories, making it easier for users to gain insights when analyzing the relationships between product groups.

[0137] (Information Addition Section) The information addition unit 70 associates the feature / evaluation information related to the same product with the same product. Specifically, the information addition unit 70 associates the feature / evaluation information associated with the product with the name-matched product master and stores it in the storage device 1b.

[0138] The information addition unit 70 associates the product graph generated by the product graph generation unit 60 with the corresponding product and the merged product master data and stores it in the storage device 1b. The information addition unit 70 may also store the vector information calculated by the vector information calculation unit 61 in association with the corresponding product and the merged product master data and store it in the storage device 1b. The vector information calculated by the vector information calculation unit 61 is individual vector information and / or overall vector information, and is also referred to as graph vector information.

[0139] The feature / rating information, product graph, and graph vector information are included in the additional information added to the name-matched product master. The product graph and graph vector information are included in the product graph information. When associating the feature / rating information and / or product graph information with the name-matched product master, the information addition unit 70 associates at least one of the feature / rating information, product graph, and graph vector information with the name-matched product master.

[0140] The merged product master may be stored in the storage device 1b in such a manner that the product name, specification information, feature / rating information, and product graph information for each product are associated with each other. Based on this association, the product name merging system 1 may be provided with a relational database including a merged product master database generated by the integration generation unit 40, a feature / rating information database that collects feature / rating information for each product, and a product graph information database that collects product graph and / or graph vector information for each product.

[0141] [2. Operation] [2-1. Overall movement] 17 is an example of an operation flowchart of the product name consolidation system of this embodiment. First, the product name consolidation system 1 acquires two or more product data items using the product data acquisition unit 20 (S01: Acquire product data). Here, the product data is assumed to be two or more product data items (HTML data) acquired from one or more EC sites, but as mentioned above, it is not limited to this and product master data used by an organization may also be acquired.

[0142] Next, the product name matching system 1 analyzes the multiple product data and identifies multiple product data related to the same product using the same product determination unit 30 (S02: Identifying product data related to the same product). Specifically, the same product determination unit 30 determines whether the two or more acquired product data relate to the same product. If the same product determination unit 30 determines that the two or more acquired product data do not relate to the same product, the process returns to S01. If the same product determination unit 30 determines that the two or more acquired product data relate to the same product, the process proceeds to the next step S03.

[0143] If it is determined that the two or more acquired product data relate to the same product, the integrating / generating unit 40 analyzes the two or more product data and classifies the product information included in the product data by category of the product information (S03: Categorization of product information). In one example, the integrating / generating unit 40 estimates, classifies, and divides the product information by category using AI 41. The integrating / generating unit 40 then integrates the product information included in the two or more product data relating to the same product by category of the product information to generate a merged product master (S04: Generation of merged product master). This integrates the product data relating to the same product, eliminating the need to manually input one product data into another.

[0144] Furthermore, in S04, the integrating / generating unit 40 eliminates duplication of product information for each divided category. This is because identical information is unnecessary for the same product information category. Furthermore, in S04, the integrating / generating unit 40 normalizes the expression of product information for each divided category. That is, the integrating / generating unit 40 standardizes the units included in the product information within a category using the normalization dictionary 42 and the normalization AI 43, and standardizes the variations in the words and phrases that make up the product information. The units, words, and phrases to be standardized may be defined in the normalization dictionary 42 or may be determined by the normalization AI 43.

[0145] The feature / evaluation information generation unit 50 generates feature / evaluation information for the product from data indicating the feature of the product included in the product data (S05: Generate feature / evaluation information). In one example, the identification unit 51 identifies and extracts product description text included in a portion or range tagged with a predetermined tag from text data indicating the feature of the product included in each product data. The identification unit 51 then performs natural language analysis on the extracted product description text using AI to break it down into words and phrases and identify the words and phrases indicating the product feature. The generation unit 52 generates feature / evaluation information associated with the words. More specifically, the category estimation unit 53 may estimate a category of product information corresponding to the identified word or phrase, and the generation unit 52 may generate feature / evaluation information associated with the word, phrase, and product information category.

[0146] The product graph generation unit 60 generates a product graph based on the generated feature / rating information (S06: Generate product graph). Specifically, the product graph generation unit 60 causes the vector information calculation unit 61 to calculate vector information (i.e., graph vector information) based on the feature / rating information, and the distance calculation unit 62 to calculate the distance between products based on the vector information. Then, the graph generation unit 63 generates a product graph for the same product category and / or a product graph for multiple product categories based on the calculated distance.

[0147] The information addition unit 70 associates the generated feature / rating information and / or product graph information with the product corresponding to the feature / rating information (S07: Addition of feature / rating information and / or product graph information). As a result, the feature / rating information and / or product graph information, which are additional information, are added to the merged product master, and comprehensive information about each product can be obtained from the merged product master.

[0148] In the above, the product graph information is associated with the name-matched product master, but this association is not necessary. Also, instead of the product graph, the vector information calculated by the vector information calculation unit 61 may be associated with the corresponding product and name-matched product master by the information addition unit 70.

[0149] [2-2. Name-matching product master generation operation] 18 is an example of an operation flowchart of the name-matched product master data generation device of this embodiment. Here, the name-matched product master data generation device 10 is configured to include a duplicate product determination unit 30 and an integration generation unit 40, and product data obtained by the product data acquisition unit 20, etc. is input to the name-matched product master data generation device 10. Furthermore, the product data is assumed to be two or more pieces of product data (HTML data) obtained from one or more e-commerce sites, but as mentioned above, it is not limited to this and a product master data used within an organization may also be obtained.

[0150] First, the pattern generation unit 31 generates a plurality of commodity identification latent patterns for each of a plurality of commodity data based on the commodity identification latent information (S21: Generate a plurality of commodity identification latent patterns). Specifically, the pattern generation unit 31 identifies and extracts the commodity identification latent information included in the commodity data using the program 311, AI, a learned model, or a combination thereof (S211: Identify and extract commodity identification latent information). Then, the pattern generation unit 31 identifies one or more pieces of unit information included in the commodity identification latent information using the program 311, AI, a learned model, or a combination thereof (S212: Identify unit information). Furthermore, the pattern generation unit 31 eliminates and / or converts predetermined unit information included in the commodity identification latent information based on the dictionary 313 or AI (S213: Eliminate and / or convert predetermined unit information). The predetermined unit information to be eliminated may be, for example, redundant information, and the unit information to be converted may be, for example, the manufacturer name or product name of the commodity. Conversions include conversion from half-width to full-width, full-width to half-width, English to Japanese, Japanese to English, manufacturer name, product name abbreviation, and nickname to official name.

[0151] In this way, after removing and / or converting predetermined unit information, multiple product identification latent patterns are generated based on the converted unit information and other unit information included in the product identification latent information (S214: Generate multiple product identification latent patterns). The product identification latent pattern can be determined based on the granularity of the identity of the product desired by the user. For example, if the granularity of identity is at the product name (brand name) level, i.e., if products are determined to be the same if the product names are the same, the pattern generation unit 31 may generate at least one pattern from the first to sixth patterns. If the granularity of identity is at the product identification code level, such as a JAN code including the sales format, i.e., if products are determined to be the same if the product names, manufacturer names, and various specification information are the same, the pattern generation unit 31 generates at least the first to seventh patterns. In order to improve the accuracy of identity determination, it is preferable to generate multiple seventh patterns related to the specification information.

[0152] Next, the vector information generation unit 32 generates vector information of each product identifying latent pattern for each product data (S22: Generate vector information). Specifically, the vector information generation unit 32 uses a known tool to convert the character string of the product identifying latent pattern into vector information, which is a collection of N-dimensional numerical values.

[0153] The similarity calculation unit 33 calculates the similarity between the two products for each product identification latent pattern based on the vector information of the product data of the two products (S23: Calculate similarity). The similarity here is a cosine similarity, and the closer the similarity is to 0, the less similar the products are, and the closer the similarity is to 1, the more similar the products are.

[0154] The determination unit 34 determines whether the two products are the same based on the similarity (S24: Are they the same product?). The determination unit 34 determines whether the products are the same based on whether the similarity of each pattern is equal to or greater than a predetermined threshold. The predetermined threshold and the number and type of patterns that are determined to be the same product can be determined according to the granularity of product identity desired by the user.

[0155] In one example, if the product names are the same, the products are determined to be the same product; if the similarity of at least one of the first to sixth patterns is above a predetermined threshold (e.g., 0.7, 0.8, or 0.9), the products are determined to be the same product; and if the similarity of all patterns is below the predetermined threshold, the products are determined to be different products.

[0156] In another example, if products are determined to be the same if the product name, manufacturer name, and various specification information are the same, they are determined to be the same product if all of the similarities of the first to seventh patterns are above a predetermined threshold (for example, 0.7, 0.8, or 0.9), and they are determined to be different products if any of the patterns are below the predetermined threshold.

[0157] If the determination unit 34 determines that the two items are not the same item (NO in S24), the process returns to the input of item data prior to S21 (e.g., S01). If the determination unit 34 determines that the items are the same item (YES in S24), the integrating and generating unit 40 classifies the item information included in the two items of item data determined to be the same by the determination unit 34 by item information category (S03), and integrates the information to generate a merged item master (S04). S03 and S04 are the same as those in FIG. 17, so their explanation will be omitted. Note that in S04, the integrating and generating unit 40 may eliminate duplicates of item information for each divided category and / or normalize the representation of the item information for each divided category.

[0158] [3. Actions and Effects] (1) The product name matching system 1 of this embodiment includes a product data acquisition unit 20 that acquires multiple product data, an identical product determination unit 30 that analyzes the multiple product data and identifies multiple product data related to the same product, an integration generation unit 40 that analyzes the multiple product data, classifies product information included in the product data by category of the product information, and for each product, integrates product information included in multiple product data related to the same product by category of the product information to generate a matched product master, a feature / evaluation information generation unit 50 that generates product feature / evaluation information from data indicating product features included in the product data, and an information addition unit 70 that associates additional information including feature / evaluation information related to the same product with the same product.

[0159] This saves the effort and time required to transmit product information and organize data, and the product information or additional information can provide product-related information necessary for marketing analysis such as demand forecasting, product development, or product recommendations.

[0160] (2) The characteristic / evaluation information generation unit 50 includes an identification unit 51 that identifies words that indicate the characteristics of the product from text data that indicates the characteristics of the product included in the product data, and a generation unit 52 that generates characteristic / evaluation information in association with the words.

[0161] This allows product features from various perspectives or angles to be stored as data associated with the aggregated product master, making it possible to provide product-related information from these perspectives or angles. This, for example, can improve product searchability and provide product-related information necessary for marketing analysis, product development, or product recommendations. In one example, rather than limiting product groups to wine, it is possible to extract products from other categories, such as wine, sake, and cocktails, enabling a wide range of analysis and recommendations. In another example, on an e-commerce site using the aggregated product master, product searchability can be improved by using feature / rating information as search tags.

[0162] (3) The feature / evaluation information generating unit 50 has a category estimating unit 53 that estimates the category of product information corresponding to a word, and the generating unit 52 generates feature / evaluation information in association with the word and the category.

[0163] This enables classification of categories for words identified by the identification unit 51. For example, if the product is a refreshing beverage, the word "refreshing" can be classified into categories such as impression and flavor. Furthermore, the categories can be hierarchized according to the size of the concept. As a result, the aggregated product master database can be searched using feature / evaluation information of various granularities as search keys, thereby improving search accuracy and generating product graphs from various angles, leading to new insights for the user.

[0164] (4) The feature / evaluation information is information included in at least one of the product information categories of the product's impression, atmosphere, taste, texture, quality, and use. This allows the product's features to be presented from a subjective or sensory perspective, which can provide hints for marketing analysis, product development, or product recommendations.

[0165] (5) A product graph generation unit 60 is provided that generates a product graph showing the relationships between products based on the feature / rating information, and the additional information includes a product graph or graph vector information for generating the product graph. This allows the relationships between products to be presented to the user, which can be used for marketing analysis such as demand forecasting, product development, or product recommendations, and can provide users with new insights about products.

[0166] (6) The product graph generation unit 60 generates a product graph for the same product category. This makes it possible to present the relationships between products in the same product category, making it easier to analyze a group of products in the same product category.

[0167] (7) The product graph generation unit 60 generates a product graph between multiple product categories. This makes it possible to present cross-sectional relationships between products, not limited to a specific product category, and provides users with new insights that would not be possible with a single product category. For example, by presenting products with the same taste or texture in the product graph, even across multiple product categories, retailers can be provided with hints for shelf allocation and product ordering, and manufacturers can be provided with hints for product development. Furthermore, online retailers can recommend products with a consistent taste or texture.

[0168] (8) The product graph generation unit 60 includes a vector information calculation unit 61 that calculates vector information based on feature / rating information, and a distance calculation unit 62 that calculates the distance between products based on the vector information. This makes it possible to prepare materials for generating various product graphs.

[0169] (9) Any of the multiple product data includes at least one of the product name, specification information, product logistics information, transaction information, customer information, and purchase information. This allows for automatic linking of product information and additional information with at least one of the logistics information, transaction information, customer information, and purchase information for the same product, eliminating the effort required to acquire and input information. Furthermore, the inclusion of product logistics information, transaction information, customer information, or purchase information in the aggregated product master enables more advanced marketing analysis, product development, and product recommendations.

[0170] For example, if product data includes purchasing information, it is possible to identify best-selling products from the purchasing information and then analyze whether the product information or additional information contained in the aggregated product master is responsible for the product's success. Based on the results of this analysis, retailers can be provided with hints for shelf allocation and product ordering, and manufacturers can be provided with hints for product development. For example, POS data, which includes product names and sales data for those products, is known as purchasing information. However, POS data does not include product information such as product specifications. Collecting product information and additional information and entering it into the product master requires time and effort. This embodiment eliminates this collection and effort. Furthermore, sales trends can be identified from purchasing information. Because the aggregated product master contains product information and additional information from various perspectives, the commonalities between the product information and additional information of best-selling products can be analyzed.

[0171] In addition, by including all product name, specification information, product logistics information, transaction information, customer information, and purchasing information, it is possible to integrate all information related to the product, from material procurement to manufacturing and distribution.

[0172] (10) The multiple product data includes product data from two or more different organizations. This allows a unified product master to be obtained across different organizations, making it possible to quickly and appropriately share product information between different organizations without the effort of manually inputting product information and additional information.

[0173] (11) The identical product determination unit 30 includes a pattern generation unit 31 that generates multiple product identification latent patterns for each of multiple product data based on product identification latent information including at least product names included in the product data, a vector information generation unit 32 that generates vector information for each product identification latent pattern for each of the product data, a similarity calculation unit 33 that calculates the similarity between two products based on the vector information of the product data for the two products for each product identification latent pattern, and a determination unit 34 that determines whether the two products are identical based on the similarity. This makes it possible to generate a name-matched product master.

[0174] (12) The integrating / generating unit 40 integrates the product information included in two pieces of product data that have been determined to be identical by the determining unit, for each category of the product information, and generates a merged product master, which is a product master for the identical products. This allows multiple product data to be automatically merged, thereby saving the effort and time required to integrate product data between different organizations and facilitating the transmission of product information.

[0175] (13) The pattern generation unit 31 identifies one or more unit information included in the commodity identification latent information and generates multiple commodity identification latent patterns based on the unit information. This generates multiple commodity identification latent patterns based on the smallest unit of information included in the commodity identification latent information, making it possible to determine whether the products are the same from multiple perspectives and improving the accuracy of determining whether the products are the same. In other words, it is possible to generate multiple commodity identification latent patterns in accordance with the actual handling practices of the products, thereby improving the accuracy of the determination.

[0176] (14) The pattern generation unit 31 removes or converts predetermined unit information included in the commodity identification latent information, and generates a commodity identification latent pattern based on the converted unit information and other unit information included in the commodity identification latent information after the removal. This makes it possible to remove unit information included in the commodity identification latent information that does not affect the same commodity determination, or to convert spelling variations that affect the same commodity determination, thereby improving the accuracy of determining whether the commodity is the same or not.

[0177] (15) The pattern generation unit 31 converts the abbreviation of a product, which is unit information included in the product identification latent information, into the official name of the product. This improves the accuracy of determining whether or not two products are the same. In particular, certain businesses, such as wholesalers, have a practice of entering the abbreviation of a product, rather than the official name, into the product master data for that product, and identifying products by the abbreviation. When a product master data set, which is product data for an industry with such a practice, is acquired, generating a product identification latent pattern based on the abbreviation of the product may reduce the accuracy of determining whether or not two products are the same. In contrast, in this embodiment, the abbreviation of a product is converted into the official name of the product, thereby improving the accuracy of determining whether or not two products are the same.

[0178] (16) The product identification latent information further includes product specification information, and the pattern generation unit 31 generates a product identification latent pattern that further includes specification information. This improves the accuracy of identifying identical products. For example, even if a product has the same name and is bottled, tea may be sold in various sales formats, such as a 500 ml bottle sold individually, a box of 12 bottles, or a 2 L bottle sold individually. Each product is assigned a product identification code, such as a JAN code, appropriate for its sales format. In other words, even if the brand name is the same, different sales formats identify the products as different. Even in such cases, it is possible to determine whether the products are identical with the accuracy of the product identification code. In other words, it is possible to determine whether the products are identical with a finer granularity than determining whether the products are identical based on the manufacturer or brand name, regardless of specification information such as volume, quantity, weight, size, or container type, thereby improving the accuracy of the determination.

[0179] (17) The integrating unit 40 treats the specification information as product information, classifies and integrates the product information by category, and generates a merged product master. This allows all specification information to be aggregated into the merged product master.

[0180] (18) The multiple product identification latent patterns include at least one of a first pattern obtained by excluding or converting predetermined unit information from the product identification latent information and excluding specification information included in the product identification latent information, a second pattern obtained by converting the first pattern into another notation format, a third pattern obtained by excluding the manufacturer name included in the product identification latent information from the first pattern, a fourth pattern obtained by excluding or converting predetermined unit information from the product identification latent information and excluding numerical information included in the product identification latent information, a fifth pattern including only the manufacturer name included in the product identification latent information, a sixth pattern including only the product name included in the product identification latent information, and a seventh pattern including only the product specification information included in the product identification latent information. This makes it possible to improve the accuracy of identifying identical products in accordance with the method of transmitting product information in multiple organizations.

[0181] (19) The integrating unit 40 eliminates duplication of product information for each category during integration. This makes the name-aggregated product master easier to use.

[0182] (20) The integrating unit 40 normalizes the representation of product information for each category during integration, making the name-aggregated product master easier to use.

[0183] (21) Of the product data of two products, one is reference product data that serves as the standard for determining whether the products are identical, and the other is undetermined product data that has not been determined by the determination unit 34. When the determination unit 34 determines that the product indicated by the undetermined product data and the product indicated by the reference product data are the same product, the undetermined product data is set as one of the reference product data. This allows for the accumulation of training data (correct answer data) that can be used for learning a learning model, such as excluding and converting unit information.

[0184] 4. Other Embodiments In other embodiments of the present invention, the present invention may be a program that realizes the functions of the embodiments of the present invention described above and the information processing shown in the flowcharts, or a computer-readable storage medium that stores the program.In still other embodiments, the present invention may be a method that realizes the functions of the embodiments of the present invention described above and the information processing shown in the flowcharts.In still other embodiments, the present invention may be a server that can supply a program that realizes the functions of the embodiments of the present invention described above and the information processing shown in the flowcharts to a computer.In still other embodiments, the present invention may be a virtual machine that realizes the functions of the embodiments of the present invention described above and the information processing shown in the flowcharts.

[0185] In the processes or operations described above, the processes or operations can be freely changed as long as no inconsistencies in the processes or operations occur, such as the use of data that should not yet be available in a certain step. Furthermore, the embodiments described above are merely examples for explaining the present invention, and the present invention is not limited to these embodiments. The present invention can be embodied in various forms without departing from the spirit of the invention.

[0186] In the above embodiment, data and information related to products are handled, but services may be handled instead of products. For example, service data related to services may be handled instead of product data. In this case, the functions of each unit of the system 1 and the device 10 can be replaced with functions related to services rather than products. [Explanation of symbols]

[0187] 1 Product name matching system 1a processor 1b Storage device 1b-1 RAM 1b-2 ROM 1b-3 External memory 1c Input device 1d output device 1e Communication equipment 1F Bus 10. Name-matching product master generation device 20 Product Data Acquisition Department 30 Same product determination department 31 Pattern generation section 311 Program 313 Dictionary 313a Exclusion Dictionary 313b Conversion Dictionary 32 Vector information generation unit 33 Similarity calculation unit 34 Judgment section 40 Integrated generation section 41 AI 42 Normalization Dictionary 43 AI for normalization 50 Feature / Evaluation Information Generation Unit 51 Specific section 52 Generation part 53 Category Estimation Unit 53a Dictionary 53b Natural Language Processing Library 60 Product graph generation unit 61 Vector information calculation unit 62 Distance calculation unit 63 Graph Generation Unit 70 Information Addition Section C10~C17 Product information category (column) G0 Product Image G1 Product identification latent information G2 Spec Information G3 Information showing product features / ratings M Name-matched product master P1~P7 1st pattern~7th pattern V1~V7 Vector information for the 1st to 7th patterns

Claims

1. a product data acquisition unit that acquires a plurality of product data; an identical product determination unit that analyzes the plurality of product data and identifies the plurality of product data related to the same product; an integration / generation unit that analyzes the plurality of product data, classifies product information included in the product data by category of the product information, and integrates, for each of the identical products, the product information included in the plurality of product data related to the identical product by category of the product information to generate a merged product master; a characteristic / evaluation information generating unit that generates characteristic / evaluation information indicating a characteristic or evaluation of the product from data indicating the characteristic of the product included in the product data; an information addition unit that associates additional information including the feature / evaluation information related to the identical product with the identical product in the name-matched product master; a product graph generation unit that generates a product graph showing the relationships between the products based on the feature / rating information; Equipped with the additional information includes the product graph or graph vector information for generating the product graph; Product name matching system.

2. The feature / evaluation information generation unit an identification unit that identifies words that indicate characteristics of the product from text data that indicates characteristics of the product and is included in the product data; a generation unit that generates the feature / evaluation information in association with the word; having The product name matching system according to claim 1.

3. The feature / evaluation information generation unit a category estimation unit that estimates a category of the product information corresponding to the word; the generation unit generates the feature / evaluation information in association with the word and the category. The product name matching system according to claim 2.

4. The characteristic / evaluation information is information included in at least one of the product information categories of the impression, atmosphere, taste, texture, quality, and use of the product. The product name matching system according to claim 3.

5. the product graph generation unit generates the product graph of the same product category; The product name matching system according to claim 1.

6. the product graph generation unit generates the product graph among a plurality of product categories; The product name matching system according to claim 1 or 5.

7. The product graph generation unit a vector information calculation unit that calculates vector information based on the feature / evaluation information; a distance calculation unit that calculates the distance between the products based on the vector information; having The product name matching system according to any one of claims 1 to 6.

8. Any of the plurality of product data includes at least one of a product name, specification information, logistics information of the product, transaction information, customer information, and purchase information. The product name matching system according to any one of claims 1 to 7.

9. The plurality of product data includes product data of two or more different organizations. The product name matching system according to any one of claims 1 to 8.

10. The identical product determination unit a pattern generating unit that generates a plurality of commodity identification latent patterns for each of the plurality of commodity data based on commodity identification latent information including at least a commodity name included in the commodity data; a vector information generating unit that generates vector information of each of the commodity identification latent patterns for each of the commodity data; a similarity calculation unit that calculates a similarity between the two products based on vector information of the product data of the two products for each of the product identification latent patterns; a determination unit that determines whether the two products are the same based on the similarity; having The product name matching system according to any one of claims 1 to 9.

11. the integration / generation unit integrates the product information included in the product data of the two products determined to be identical by the determination unit for each category of the product information, and generates a merged product master that is a product master of the identical products. The product name matching system according to claim 10.

12. A product data acquisition unit that acquires multiple product data; an identical product determination unit that analyzes the plurality of product data and identifies the plurality of product data related to the same product; an integration / generation unit that analyzes the plurality of product data, classifies product information included in the product data by category of the product information, and integrates, for each of the identical products, the product information included in the plurality of product data related to the identical product by category of the product information to generate a merged product master; a characteristic / evaluation information generating unit that generates characteristic / evaluation information indicating a characteristic or evaluation of the product from data indicating the characteristic of the product included in the product data; an information addition unit that associates additional information including the feature / evaluation information related to the identical product with the identical product in the name-matched product master; Equipped with The identical product determination unit a pattern generating unit that generates a plurality of commodity identification latent patterns for each of the plurality of commodity data based on commodity identification latent information including at least a commodity name included in the commodity data; a vector information generating unit that generates vector information of each of the commodity identification latent patterns for each of the commodity data; a similarity calculation unit that calculates a similarity between the two products based on vector information of the product data of the two products for each of the product identification latent patterns; a determination unit that determines whether the two products are the same based on the similarity; having Product name matching system.

13. a product data acquisition step of acquiring a plurality of product data; an identical product determination step of analyzing the plurality of product data and identifying the plurality of product data relating to the same product; an integration generation step of analyzing the plurality of product data, classifying product information included in the product data by category of the product information, and integrating the product information included in the plurality of product data related to the same product for each of the same products by category of the product information to generate a merged product master; a characteristic / evaluation information generating step of generating characteristic / evaluation information indicating a characteristic or evaluation of the product from data indicating the characteristic of the product included in the product data; an information addition step of associating additional information including the feature / evaluation information related to the identical product with the identical product in the name-matched product master; a product graph generation step of generating a product graph showing the relationships between the products based on the feature / rating information; The computer executes the additional information includes the product graph or graph vector information for generating the product graph; How to generate a merged product master.

14. A program that causes a computer to execute the method according to claim 13.

Citation Information

Patent Citations

  • Recommendation method and system based on commodity knowledge graph feature learning

    CN111369318A

  • Surface treating method and tool

    JP1989027850A

  • Information processor, information processing method, and program

    JP2010079657A

  • Commodity web page analyzer, commodity web page analysis method, and program for commodity web page analyzer

    JP2013101415A

  • Computer-assisted name identification device, computer-assisted name identification system, method, and program

    JP2015130040A